A method, device and computer-readable storage medium for node partitioning of a computational graph
By dividing the nodes supporting machine learning processors into subgraphs in the calculation graph and processing unsupported nodes separately, the problem of individual calculations of each node in the calculation graph in the prior art is solved, and more efficient calculation and optimization are achieved.
Patent Information
- Application Number
- CN202010340740.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-04-26
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2040-10-03
AI Technical Summary
In the prior art, machine learning processors in deep learning networks cannot support all operators, resulting in each node in the calculation graph being individually calculated, resulting in high input and output overhead and the underlying layer being unable to optimize.
By dividing nodes that support machine learning processors and can be fused into the corresponding subgraph in the calculation graph, and processing unsupported nodes separately, a convex subgraph division method with the largest weak connectivity is used to realize fusion computing.
Reduces input and output overhead, improves computing efficiency, allows the underlying layer to perform global optimization, and enhances performance and computing speed.
Smart Images

Figure CN113553287B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computers, and more specifically, to a method, device, and computer-readable storage medium for node partitioning of a computational graph. Background Art
[0002] Pytorch is a popular deep learning framework. Since there are operators in the deep learning network that cannot be supported by the machine learning processor, it is impossible to guarantee that the complete network model can run on the machine learning processor. In the computational graph of the existing technology, each node or operator needs to be calculated one by one, resulting in high input and output overhead, and the bottom layer cannot be optimized in conjunction with the context.
[0003] Under such circumstances, there is an urgent need in the art to solve at least some of the above-mentioned technical problems. Summary of the invention
[0004] The inventors of the present disclosure recognize that in a computation graph, if fusion computation optimization is not performed, the network computation speed will be limited. If the fusion mode of a machine learning processor is to be called for a Pytorch network, the network cannot be fused into a fusion node or operator because there are nodes or operators that are not supported by the machine learning processor.
[0005] The contradictions in the prior art determine that an algorithm needs to be used to select a set of nodes or operators that can be fused in the network, and the set of nodes or operators respectively fuses multiple nodes or operators.
[0006] According to a first aspect of the present disclosure, a method for node partitioning of a computational graph is provided, which may include:
[0007] Nodes that support machine learning processors and can be fused together are divided into corresponding subgraphs; nodes that do not support machine learning processors are processed separately.
[0008] According to a second aspect of the present disclosure, a node partitioning device corresponding to an operator is provided, which may include: a processor configured to execute program instructions; and a memory configured to store program instructions, so that when the program instructions are loaded and executed by the processor, the device executes the above method.
[0009] According to a third aspect of the present disclosure, a computer-readable storage medium is provided, in which program instructions are stored, and the program instructions are suitable for being loaded by a processor and executing the above method.
[0010] With the help of the above technical solution, nodes or operators that support machine learning processors and can be fused together can be divided into corresponding subgraphs. After fusion, the bottom-level optimization can take a global view. It avoids calculating each node (operator) one by one, reduces input and output (IO) overhead, and the bottom-level can also be optimized in context. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above features of the present disclosure may be better understood, and its numerous objects, features and advantages may become apparent to those skilled in the art by taking into account the accompanying drawings, wherein like reference numerals represent like elements, wherein:
[0012] Figure 1 is a schematic diagram showing a method for dividing nodes of a computational graph according to an embodiment of the present disclosure;
[0013] Figure 2 is a schematic diagram showing a node calculation method in the prior art;
[0014] Figure 3A is a schematic diagram showing a method for dividing nodes of a computational graph according to an embodiment of the present disclosure;
[0015] Figure 3B is a schematic diagram showing a method for dividing nodes of a computational graph according to an embodiment of the present disclosure;
[0016] Figure 3C is a schematic diagram showing a method for dividing nodes of a computational graph according to an embodiment of the present disclosure;
[0017] Figure 3D is a schematic diagram showing a method for dividing nodes of a computational graph according to an embodiment of the present disclosure;
[0018] Figure 4A is a schematic diagram showing a method for dividing nodes of a computational graph according to an embodiment of the present disclosure;
[0019] Figure 4B is a schematic diagram showing a method for dividing nodes of a computational graph according to an embodiment of the present disclosure;
[0020] Figure 4C is a schematic diagram showing a method for dividing nodes of a computational graph according to an embodiment of the present disclosure;
[0021] Figure 4D is a schematic diagram showing a method for dividing nodes of a computational graph according to an embodiment of the present disclosure;
[0022] Figure 5 A structural diagram schematically showing a combined processing device according to an embodiment of the present disclosure; and
[0023] Figure 6 The schematic diagram shows the structure of a board according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0024] Regarding the terms "operator" and "node" mentioned in the embodiments of the present disclosure, it should be noted that the term "operator" is from the computer computing level (or from the software level or algorithm level), and the term "node" is a more figurative term (from the graphics level or a more intuitive level). In terms of the content referred to, the terms "operator" and "node" are actually the same thing. That is, in the embodiments of the present disclosure, it can be considered that the terms "operator" and "node" have the same meaning and are described from different aspects.
[0025] The term "maximum weakly connected convex subgraph" appears in the embodiments of the present disclosure, the purpose is to divide the original graph with as few weakly connected convex subgraphs as possible. The definition of "weak connectivity" is that if all directed edges of a directed graph are replaced by undirected edges, and the resulting undirected graph is a connected graph, then the directed graph has weak connectivity. The definition of "convex subgraph" is that for any two nodes in the subgraph, any path between them only passes through the nodes in the subgraph and does not pass through other nodes in the original graph, then this is a convex subgraph.
[0026] By running the Pytorch network layer by layer, each "operator" is regarded as a "node" of a computational graph, and the direction of data flow is regarded as a directed edge. The entire static acyclic directed computational graph composed of these directed edges is traversed (the C++ representation can be obtained at the bottom layer in Pytorch) to obtain the running device information of each node, that is, whether the running device of each node is a running device that supports machine learning processors, and if so, which machine learning processor it is.
[0027] In one embodiment of the present disclosure, Figure 1 1 is a schematic diagram showing a method 100 for partitioning nodes of a computation graph according to an embodiment of the present disclosure. The method 100 for partitioning nodes of a computation graph may include the following steps, for example, the following steps are performed by a processor:
[0028] Step 102, the nodes that support the machine learning processor and can be fused together are divided into corresponding subgraphs; the inventor divides the subgraphs for fusion according to the above information (the operating device information of each node). The premise that the divided subgraphs can support the machine learning processor and can be fused calculations is: (1) for any two nodes AB in the fusion subgraph, there is no calculation path from A to B passing through nodes outside the subgraph, (2) the fusion subgraph has weak connectivity, that is, if the subgraph is converted into an acyclic graph, it is a connected graph.
[0029] Step 104, the nodes that do not support the machine learning processor are processed separately.
[0030] In one embodiment of the present disclosure, the machine learning processor may include one or more machine learning processors. For example, the machine learning processor may include a first machine learning processor and a second machine learning processor, etc.
[0031] In one embodiment of the present disclosure, step 102 of dividing the nodes that support the machine learning processor and can be fused together into corresponding subgraphs may include:
[0032] Step 106: Divide the nodes that support the machine learning processor and can be fused together into the convex subgraph corresponding to the maximum weak connection.
[0033] In one embodiment of the present disclosure, step 106, dividing the nodes that support the machine learning processor and can be fused together into the corresponding convex subgraph with the largest weak connection may include:
[0034] Step 108, taking a node without a predecessor node as the starting node, traversing the successor nodes of the starting node, if there is no other path back to the starting node during the reverse traversal from the successor node to the starting node, the starting node and the successor node are divided into the same subgraph. Because the fusion process is carried out step by step, there may be multiple successor nodes in the fusion process from the starting node to the successor node.
[0035] Optionally, the successor node of the starting node includes one or more successor nodes.
[0036] Optionally, when a starting node includes multiple successor nodes, when traversing the successor nodes of the starting node, the traversal can be carried out in sequence according to the logical order of the operators corresponding to the successor nodes. Optionally, the traversal can also be carried out in the reverse order of the logical order of the operators corresponding to the successor nodes. The present disclosure does not impose any restrictions on this.
[0037] Optionally, the reverse traversal from the successor node to the start node refers to a reverse depth traversal.
[0038] Optionally, when there is no other path returning to the starting node during the reverse traversal from the successor node to the starting node, the starting node and the successor node are divided into the same subgraph, which means that each successor node of the starting node is visited in turn, and each visited successor node is traversed in reverse. If there is no other path returning from the successor node to the starting node, the successor node and the starting node are divided into the same subgraph, that is, each successor node is judged individually, and the path from the successor node to the starting node in reverse is unique.
[0039] In one embodiment of the present disclosure, step 106, dividing the nodes that support the machine learning processor and can be fused together into the corresponding convex subgraph with the largest weak connection, may also include:
[0040] Step 110, merge the subgraphs into a new starting node, traverse the successor nodes of the new starting node along the directed edges starting from the new starting node, and if there is no other path back to the new starting node during the reverse traversal from the successor nodes of the new starting node to the new starting node, divide the new starting node and the successor nodes of the new starting node into a new subgraph, and repeat the above steps until all nodes that support the machine learning processor and can be fused together are divided into the same maximum weakly connected convex subgraph.
[0041] In one embodiment of the present disclosure, step 106, dividing the nodes that support the machine learning processor and can be fused together into the corresponding convex subgraph with the largest weak connection may include:
[0042] Step 112, taking a node without a predecessor node as the starting node, traversing the successor node corresponding to the starting node along the directed edge, in the case where there are other paths to return to the starting node during the reverse traversal process from the successor node to the starting node, the access from the starting node to the successor node is terminated. It should be noted that in the case where there are other paths to return to the starting node during the reverse traversal process from the successor node to the starting node, the access from the starting node to the successor node is terminated, which means that each successor node of the starting node is visited in turn, and each visited successor node is reversely traversed. If there are other paths to return to the starting node from the successor node, the access from the starting node to the successor node is terminated. That is, each successor node is judged separately, and the situation that there are other paths to return to the starting node means that in the reverse traversal process from the successor node to the starting node, other paths are experienced, that is, the path from the successor node to the starting node is not unique. The situation that the path from the successor node to the starting node is not unique will be specifically described in the following embodiments 1 and 2, and will not be repeated here.
[0043] In one embodiment of the present disclosure, step 106, dividing the nodes that support the machine learning processor and can be fused together into the corresponding convex subgraph with the largest weak connection may include:
[0044] Step 114, taking the successor node as the new starting node, traversing the successor nodes of the new starting node along the directed edges starting from the new starting node; in the case that there is no other path returning to the new starting node during the reverse traversal from the successor node of the new starting node to the new starting node, the new starting node and the successor nodes of the new starting node are divided into the same subgraph; in the case that there is other path returning to the new starting node during the reverse traversal from the successor node of the new starting node to the new starting node, the visit from the new starting node to the successor nodes of the new starting node is terminated.
[0045] Figure 2 Schematic diagram showing a node calculation method in the prior art. Figure 2 , Figures 3A-3D as well as Figures 4A-4D The nodes that support the machine learning processor are indicated by a solid circle "●", and the nodes that do not support the machine learning processor are indicated by a hollow circle "○". Figure 2 In the figure, node E is a node without a predecessor calculation, node A is the successor node of node E, node D and node B are the successor nodes of node A, and node C is the successor node of node D and node B. In the schematic diagram of the calculation method of the prior art, node E is calculated first, and then node A is calculated along the directed edge EA, and then starting from node A, two directed edges AD and AB are divided, and node D and node B are calculated respectively, and finally node C is calculated. In the case of the prior art, the input and output (IO) overhead is relatively large and the calculation efficiency is relatively low. For example, the output of node E is the input of node A, the output of node A is the input of node D and node B, and the output of node D and node B is the input of node C, and node D is a node that does not support the machine learning processor. When the calculation runs to node D along the directed edge AD, because node D does not support the machine learning processor, the data transmitted from node A needs to be copied to, for example, the central processing unit for processing. For example, after the central processing unit completes the processing, the processed data is transferred from node D to node C along the directed edge DC, or the data is copied from node D to node C.
[0046] The prior art does not consider fusing these nodes, such as node E, node A, etc., to improve computing efficiency and reduce IO overhead.
[0047] According to the concept of the present disclosure, a union-find set (or other similar data structures that can store nodes of the same type) can be used to record the segmentation results.
[0048] The steps of division can be as follows:
[0049] (1) Starting from all nodes without predecessors in the computation graph, breadth-wise traverse all remaining nodes in the graph (if there are nodes that are deleted in subsequent steps, such as node D, they are skipped).
[0050] (2) For each arbitrary node, such as node A, if the node is a node that supports a machine learning processor, such as one that can run on an MLU, all successor nodes starting from it are traversed (if there are directed edges added to this node in subsequent steps, they must also be traversed). It should be noted that nodes are connected to each other through directed edges. As mentioned in the embodiment, traversal along directed edges actually refers to traversal along directed edges from one node to the next node, or traversal from a starting node to its successor node.
[0051] (3) For the endpoint of the directed edge visited, such as node B, check whether another computational path can be found from A to B through reverse depth traversal (traversal from the endpoint of the directed edge to the starting point). If so, end the traversal of this directed edge. If not, merge nodes A and B into a set AB (indicating that they are in a fused subgraph), add all the incoming edges of node B, such as AB, and the outgoing edges, such as BC, to node A, delete node B and its directed edges, and complete the traversal of this edge.
[0052] The rationality of the node (algorithm) division disclosed in this disclosure lies in:
[0053] Correctness: Connectivity is a natural consequence of the algorithm’s process of adding nodes;
[0054] Convexity (for any two nodes A and B in the fusion subgraph, there is no computational path from node A to B that passes through nodes outside the subgraph) can be proved by contradiction;
[0055] Optimality: can be proved by contradiction;
[0056] Stability: The topological sorting of nodes in a directed graph is stable and uses a fixed traversal order.
[0057] Combine the following Figures 3A-3D and Figures 4A-4D Describe the steps of fusion.
[0058] It should be pointed out that in various embodiments of the present disclosure, the solid circle “●” represents a node that supports the machine learning processor, and is a node that needs to be considered for further integration later, while the hollow circle “○” represents a node that does not support the machine learning processor and is a node that needs to be processed separately. For example, the node that does not support the machine learning processor may be processed by a central processing unit, a voice processor, a programmable logic processor and / or an image processor.
[0059] Example 1
[0060] When performing fusion calculation, start from node F (node F is a node without a predecessor calculation node, that is, the starting node in this embodiment of the present disclosure), and there are two directed edges starting from node F, one is the directed edge FG leading to node G, and the other is the directed edge FH leading to node H. It can be considered that node G and node H are the successor nodes of node F, node I is the successor node of node H, node J is the successor node of node G and node I, and node K is the successor node of node J.
[0061] For example, taking node F without a predecessor node as the starting node, traversing the successor nodes along the directed edges is actually traversing from one node to its successor nodes. The traversal along the directed edges described in the specification of the present disclosure is actually the same as the traversal from one node to the next node: node G and node H. In the case where there is no other path back to the starting node (node F) during the reverse traversal from the successor nodes, such as nodes G and node H, to the starting node (node F), the starting node (node F) and the successor nodes (such as nodes G and node H) can be divided into the same subgraph {F, G, H}, as in Figure 3B As shown, or nodes F, G, H are merged into a set FGH. The subgraph {F, G, H} is used as the new starting node. When merging from the new starting node {F, G, H} to the successor node J and the successor node I, it is found that it is possible to pass through node I (such as in Figure 3B As shown in FIG. 1 , the node without a predecessor node (e.g., {F, G, H}) is used as the starting node, and the successor node J is traversed along the directed edge. In the case where there is another path (e.g., via node I) to return to the starting node (e.g., {F, G, H}) during the reverse traversal from the successor node J to the starting node (e.g., {F, G, H}), the access from the starting node (e.g., {F, G, H}) to the successor node J ends. As shown in FIG. Figure 3BAs shown. Optionally, the passed node may be a node that does not support the machine learning processor, or a node that supports the machine learning processor. If it is a node that supports the machine learning processor, it may be the same device as the starting node {F, G, H}, or it may be a different device. This application does not impose any restrictions on this. As long as there are other paths to return to the starting node during the reverse traversal of the successor node, it has nothing to do with the node device information in this path. Continue to visit another successor node I of the starting node. Because node I is a node that does not support the machine learning processor, that is, a node that needs to be processed separately, it is impossible to be divided into the same subgraph with the new starting node {F, G, H} or node J. In an optional embodiment, if node I is a node that supports a machine learning processor and there is no other path from the successor node I back to the starting node, then node I may be divided into a subgraph with the starting node {F, G, H} (this situation is not marked in the figure); in a possible embodiment, when the successor node I is a node that supports a machine learning processor and the device information of node I is consistent with the device information of the starting node {F, G, H}, then node I may be divided into a subgraph with the actual node (this situation is not marked in the figure).
[0062] At this point, we can consider nodes F, G, and H to be in the largest weakly connected convex subgraph, and node J can be considered a node without a predecessor calculation. Node J can be considered a new starting node for the subsequent fusion calculation. It should be noted that the successor node and the node without a predecessor calculation are relative. For example, Figure 3C In , the successor node J is the successor node of nodes F, G, and H. However, when the successor node J cannot be merged with nodes F, G, and H in the same subgraph, the successor node J becomes a node without a predecessor calculation in the subsequent fusion calculation, that is, a new starting node. That is, node J and the new starting node {F, G, H} are not divided into the same subgraph. At this time, node J can be used as the next new starting node to start the subsequent fusion step. For example, Figure 3C As shown in , the new starting node J and its successor node K can be divided into the same subgraph. This is because in the reverse traversal process from the successor node K to the starting node J, there is no other path back to the starting node J. Therefore, the starting node J and the successor node K can be divided into the same subgraph, or the nodes J and K are merged into a set JK. That is, it can be considered that nodes J and K are divided into another maximum weakly connected convex subgraph.
[0063] Finally, in Figure 3D As shown in , nodes F, G, and H form a maximally weakly connected convex subgraph, nodes J and K form another maximally weakly connected convex subgraph, and node I is a node that does not support the machine learning processor and needs to be processed separately.
[0064] exist Figures 3A to 3D In the embodiment of FIG. 5 , node K is the successor node of node J, but node K is also the only subsequent node of node J. In different implementations, the successor node may be unique or may not be unique.
[0065] Example 2
[0066] When performing fusion calculation, start from node F (node F is a node without a predecessor calculation node, that is, the starting node in this embodiment of the present disclosure), and there are two directed edges starting from node F, one is the directed edge FG leading to node G, and the other is the directed edge FH leading to node H. It can be considered that node G and node H are the successor nodes of node F, node L is the successor node of node H, node I is the successor node of node L, node J is the successor node of node G and node I, and node K is the successor node of node J.
[0067] For example, taking node F without a predecessor node as the starting node, traversing the successor nodes, such as node G and node H, along the directed edges FG and FH, when there is no other path back to the starting node (node F) during the reverse traversal from the successor nodes, such as node G and node H, to the starting node (node F), the starting node (node F) and the successor nodes (such as node G and node H) are divided into the same subgraph {F, G, H}, as in Figure 4A As shown. In other words, nodes F, G, and H are merged into a set FGH.
[0068] Take the subgraph {F, G, H} as the new starting node, merge the subgraph {F, G, H} into the new starting node, and traverse the successor nodes of the new starting node along the directed edges starting from the new starting node, that is, merge from the new starting node {F, G, H} to the successor node J and the successor node L. When there is no other path back to the new starting node {F, G, H} during the reverse traversal from the successor node L of the new starting node {F, G, H} to the new starting node {F, G, H}, the new starting node {F, G, H} and the successor node L of the new starting node are divided into the new subgraph {F, G, H, L}. In other words, the nodes F, G, H, L are merged into a set FGHL. FIG. 4A to FIG. 4DIn the embodiment, according to the logical order of execution of the operators corresponding to the nodes, it is assumed that node L has a higher processing order than node J, so it is first determined whether node L is to be fused. Repeat the above steps until all nodes that support the machine learning processor and can be fused together are divided into the same maximally weakly connected convex subgraph. In the process of traversing from the successor node J to the starting node {F, G, H, L} in reverse, it is found that it is possible to pass through the node I that does not support the machine learning processor. At this time, it can be considered that the successor node J and the new starting node {F, G, H, L} cannot be divided into the same subgraph. Because node I is a node that does not support the machine learning processor, that is, a node that needs to be processed separately, it is impossible to be divided into the same subgraph with the new starting node {F, G, H, L} or node J. At this time, it can be considered that nodes F, G, H, L are divided into the maximally weakly connected convex subgraph. That is, it can be considered that the subgraph {F, G, H, L} is a weakly connected convex subgraph larger than the subgraph {F, G, H}.
[0069] Taking a node without a predecessor node (e.g. {F, G, H, L}) as the starting node, traversing the successor node J along the directed edge, if there is another path (e.g. via node I) returning to the starting node (e.g. {F, G, H, L}) during the reverse traversal from the successor node J to the starting node (e.g. {F, G, H, L}), then the visit from the starting node (e.g. {F, G, H, L}) to the successor node J ends. Figure 4B As shown. This means that the nodes {F, G, H, L} and the successor node J cannot be merged into one word graph.
[0070] Node J and the new starting node {F, G, H, L} are not in the same subgraph. In this case, node J can be used as the next new starting node to start the subsequent fusion steps. Figure 4B As shown in , the new starting node J and its successor node K can be divided into the same subgraph. This is because in the reverse traversal process from the successor node K to the starting node J, there is no other path back to the starting node J, so the starting node J and the successor node K can be divided into the same subgraph. It can be considered that nodes J and K are divided into another maximally weakly connected convex subgraph. That is, with the successor node (such as node J) as the new starting node, along the directed edge starting from the new starting node, such as JK, traverse the successor nodes of the new starting node (such as node K);
[0071] In the case that there is no other path back to the new starting node during the reverse traversal from the successor node of the new starting node (e.g., node K) to the new starting node (e.g., node J), the new starting node (e.g., node J) and the successor node of the new starting node (e.g., node K) are divided into the same subgraph (e.g., {J, K}), or nodes J and K are merged into a set JK. Figure 4C shown.
[0072] Finally, in Figure 4D As shown in , nodes F, G, H, and L form a maximally weakly connected convex subgraph, nodes J and K form another maximally weakly connected convex subgraph, and node I is a node that does not support the machine learning processor and needs to be processed separately.
[0073] In one embodiment of the present disclosure, the subgraph and the maximum weakly connected convex subgraph may be directed graphs.
[0074] In one embodiment of the present disclosure, the node partitioning method is used in a Pytorch framework.
[0075] In one embodiment of the present disclosure, the partitioning result of the node partitioning method is stored in a union-find set, a common set or a linked list.
[0076] This algorithm, which was not originally available in Pytorch, can support CPU and MLU heterogeneous hybrid network computing. The original network can be divided to obtain the least and largest possible fusion subgraphs, and the operators supported by MLU can be fused as much as possible. Each subgraph meets the premise that it can be fused, thereby maximizing the performance enhancement through fusion computing and improving the computing speed, so as to achieve the goal of maximizing computing efficiency even if there are operators not supported by MLU.
[0077] Figure 5 is a structural diagram showing a combined processing device 500 according to an embodiment of the present disclosure. As shown in the figure, the combined processing device 500 includes a node division device 502 having the aforementioned computational graph, which can be configured to execute the node division method of the computational graph described in conjunction with the accompanying drawings. In one or more embodiments, the node division device 502 of the computational graph may also be a chip or an integrated circuit for calculating the gradient of input data. In addition, the combined processing device also includes a universal interconnection interface 504 and other processing devices 506. The node division device 502 of the computational graph according to the present disclosure can interact with other processing devices 506 through the universal interconnection interface 504 to jointly complete the operation specified by the user.
[0078] According to the scheme of the present disclosure, the other processing device may include one or more types of processors in general and / or special-purpose processors such as a central processing unit ("CPU"), a graphics processing unit ("GPU"), an artificial intelligence processor, etc., and the number thereof may not be limited but determined according to actual needs. In one or more embodiments, the other processing device may include the aforementioned benchmark hardware platform or benchmark computing device, so that it can form a test system with the node division device of the computational graph including the test hardware platform. In one or more embodiments, the other processing device can serve as an interface between the node division device of the computational graph of the present disclosure (which can be embodied as an artificial intelligence-related computing device) and external data and control, and perform basic control including but not limited to data handling, and complete the start and stop of the machine learning computing device; other processing devices can also cooperate with machine learning-related computing devices to jointly complete computing tasks.
[0079] According to the scheme of the present disclosure, the universal interconnection interface can be used to transmit data and control instructions between the node division device of the computation graph and other processing devices. For example, the node division device of the computation graph can obtain the required input data from other processing devices via the universal interconnection interface, and write it into the storage device (or memory) on the chip of the node division device of the computation graph. Further, the node division device of the computation graph can obtain control instructions from other processing devices via the universal interconnection interface, and write them into the control cache on the chip of the node division device of the computation graph. Alternatively or optionally, the universal interconnection interface can also read the data in the storage module of the node division device of the computation graph and transmit it to other processing devices.
[0080] Optionally, the combined processing device may further include a storage device 508, which may be connected to the node partitioning device of the computation graph and other processing devices, respectively. In one or more embodiments, the storage device may be used to store data of the node partitioning device of the computation graph and other processing devices, especially data that cannot be fully stored in the internal or on-chip storage device of the node partitioning device of the computation graph or other processing devices.
[0081] According to different application scenarios, the combined processing device disclosed in the present invention can be used as a SOC chip system for mobile phones, robots, drones, video surveillance equipment and other devices, effectively reducing the core area of the control part, improving the processing speed, and reducing the overall power consumption. In this case, the universal interconnection interface of the combined processing device is connected to certain components of the device. Certain components such as cameras, displays, mice, keyboards, network cards or wifi interfaces.
[0082] In some embodiments, the present disclosure further discloses a chip, which includes the node partitioning device or the combination processing device of the above-mentioned computation graph. In other embodiments, the present disclosure further discloses a chip packaging structure, which includes the above-mentioned chip.
[0083] In some embodiments, the present disclosure further discloses a board, which includes the above chip packaging structure. Figure 6 , which provides the aforementioned exemplary board card, which, in addition to the aforementioned chip 602 , may also include other supporting components, including but not limited to: a storage device 604 , an interface device 606 and a control device 608 .
[0084] The memory device is connected to the chip in the chip package structure through a bus for storing data. The memory device may include multiple groups of memory cells 610. Each group of memory cells is connected to the chip through a bus. It is understood that each group of memory cells may be DDR SDRAM ("Double Data Rate SDRAM, double rate synchronous dynamic random access memory").
[0085] DDR can double the speed of SDRAM without increasing the clock frequency. DDR allows data to be read out on the rising and falling edges of the clock pulse. The speed of DDR is twice that of standard SDRAM. In one embodiment, the storage device may include 4 groups of storage units. Each group of storage units may include multiple DDR4 particles (chips). In one embodiment, the chip may include 4 72-bit DDR4 controllers, of which 64 bits are used for data transmission and 8 bits are used for ECC verification.
[0086] In one embodiment, each group of storage units includes a plurality of double rate synchronous dynamic random access memories arranged in parallel. DDR can transmit data twice in one clock cycle. A controller for controlling DDR is arranged in the chip to control the data transmission and data storage of each storage unit.
[0087] The interface device is electrically connected to the chip in the chip packaging structure. The interface device is used to realize data transmission between the chip and an external device 612 (such as a server or a computer). For example, in one embodiment, the interface device can be a standard PCIE interface. For example, the data to be processed is transmitted to the chip by the server through the standard PCIE interface to realize data transfer. In another embodiment, the interface device can also be other interfaces. The present disclosure does not limit the specific manifestations of the above-mentioned other interfaces. The interface unit can realize the switching function. In addition, the calculation results of the chip are still transmitted back to the external device (such as a server) by the interface device.
[0088] The control device is electrically connected to the chip. The control device is used to monitor the state of the chip. Specifically, the chip and the control device can be electrically connected through an SPI interface. The control device may include a single-chip microcomputer (Micro Controller Unit, MCU). In one or more embodiments, the chip may include multiple processing chips, multiple processing cores or multiple processing circuits, which can drive multiple loads. Therefore, the chip can be in different working states such as multi-load and light load. The control device can realize the regulation of the working state of multiple processing chips, multiple processing and / or multiple processing circuits in the chip.
[0089] In some embodiments, the present disclosure further discloses an electronic device or apparatus, which includes the above-mentioned board. According to different application scenarios, the electronic device or apparatus may include a data processing device, a robot, a computer, a printer, a scanner, a tablet computer, a smart terminal, a mobile phone, a driving recorder, a navigator, a sensor, a camera, a server, a cloud server, a camera, a video camera, a projector, a watch, a headset, a mobile storage, a wearable device, a means of transportation, a household appliance, and / or a medical device. The means of transportation include airplanes, ships and / or vehicles; the household appliances include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, electric lights, gas stoves, and range hoods; the medical equipment includes an MRI, an ultrasound machine and / or an electrocardiograph.
[0090] According to different application scenarios, the node division device of the computation graph disclosed in the present invention, or the combined processing device containing the node division device of the computation graph, the chip for calculating the gradient of input data, and the corresponding computer-readable storage medium, the integrated circuit for calculating the gradient of input data can be applied to data processing devices, robots, computers, printers, scanners, tablet computers, smart terminals, mobile phones, driving recorders, navigators, sensors, cameras, servers, cloud servers, cameras, video cameras, projectors, watches, headphones, mobile storage, wearable devices, transportation tools, household appliances, and / or medical equipment and the like. Transportation tools include airplanes, ships and / or vehicles; household appliances include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, electric lights, gas stoves, and range hoods; medical equipment includes magnetic resonance imaging, ultrasound machines and / or electrocardiographs.
[0091] In addition, it should be pointed out that the term "data" mentioned in the application documents of the present disclosure should be understood in a broad sense and may include graphics, images, videos, audio, etc.
[0092] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present disclosure is not limited by the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present disclosure.
[0093] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0094] In the several embodiments provided in the present disclosure, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are only schematic, such as the division of units, which is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, optical, acoustic, magnetic or other forms.
[0095] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0096] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software program module.
[0097] If the integrated unit is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, when the technical solution of the present disclosure can be embodied in the form of a software product, the computer software product is stored in a memory, including a number of instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the various embodiments of the present disclosure. The aforementioned memory includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk and other media that can store program codes.
[0098] In the above embodiments of the present disclosure, the description of each embodiment has its own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant description of other embodiments. The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, all possible combinations of the technical features in the above embodiments are not described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0099] The foregoing content can be better understood in accordance with the following terms:
[0100] Clause A1. A method for partitioning nodes of a computational graph, comprising:
[0101] Divide nodes that support machine learning processors and can be fused together into corresponding subgraphs;
[0102] Nodes that do not support machine learning processors are treated separately.
[0103] Item A2. According to the method of Item A1, the machine learning processor includes one or more machine learning processors.
[0104] Clause A3. The method of clause A2, wherein dividing the nodes supporting the machine learning processor and capable of being fused together into corresponding subgraphs comprises:
[0105] Nodes that support machine learning processors and can be fused together are divided into convex subgraphs corresponding to the maximum weak connectivity.
[0106] Clause A4. According to the method of clause A3, the nodes that support the machine learning processor and can be fused together are divided into convex subgraphs corresponding to the maximum weak connectivity, including:
[0107] Take a node without a predecessor node as the starting node, traverse the successor nodes along the directed edges, and if there is no other path back to the starting node during the reverse traversal from the successor node to the starting node, the starting node and the successor node are divided into the same subgraph.
[0108] Clause A5. According to the method of clause A4, the nodes that support the machine learning processor and can be fused together are divided into convex subgraphs corresponding to the maximum weak connectivity, and further include:
[0109] The subgraphs are merged into a new starting node, and the successor nodes of the new starting node are traversed along the directed edges starting from the new starting node. When there is no other path returning to the new starting node during the reverse traversal from the successor nodes of the new starting node to the new starting node, the new starting node and the successor nodes of the new starting node are divided into a new subgraph, and the above steps are repeated until all nodes that support machine learning processors and can be fused together are divided into the same maximum weakly connected convex subgraph.
[0110] Clause A6. According to the method of clause A5, the nodes that support the machine learning processor and can be fused together are divided into convex subgraphs corresponding to the maximum weak connectivity, including:
[0111] Take a node without a predecessor node as the starting node, traverse the successor nodes along the directed edges, and if there is another path back to the starting node during the reverse traversal from the successor node to the starting node, the visit from the starting node to the successor node is completed.
[0112] Clause A7. According to the method of clause A6, the nodes that support the machine learning processor and can be fused together are divided into convex subgraphs corresponding to the maximum weak connectivity, including:
[0113] Take the successor node as the new starting node and traverse the successor nodes of the new starting node along the directed edge starting from the new starting node;
[0114] In the case that there is no other path returning to the new starting node during the reverse traversal from the successor node of the new starting node to the new starting node, the new starting node and the successor node of the new starting node are divided into the same subgraph;
[0115] In the case where there is another path returning to the new start node during the reverse traversal from the successor node of the new start node to the new start node, the visit from the new start node to the successor node of the new start node is completed.
[0116] Clause A8. A method according to clauses A1-7, wherein separately processing nodes that do not support a machine learning processor comprises:
[0117] Nodes that do not support machine learning processors will be processed by central processing units, voice processors, programmable logic processors, and / or image processors.
[0118] Item A9. A method according to Item A1-7, wherein the subgraph and the maximum weakly connected convex subgraph are directed graphs.
[0119] Item A10. The method according to Item A1-7, wherein the node partitioning method is used for the Pytorch framework.
[0120] Clause A11. According to the method of Clause A1-7, the partitioning results of the node partitioning method are stored in a union-find set, a common set or a linked list.
[0121] Clause A12. A node partitioning device for a computational graph may include:
[0122] a processor configured to execute program instructions; and
[0123] A memory configured to store program instructions which, when loaded and executed by a processor, cause the apparatus to perform a method according to clause A1-11.
[0124] Clause A13. A computer-readable storage medium having program instructions stored therein, the program instructions being suitable for being loaded by a processor and executing the method according to clause A1-11.
Claims
1. A method for node partitioning of a computational graph. include: Divide nodes that support machine learning processors and can be fused together into corresponding subgraphs; as well as Nodes that do not support machine learning processors will be processed separately; The nodes that support the machine learning processor and can be fused together are divided into corresponding subgraphs, including: Divide nodes that support machine learning processors and can be fused together into convex subgraphs corresponding to the largest weak connectivity; The partitioning of nodes that support the machine learning processor and can be fused together into convex subgraphs corresponding to the maximum weak connectivity includes: A node without a predecessor node is taken as a starting node, and the successor nodes of the starting node are traversed along directed edges. When there is no other path returning to the starting node during the reverse traversal from the successor nodes to the starting node, the starting node and the successor nodes are divided into the same subgraph.
2. The node partitioning method according to claim 1, wherein the nodes that support the machine learning processor and can be fused together are partitioned into the convex subgraph corresponding to the maximum weakly connected include: The subgraphs are merged into a new starting node, and the successor nodes of the new starting node are traversed along the directed edges starting from the new starting node. When there is no other path returning to the new starting node during the reverse traversal from the successor nodes of the new starting node to the new starting node, the new starting node and the successor nodes of the new starting node are divided into a new subgraph, and the above steps are repeated until all nodes that support the machine learning processor and can be fused together are divided into the same maximum weakly connected convex subgraph.
3. The node partitioning method according to claim 2, wherein the nodes that support the machine learning processor and can be fused together are partitioned into the convex subgraph corresponding to the maximum weakly connected include: Taking a node without a predecessor node as the starting node, traverse the successor nodes of the starting node along the directed edges. If there are other paths returning to the starting node during the reverse traversal from the successor nodes to the starting node, the access from the starting node to the successor nodes is terminated.
4. The node partitioning method according to claim 3, wherein the nodes that support the machine learning processor and can be fused together are partitioned into the convex subgraph corresponding to the maximum weakly connected include: Taking the successor node as a new starting node, traversing the successor nodes of the new starting node along the directed edge starting from the new starting node; In the case that there is no other path returning to the new starting node during the reverse traversal process from the successor node of the new starting node to the new starting node, the new starting node and the successor node of the new starting node are divided into the same subgraph; In the case where there is another path returning to the new starting node during the reverse traversal process from the successor node of the new starting node to the new starting node, the access from the new starting node to the successor node of the new starting node ends.
5. The node partitioning method according to claim 4, It is characterized in that The machine learning processor includes one or more machine learning processors, wherein nodes corresponding to the same type of machine learning processors can be integrated into the same subgraph.
6. The node partitioning method according to any one of claims 1 to 5, wherein the nodes that do not support the machine learning processor are processed separately. include: Nodes that do not support machine learning processors will be processed by central processing units, voice processors, programmable logic processors, and / or image processors.
7. A node partitioning method according to any one of claims 1 to 5, wherein the subgraph and the maximum weakly connected convex subgraph are directed graphs.
8. The node partitioning method according to any one of claims 1 to 5, wherein the node partitioning method is used for a Pytorch framework.
9. According to the node partitioning method according to any one of claims 1 to 5, the partitioning result of the node partitioning method is stored in a union-find set, a common set or a linked list.
10. A node partitioning device for a computational graph, include: a processor configured to execute program instructions; as well as A memory configured to store the program instructions, which, when loaded and executed by the processor, causes the apparatus to perform the method according to any one of claims 1 to 9.
11. A computer-readable storage medium storing program instructions, wherein the program instructions are suitable for being loaded by a processor and executing the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Neural network pruning method and device, computer equipment and storage medium
CN110689116A