Computation graph segmentation method for heterogeneous equipment

By defining the conditional table of computing subgraphs and traversing the computing graph from bottom to top, the problem that the computing graph slicing method in the prior art cannot effectively handle multiple outputs, and the efficient slicing of computing graphs on heterogeneous devices and accelerated inference of deep learning models is realized.

CN120068965APending Publication Date: 2025-05-30SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510189891.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing computational graph segmentation method cannot effectively solve the data dependence problem when facing multiple outputs, resulting in the inability to calculate the sub-graph, affecting the calculation efficiency and accuracy. At the same time, some methods need to add a large amount of additional label information, resulting in large workload and redundant workload.

Method used

A calculation graph segmentation method for heterogeneous devices is proposed. By obtaining the calculation graph corresponding to the deep learning model, defining the condition table of the calculation subgraph, and traversing the calculation graph from the bottom to the top, creating a subgraph list and updating the condition table until all calculation nodes are added to the corresponding calculation subgraph.

Benefits of technology

It realizes efficient segmentation of computational graphs on heterogeneous devices, and is suitable for scenarios where deep learning models are used for use of acceleration cards, improves computing efficiency and accuracy, and reduces the need for additional labeling of information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068965A_ABST
    Figure CN120068965A_ABST
Patent Text Reader

Abstract

The invention discloses a computational graph segmentation method for heterogeneous equipment, and relates to the technical field of data management. Comprising the steps of 1, obtaining a computational graph corresponding to a deep learning model, 2, defining a condition table of a computational sub-graph, 3, traversing the computational graph from bottom to top, 4, creating a sub-graph list, traversing each father node of a currently accessed computational node, creating a limited sub-graph list, and 5, creating a limited sub-graph list. Adding the value of the condition table of the calculation sub-graph to which each father node belongs to a limited sub-graph list, searching the father node consistent with the back end of the currently accessed calculation node, if the calculation sub-graph in which the father node consistent with the back end is located is not in the limited sub-graph list, adding the calculation sub-graph in which the father node consistent with the back end is located to the sub-graph list, and if the calculation sub-graph in which the father node consistent with the back end is located is not in the limited sub-graph list, adding the father node consistent with the back end to the limited sub-graph list. The method comprises the following steps of: (1) traversing the computing nodes in the computing graph, (2) adding the currently accessed computing nodes into the computing sub-graphs in the sub-graph list, (3) adding the currently accessed computing nodes into the computing sub-graphs in the sub-graph list, (4) updating a condition table of the computing sub-graphs in the sub-graph list, (7) continuing to traverse the remaining computing nodes in the computing graph until all the computing nodes are added into the corresponding computing sub-graphs, and (8) executing all the computing sub-graphs. The method is used for realizing an accelerated reasoning process of a deep learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention discloses a computational graph partitioning method for heterogeneous devices, which relates to the technical field of data management. Background Art

[0002] Benefiting from the explosive growth of electronic data and hardware computing power, the performance of deep learning models, especially large models, is getting better and better. Deep learning models have been widely used and quickly implemented and popularized in many fields. Due to the large amount of computation, deep learning models often need to use acceleration cards for inference, with GPU cards being the mainstream. When using acceleration cards for inference, part of the computation needs to be performed on the host side, and part of the computationally intensive part needs to be performed on the device side. Therefore, it is very important to reasonably partition the computational tasks. Computational graph partitioning can be used to partition computational tasks, but there are still problems to varying degrees at present. For example, in the face of multiple outputs, the data dependence problem cannot be effectively solved, resulting in the subgraphs divided may not be computable, affecting the computational efficiency and accuracy. In addition, some computational graph partitioning methods require adding a lot of additional annotation information to assist in partitioning the computational graph, resulting in a large amount of work and redundancy. Summary of the Invention

[0003] In view of the problems of the prior art, the present invention provides a computational graph partitioning method for heterogeneous devices, which is applicable to the scenario of using acceleration cards to infer deep learning models, and can be used in deep learning tasks such as natural language processing, computer vision, and large models, and has high practical value and innovation value.

[0004] The specific solution proposed by the present invention is as follows:

[0005] The present invention provides a computational graph partitioning method for heterogeneous devices, including:

[0006] Step 1: Obtain the computational graph corresponding to the deep learning model, represent the computational graph as a directed acyclic graph, determine the device type on which each computational node in the computational graph runs, and perform backend annotation on the computational nodes.

[0007] Step 2: Define a condition table for computational subgraphs, where the keys in the condition table represent computational subgraphs, and the values represent all computational subgraphs that cannot be merged with the computational subgraphs represented by the keys.

[0008] Step 3: Traverse the computational graph from bottom to top: starting from the output computational node of the directed acyclic graph, recursively visit each computational node in the graph.

[0009] Step 4: Create a sub-graph list. Traverse each parent node of the currently visited computing node, create a restricted sub-graph list, add the values of the condition tables of the computing sub-graphs to which each parent node belongs to the restricted sub-graph list. Search for the parent node with the same backend as the currently visited computing node. If the computing sub-graph where the parent node with the same backend is located is not in the restricted sub-graph list, add the computing sub-graph where the parent node with the same backend is located to the sub-graph list.

[0010] Step 5: Add the currently visited computing node to the computing sub-graph in the sub-graph list.

[0011] Step 6: Update the condition table of the computing sub-graph in the sub-graph list.

[0012] Step 7: Continue to traverse the remaining computing nodes in the computing graph until all computing nodes are added to the corresponding computing sub-graphs.

[0013] Step 8: According to the obtained computing sub-graphs and condition lists, traverse each computing sub-graph and execute all computing sub-graphs to implement the accelerated inference process of the deep learning model.

[0014] Further, in step 5 of the method for partitioning a computing graph for heterogeneous devices, it includes: judging the situation of the computing sub-graphs in the sub-graph list, and adding the currently visited computing node to the computing sub-graph in the sub-graph list according to the judgment result. If there are multiple computing sub-graphs in the sub-graph list, merge all computing sub-graphs and add the currently visited computing node to the merged computing sub-graph; if there is only one computing sub-graph in the sub-graph list, add the currently visited computing node to the computing sub-graph; if the sub-graph list is empty, create a new computing sub-graph and add the currently visited computing node to the new computing sub-graph.

[0015] Further, in step 6 of the method for partitioning a computing graph for heterogeneous devices, if there are multiple computing sub-graphs in the sub-graph list, merge the computing sub-graphs, update the condition list of the merged computing sub-graph, and the value of the condition list of the merged computing sub-graph is the intersection of the condition list values of all the merged computing sub-graphs.

[0016] Further, when traversing each computing sub-graph in step 8 of the method for partitioning a computing graph for heterogeneous devices, judge whether the condition list of the current computing sub-graph is empty. If so, execute the computing task on the backend corresponding to the computing sub-graph, and delete the computing sub-graph from the condition lists of all computing sub-graphs after execution.

[0017] The present invention also provides a device for partitioning a computing graph for heterogeneous devices, including a computing graph management module, a condition table management module, a traversal access module, a sub-graph management module, and a model inference module.

[0018] The computation graph management module obtains the computation graph corresponding to the deep learning model, represents the computation graph as a directed acyclic graph, determines the device type on which each computation node on the computation graph runs, and performs backend annotation on the computation nodes.

[0019] The conditional table management module defines the conditional table of the computation subgraph. In the conditional table, the computation subgraph is represented by a key, and all computation subgraphs that cannot be merged with the computation subgraph represented by the key are represented by values.

[0020] The traversal access module traverses the computation graph from bottom to top: starting from the output computation node of the directed acyclic graph, recursively accesses each computation node in the graph.

[0021] The subgraph management module creates a subgraph list, traverses each parent node of the currently accessed computation node, creates a restricted subgraph list, adds the values of the conditional tables of the computation subgraphs to which each parent node belongs to the restricted subgraph list, searches for a parent node with the same backend as the currently accessed computation node. If the computation subgraph where the parent node with the same backend is located is not in the restricted subgraph list, then adds the computation subgraph where the parent node with the same backend is located to the subgraph list.

[0022] Add the currently accessed computation node to the computation subgraph in the subgraph list.

[0023] Update the conditional table of the computation subgraph in the subgraph list.

[0024] Continue to traverse the remaining computation nodes in the computation graph until all computation nodes are added to the corresponding computation subgraphs.

[0025] The model inference module traverses each computation subgraph according to the obtained computation subgraphs and the condition list, and executes all computation subgraphs to implement the accelerated inference process of the deep learning model.

[0026] Furthermore, the subgraph management module of the computation graph splitting device for heterogeneous devices determines the situation of the computation subgraphs in the subgraph list, and adds the currently accessed computation node to the computation subgraph in the subgraph list according to the judgment result. Among them, if there are multiple computation subgraphs in the subgraph list, then merges all computation subgraphs, and adds the currently accessed computation node to the merged computation subgraph; if there is only one computation subgraph in the subgraph list, then adds the currently accessed computation node to the computation subgraph; if the subgraph list is empty, then creates a new computation subgraph and adds the currently accessed computation node to the new computation subgraph.

[0027] Furthermore, if there are multiple computation subgraphs in the subgraph list of the computation graph splitting device for heterogeneous devices, then the subgraph management module merges the computation subgraphs, updates the condition list of the merged computation subgraph, and the value of the condition list of the merged computation subgraph is the intersection of the condition list values of all the merged computation subgraphs.

[0028] Furthermore, when the model inference module of the computing graph splitting device for heterogeneous devices traverses each computing sub-graph, it determines whether the condition list of the current computing sub-graph is empty. If it is, it executes the computing task on the backend corresponding to the computing sub-graph, and deletes the computing sub-graph from the condition lists of all computing sub-graphs after execution.

[0029] The beneficial effects of the present invention are:

[0030] According to the computing properties of nodes, the running backends are labeled for the nodes on the computing graph, and the condition table is initialized. Then, traverse the computing graph from bottom to top. For each node, create a list of sub-graphs that can be added, add the current node to the sub-graph, and update the condition table. Repeat the above steps until all nodes are added to a certain sub-graph. Through the above method, the present invention realizes efficient splitting of the computing graph on heterogeneous devices, is applicable to scenarios of using acceleration cards to infer deep learning models, and can be used in deep learning tasks such as natural language processing, computer vision, and large models, with high practical value and innovation value. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is a schematic diagram of the application process of the method of the present invention.

[0032] Figure 2 It is a schematic diagram of the computing graph. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments cited are not intended to limit the present invention.

[0034] Embodiment 1

[0035] The present invention provides a method for splitting a computing graph for heterogeneous devices, including:

[0036] Step 1: Obtain the computing graph corresponding to the deep learning model, represent the computing graph as a directed acyclic graph, determine the device type on which each computing node on the computing graph runs, and perform backend labeling on the computing nodes. Perform backend labeling on the computing nodes on the computing graph, such as CPU, GPU, etc.

[0037] For example Figure 2 in, computing nodes A, B, and F in the directed acyclic graph are one backend, and computing nodes C, D, E, and G are another backend.

[0038] Step 2: Define the condition table of the computing sub-graph. In the condition table, the key represents the computing sub-graph, and the value represents all computing sub-graphs that cannot be merged with the computing sub-graph represented by the key. Among them, the condition table is a dictionary. When initializing the condition table, the value in the condition table is empty.

[0039] Step 3: Traverse the computation graph from bottom to top: Starting from the output computation node of the directed acyclic graph, recursively visit each computation node in the graph. During the visit, when all the parent nodes of a computation node have been assigned to a certain computation sub-graph, or the computation node has no parent nodes, consider adding the computation node to the computation sub-graph for subsequent processes. For example Figure 2 in, traverse the nodes upward from the computation node G until the computation node A meets the above conditions, and consider adding the A computation node to a certain computation sub-graph.

[0040] Step 4: Create a sub-graph list, traverse each parent node of the currently visited computation node, create a restricted sub-graph list, add the values of the condition tables of the computation sub-graphs to which each parent node belongs to the restricted sub-graph list, find the parent node whose backend is consistent with the currently visited computation node, and if the computation sub-graph where the parent node with the consistent backend is not in the restricted sub-graph list, add the computation sub-graph where the parent node with the consistent backend is to the sub-graph list.

[0041] For example Figure 2 in, assume that the currently visited computation node is F, and its parent nodes B and E have been assigned to appropriate computation sub-graphs. At this time, B is in the computation sub- Figure 1 graph, including nodes A and B. E is in the computation sub- Figure 2 graph, including nodes C, D, and E. The values of the condition tables of the computation sub- Figure 1 graph and the computation sub- Figure 2 graph are both empty. At this time, the condition table is as follows: {"computation sub- Figure 1 graph": [], "computation sub- Figure 2 graph": []}, so the restricted sub-graph list is empty. Since F is consistent with the backend of the computation sub- Figure 1 graph, and the computation sub- Figure 1 graph is not in the restricted sub-graph list, the computation sub- Figure 1 graph can be added to the sub-graph list. The backend of the computation sub- Figure 2 graph is inconsistent and is not added to the sub-graph list.

[0042] Step 5: Add the currently visited computation node to the computation sub-graph in the sub-graph list. Further, the situation of the computation sub-graphs in the sub-graph list can be judged, and according to the judgment result, the currently visited computation node is added to the computation sub-graph in the sub-graph list. Among them, if there are multiple computation sub-graphs in the sub-graph list, all the computation sub-graphs are merged, and the currently visited computation node is added to the merged computation sub-graph; if there is only one computation sub-graph in the sub-graph list, the currently visited computation node is added to the computation sub-graph; if the sub-graph list is empty, a new computation sub-graph is created, and the currently visited computation node is added to the new computation sub-graph.

[0043] Step 6: Update the condition table of the computing subgraphs in the subgraph list. If there are multiple computing subgraphs in the subgraph list, merge the computing subgraphs, update the condition list of the merged computing subgraph, and the value of the condition list of the merged computing subgraph is the intersection of the condition list values of all the merged computing subgraphs. Additionally, if the currently visited computing node is added to a computing subgraph, the computing subgraph to which the parent node of the currently visited computing node belongs can be added to the condition list of the subgraph where the computing node is located.

[0044] Step 7: Continue to traverse the remaining computing nodes in the computing graph until all computing nodes are added to the corresponding computing subgraphs.

[0045] Step 8: According to the obtained computing subgraphs and condition lists, traverse each computing subgraph and execute all computing subgraphs to implement the accelerated inference process of the deep learning model.

[0046] Furthermore, in step 6 of the method for splitting a computing graph for heterogeneous devices,

[0047] Furthermore, when traversing each computing subgraph in step 8 of the method for splitting a computing graph for heterogeneous devices, determine whether the condition list of the current computing subgraph is empty. If it is, execute the computing task on the backend corresponding to the computing subgraph, and after execution, delete the computing subgraph from the condition lists of all computing subgraphs.

[0048] Embodiment 2

[0049] The present invention also provides a device for splitting a computing graph for heterogeneous devices, including a computing graph management module, a condition table management module, a traversal access module, a subgraph management module, and a model inference module.

[0050] The computing graph management module obtains the computing graph corresponding to the deep learning model, represents the computing graph as a directed acyclic graph, determines the device type on which each computing node in the computing graph runs, and performs backend annotation on the computing nodes.

[0051] The condition table management module defines the condition table of the computing subgraphs, where the key represents the computing subgraph in the condition table, and the value represents all the computing subgraphs that cannot be merged with the computing subgraph represented by the key.

[0052] The traversal access module traverses the computing graph from bottom to top: starting from the output computing node of the directed acyclic graph, recursively access each computing node in the graph.

[0053] The sub-graph management module creates a sub-graph list, traverses each parent node of the currently accessed computing node, creates a restricted sub-graph list, adds the values of the condition tables of the computing sub-graphs to which each parent node belongs to the restricted sub-graph list, and searches for a parent node with the same backend as the currently accessed computing node. If the computing sub-graph where the parent node with the same backend is located is not in the restricted sub-graph list, add the computing sub-graph where the parent node with the same backend is located to the sub-graph list.

[0054] Add the currently accessed computing node to the computing sub-graph in the sub-graph list.

[0055] Update the condition table of the computing sub-graph in the sub-graph list.

[0056] Continue to traverse the remaining computing nodes in the computing graph until all computing nodes are added to the corresponding computing sub-graphs.

[0057] The model inference module traverses each computing sub-graph according to the obtained computing sub-graphs and condition lists, and executes all computing sub-graphs to implement the accelerated inference process of the deep learning model.

[0058] Regarding the information interaction and execution process among the above-mentioned modules in the device, since they are based on the same concept as the method embodiment of the present invention, the specific content can be referred to the description in the method embodiment of the present invention, and will not be elaborated here.

[0059] Similarly, the device of the present invention annotates the running backend for the nodes on the computing graph according to the node computing properties and initializes the condition table. Then, traverse the computing graph from bottom to top. For each node, create a list of sub-graphs that can be added, add the current node to the sub-graph, and update the condition table. Repeat the above steps until all nodes are added to a certain sub-graph. Through the above method, the present invention realizes efficient partitioning of the computing graph on heterogeneous devices, is applicable to the scenario of using an acceleration card to infer a deep learning model, and can be used in deep learning tasks such as natural language processing, computer vision, and large models, with high practical value and innovation value.

[0060] It should be noted that not all steps and modules in the above-mentioned processes and device structures are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted according to needs. The system structure described in the above embodiments can be a physical structure or a logical structure, that is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities respectively, or some components in multiple independent devices may be jointly implemented.

[0061] The above-described embodiments are merely preferred embodiments cited to fully illustrate the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or transformations made by those skilled in the art on the basis of the present invention are all within the protection scope of the present invention. The protection scope of the present invention shall be subject to the claims.

Claims

1. A computational graph segmentation method for heterogeneous devices, characterized in that include: Step 1: Get the computational graph corresponding to the deep learning model, represent the computational graph as a directed acyclic graph, determine the device type running on each computing node on the computational graph, and annotate the backend of the computing node. Step 2: Define the condition table of the computational subgraph. In the condition table, the computational subgraph is represented by the key, and all computational subgraphs that cannot be merged with the computational subgraph represented by the key are represented by the value. Step 3: Traverse the computational graph from bottom to top: Starting from the output computational node of the directed acyclic graph, recursively visit each computational node in the graph. Step 4: Create a subgraph list, traverse each parent node of the currently visited computing node, create a restricted subgraph list, add the values ​​of the condition table of the computing subgraph to which each parent node belongs to the restricted subgraph list, and search for the parent node that is consistent with the backend of the currently visited computing node. If the computing subgraph where the parent node with the consistent backend is located is not in the restricted subgraph list, then add the computing subgraph where the parent node with the consistent backend is located to the subgraph list. Step 5: Add the currently visited computation node to the computation subgraph list. Step 6: Update the condition table for calculating subgraphs in the subgraph list. Step 7: Continue to traverse the remaining computing nodes in the computation graph until all computing nodes are added to the corresponding computation subgraph. Step 8: According to the obtained computational subgraphs and condition lists, traverse each computational subgraph and execute all computational subgraphs to realize the accelerated reasoning process of the deep learning model.

2. A computational graph segmentation method for heterogeneous devices according to claim 1, characterized in that Step 5 includes: judging the situation of the computational subgraph in the subgraph list, and adding the currently accessed computational node to the computational subgraph in the subgraph list according to the judgment result, wherein if there are multiple computational subgraphs in the subgraph list, all computational subgraphs are merged, and the currently accessed computational node is added to the merged computational subgraph; if there is only one computational subgraph in the subgraph list, the currently accessed computational node is added to the computational subgraph; if the subgraph list is empty, a new computational subgraph is created, and the currently accessed computational node is added to the new computational subgraph.

3. A computational graph segmentation method for heterogeneous devices according to claim 1 or 2, characterized in that In step 6, if there are multiple computation subgraphs in the subgraph list, the computation subgraphs are merged, and the condition list of the merged computation subgraph is updated. The condition list value of the merged computation subgraph is the intersection of the condition list values ​​of all merged computation subgraphs.

4. According to claim 1, a computational graph segmentation method for heterogeneous devices is characterized in that when traversing each computational subgraph in step 8, it is determined whether the condition list of the current computational subgraph is empty. If so, the computational task is executed in the backend corresponding to the computational subgraph, and after execution, the computational subgraph is deleted from the condition list of all computational subgraphs.

5. A computing graph slicing device for heterogeneous devices, characterized in that It includes calculation graph management module, condition table management module, traversal access module, subgraph management module and model reasoning module. The computational graph management module obtains the computational graph corresponding to the deep learning model, represents the computational graph as a directed acyclic graph, determines the device type running on each computing node on the computational graph, and performs backend annotation of the computing nodes. The condition table management module defines the condition table of the computational subgraph. In the condition table, the computational subgraph is represented by a key, and all computational subgraphs that cannot be merged with the computational subgraph represented by the key are represented by a value. The traversal access module traverses the computational graph from bottom to top: starting from the output computational node of the directed acyclic graph, recursively accesses each computational node in the graph, The subgraph management module creates a subgraph list, traverses each parent node of the currently accessed computing node, creates a restricted subgraph list, adds the values ​​of the condition table of the computing subgraph to which each parent node belongs to the restricted subgraph list, and searches for the parent node that is consistent with the backend of the currently accessed computing node. If the computing subgraph where the parent node with the consistent backend is located is not in the restricted subgraph list, then the computing subgraph where the parent node with the consistent backend is located is added to the subgraph list. Add the currently visited computation node to the computation subgraph list, Update the condition table for calculating subgraphs in the subgraph list, Continue to traverse the remaining computational nodes in the computational graph until all computational nodes are added to the corresponding computational subgraph. The model reasoning module traverses each computational subgraph according to the obtained computational subgraphs and condition lists, and executes all computational subgraphs to realize the accelerated reasoning process of the deep learning model.

6. The computing graph slicing device for heterogeneous devices according to claim 5, characterized in that The subgraph management module determines the status of the computational subgraph in the subgraph list, and adds the currently accessed computational node to the computational subgraph in the subgraph list according to the determination result. If there are multiple computational subgraphs in the subgraph list, all computational subgraphs are merged, and the currently accessed computational node is added to the merged computational subgraph; if there is only one computational subgraph in the subgraph list, the currently accessed computational node is added to the computational subgraph; if the subgraph list is empty, a new computational subgraph is created, and the currently accessed computational node is added to the new computational subgraph.

7. A computing graph slicing device for heterogeneous devices according to claim 5 or 6, characterized in that If there are multiple computational subgraphs in the subgraph list, the subgraph management module merges the computational subgraphs and updates the condition list of the merged computational subgraph. The condition list value of the merged computational subgraph is the intersection of the condition list values ​​of all merged computational subgraphs.

8. The computing graph slicing device for heterogeneous devices according to claim 5 is characterized in that the model When the inference module traverses each computational subgraph, it determines whether the condition list of the current computational subgraph is empty. If so, it executes the computational task on the backend corresponding to the computational subgraph. After the execution, the computational subgraph is deleted from the condition list of all computational subgraphs.

Citation Information

Cited By

  • Cloud edge collaborative scheduling and agent operation method and system based on distributed idle computing power network

    CN122226727A