Computational graph processing method and device, electronic equipment and readable storage medium
In the calculation graph processing, loop detection, address repeated write detection and head node detection are performed based on the object's entry degree, which solves the problem of high time complexity in traditional methods, and realizes more efficient calculation graph verification and processing.
Patent Information
- Application Number
- CN202510264874.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-20
AI Technical Summary
The traditional computing graph processing method requires node-by-side traversal verification, resulting in high time complexity and low efficiency.
During the topological ordering of nodes, the calculation graph is traversed, and the calculation graph is performed based on the incoming of the object, and the calculation graph is verified based on the loop detection results, the address repeated write detection results and the head node detection results.
In the process of verifying the calculation graph at one time, loop detection, address repetitive write detection and head node detection can be completed simultaneously, avoiding the high time complexity of node-by-side traversal verification in traditional methods, reducing the time complexity, and improving the efficiency of calculation graph processing.
Smart Images

Figure CN120179866A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and in particular, to a computational graph processing method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] A computational graph includes multiple computational nodes, and each computational node is connected by input and output parameters. After the computational graph is defined, the computational graph is processed such as verified and executed by the hardware or software of an electronic device. However, in traditional computational graph processing, the computational graph is usually traversed and verified node by node and edge by edge, resulting in a high time complexity. Summary of the Invention
[0003] Embodiments of this application provide a computational graph processing method, apparatus, electronic device, and computer-readable storage medium, which can reduce the time complexity.
[0004] In a first aspect, this application provides a computational graph processing method. The computational graph includes nodes and interface parameters of the nodes, and each node and each interface parameter are respectively used as an object. The method includes:
[0005] During the topological sorting of the nodes, traverse the computational graph, and perform a loop detection on the computational graph based on the in-degree of the object to obtain a loop detection result. The in-degree represents the number of forward dependencies of the object.
[0006] During the traversal of the computational graph, perform an address duplicate write detection on the nodes based on the in-degree of the parameter to obtain an address duplicate write detection result.
[0007] Perform a head node detection on the set of nodes after topological sorting to obtain a head node detection result.
[0008] Verify the computational graph based on the loop detection result, the address duplicate write detection result, and the head node detection result.
[0009] In one of the embodiments, the traversing the computational graph, performing a loop detection on the computational graph based on the in-degree of the object to obtain a loop detection result includes:
[0010] Starting from a first object in the computational graph, traverse the backward-dependent objects of the first object according to an object linked list, decrement the in-degree of the backward-dependent object by 1. If the in-degree of the backward-dependent object is obtained as 0, use the backward-dependent object as the new first object, and return to execute the step of traversing the backward-dependent objects of the first object according to the object linked list. The in-degree of the first object is 0, and the object linked list is generated based on the forward-dependent objects and backward-dependent objects of the attribute parameters of each node.
[0011] Based on the number of first objects traversed, obtain the loop detection result.
[0012] In one embodiment, the obtaining the loop detection result based on the number of first objects traversed includes:
[0013] If the number of first objects traversed is the same as the number of objects in the computation graph, the loop detection result is that there is no circular dependency in the computation graph;
[0014] If the number of first objects traversed is different from the number of objects in the computation graph, the loop detection result is that there is a circular dependency in the computation graph.
[0015] In one embodiment, during the process of traversing the computation graph, performing address duplicate writing detection on the nodes based on the in-degree of the parameter to obtain the address duplicate writing detection result includes:
[0016] During the process of traversing the computation graph, if there is an interface parameter with an output attribute and the in-degree of the interface parameter is greater than the number of forward dependencies of the interface parameter, the obtained address duplicate writing detection result is that there are at least two nodes writing to the same address;
[0017] If there is an interface parameter with an output attribute and the in-degree of the interface parameter is equal to the number of forward dependencies of the interface parameter, the obtained address duplicate writing detection result is that there are not at least two nodes writing to the same address.
[0018] In one embodiment, the performing head node detection on the node set after topological sorting to obtain the head node detection result includes:
[0019] During the process of traversing the computation graph, if the object is a node, put the interface parameter of the node into the node set of the computation graph to complete the layer-by-layer sorting of each node in the computation graph;
[0020] Traverse the interface parameters of the nodes in the node set, and determine the node to which the interface parameter with only input attributes belongs as the head node to obtain the head node detection result.
[0021] In one embodiment, after verifying the computation graph based on the loop detection result, the address duplicate writing detection result, and the head node detection result, it further includes:
[0022] Determine the first node according to the topological relationship of each node, and execute the first node; the first node is the head node when it is executed for the first time;
[0023] Determine the backward dependent nodes of the first node, decrement the in-degree of the backward dependent nodes by 1, and use the backward dependent nodes as the new first node. Return to execute the step of executing the first node until all nodes in the computational graph have been executed.
[0024] In one embodiment, the method further includes:
[0025] Traverse each node in the computationally sorted graph to obtain the interface parameters of each node;
[0026] If the interface parameters of the output attribute of the second node are the same as the interface parameters of the input attribute of the third node, increment the in-degree of the third node by 1;
[0027] If the in-degree of the third node is greater than 1, establish a dependency relationship between the third node and the second node until all nodes in the topologically sorted computational graph have been traversed to obtain the topological relationships of all nodes.
[0028] In a second aspect, the present application also provides a computational graph processing device. The computational graph includes nodes and the interface parameters of the nodes, and each node and each interface parameter are respectively used as an object. The device includes:
[0029] A loop detection module for traversing the computational graph during the topological sorting of the nodes, performing loop detection on the computational graph based on the in-degree of the object, and obtaining a loop detection result; the in-degree represents the number of forward dependencies of the object;
[0030] An address duplicate write detection module for performing address duplicate write detection on the nodes based on the in-degree of the parameter during the traversal of the computational graph, and obtaining an address duplicate write detection result;
[0031] A head node detection module for performing head node detection on the set of topologically sorted nodes to obtain a head node detection result;
[0032] A computational graph verification module for verifying the computational graph based on the loop detection result, the address duplicate write detection result, and the head node detection result.
[0033] In a third aspect, the present application also provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0034] During the topological sorting of the nodes, traverse the computational graph, perform loop detection on the computational graph based on the in-degree of the object, and obtain a loop detection result; the in-degree represents the number of forward dependencies of the object;
[0035] During the process of traversing the computational graph, address duplicate writing detection is performed on the nodes based on the in-degree of the parameters, and an address duplicate writing detection result is obtained;
[0036] Head node detection is performed on the set of nodes after topological sorting, and a head node detection result is obtained;
[0037] Based on the loop detection result, the address duplicate writing detection result, and the head node detection result, the computational graph is verified.
[0038] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0039] During the process of performing topological sorting on the nodes, the computational graph is traversed, and loop detection is performed on the computational graph based on the in-degree of the object, and a loop detection result is obtained; the in-degree represents the number of forward dependencies of the object;
[0040] During the process of traversing the computational graph, address duplicate writing detection is performed on the nodes based on the in-degree of the parameters, and an address duplicate writing detection result is obtained;
[0041] Head node detection is performed on the set of nodes after topological sorting, and a head node detection result is obtained;
[0042] Based on the loop detection result, the address duplicate writing detection result, and the head node detection result, the computational graph is verified.
[0043] In a fifth aspect, the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the following steps are implemented:
[0044] During the process of performing topological sorting on the nodes, the computational graph is traversed, and loop detection is performed on the computational graph based on the in-degree of the object, and a loop detection result is obtained; the in-degree represents the number of forward dependencies of the object;
[0045] During the process of traversing the computational graph, address duplicate writing detection is performed on the nodes based on the in-degree of the parameters, and an address duplicate writing detection result is obtained;
[0046] Head node detection is performed on the set of nodes after topological sorting, and a head node detection result is obtained;
[0047] Based on the loop detection result, the address duplicate writing detection result, and the head node detection result, the computational graph is verified.
[0048] The above calculation graph processing method, device, electronic device, computer-readable storage medium, and computer program product traverse the calculation graph during the topological sorting of nodes, perform loop detection on the calculation graph based on the in-degree of objects, and obtain a loop detection result, where the in-degree represents the number of forward dependencies of the object; during the traversal of the calculation graph, perform address duplicate write detection on the nodes based on the in-degree of parameters, and obtain an address duplicate write detection result; perform head node detection on the set of nodes after topological sorting, and obtain a head node detection result; then, based on the loop detection result, address duplicate write detection result, and head node detection result, loop detection, address duplicate write detection, and head node detection can be completed simultaneously during one verification of the calculation graph, avoiding the problem of high time complexity caused by traversing and verifying each node and each edge of the calculation graph, thereby reducing the time complexity and improving the efficiency of calculation graph processing. Description of the Drawings
[0049] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0050] Figure 1 It is a schematic flowchart of the calculation graph processing method in an embodiment;
[0051] Figure 2 It is a schematic flowchart of verifying the calculation graph in an embodiment;
[0052] Figure 3 It is a schematic flowchart of the processing flow of interface parameters as input attributes in an embodiment;
[0053] Figure 4 It is a schematic diagram of the data structure of interface parameters as input attributes in an embodiment;
[0054] Figure 5 It is a schematic flowchart of the processing flow of interface parameters as output attributes in an embodiment;
[0055] Figure 6 It is a schematic diagram of the data structure of interface parameters as output attributes in an embodiment;
[0056] Figure 7 It is a schematic diagram of connecting all objects with an in-degree of 0 in an embodiment;
[0057] Figure 8 It is a schematic flowchart of executing the calculation graph in an embodiment;
[0058] Figure 9 Schematic diagram of constructing the topological relationship between nodes in an embodiment;
[0059] Figure 10 Schematic diagram of each node before topological sorting in an embodiment;
[0060] Figure 11 Schematic diagram of each node after topological sorting in an embodiment;
[0061] Figure 12 Block diagram of the computational graph processing device in an embodiment;
[0062] Figure 13 Internal structure diagram of an electronic device in an embodiment. Detailed implementation manners
[0063] In order to make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0064] In one embodiment, as Figure 1 shown, a computational graph processing method is provided. In this embodiment, the method is exemplified by being applied to an electronic device. The electronic device can be a terminal or a server. It can be understood that the method can also be applied to a system including a terminal and a server and implemented through the interaction between the terminal and the server. Among them, the terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted device can be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0065] In this embodiment, the computational graph includes nodes and interface parameters of the nodes, and each node and each interface parameter are respectively used as an object. The computational graph processing method includes the following steps:
[0066] Step S102, during the topological sorting of the nodes, traverse the computational graph, perform a loop detection on the computational graph based on the in-degree of the object, and obtain a loop detection result; the in-degree represents the number of forward dependencies of the object.
[0067] Among them, a computational graph (Graph) is a graphical structure used to represent a computational process. The computational task is decomposed into multiple interconnected nodes (Nodes), each node represents an operation or function, and the edges (Edges) represent the data dependency relationships between the nodes. Computational graphs are widely used in fields such as deep learning, machine learning, and scientific computing, especially in automatic differentiation and optimization algorithms. The computational graph can be a computational graph in a heterogeneous framework.
[0068] An interface parameter (Parameter) is at least one of the input parameters and output parameters of a node. The input parameter is passed to the node and processed in the node, and the output parameter is given to the next node or used as the final output.
[0069] Optionally, after obtaining the defined computational graph, the electronic device processes the computational graph through the following steps:
[0070] 1. Create a context (Context): The context manages all objects in the OpenVX runtime;
[0071] 2. Create a computational graph (Graph): The computational graph is a conforming operation and consists of one or more graph nodes (Nodes);
[0072] 3. Add computational nodes (Nodes) to the graph: The nodes are basic operations and will execute specific operators;
[0073] 4. Verify the computational graph (Verify graph): Verify the consistency of the graph structure and internal data, and check for circular graphs, etc.;
[0074] 5. Execute the computational graph (Process graph): Execute the verified computational graph, schedule the Nodes in the Graph, and execute the operators (Kernels) therein to complete the computational task;
[0075] 6. Release resources: After the graph is executed, release the resources related to the graph and the context.
[0076] During the process of the electronic device verifying the computational graph (Verify graph), during the topological sorting of the nodes, traverse the computational graph, perform loop detection on the computational graph based on the in-degree of the objects, and obtain the loop detection result; during the process of traversing the computational graph, perform address duplicate write detection on the nodes based on the in-degree of the parameters, and obtain the address duplicate write detection result; perform head node detection on the set of nodes after topological sorting, and obtain the head node detection result; based on the loop detection result, the address duplicate write detection result, and the head node detection result, verify the computational graph.
[0077] Optionally, the electronic device abstracts each node and each interface parameter in the computation graph into an object (Object). Each object has an in-degree and a pointer (top) that points to the next-layer object, and an object array with a size equal to the number of objects is allocated. The object array stores object information. It can be understood that each abstraction has a pointer that connects the next-layer objects of the object with a linked list.
[0078] Step S104: During the process of traversing the computation graph, perform an address duplicate write detection on the nodes based on the in-degree of the parameters to obtain an address duplicate write detection result.
[0079] The address duplicate write detection result refers to the result of whether there are at least two nodes writing to the same address.
[0080] Optionally, during the process of traversing the computation graph, perform an address duplicate write detection on the nodes based on the in-degree of the parameters to obtain an address duplicate write detection result, including: during the process of traversing the computation graph, if there is an interface parameter with an output attribute and the in-degree of the interface parameter is greater than the number of forward dependencies of the interface parameter, the obtained address duplicate write detection result is that there are at least two nodes writing to the same address; if there is an interface parameter with an output attribute and the in-degree of the interface parameter is equal to the number of forward dependencies of the interface parameter, the obtained address duplicate write detection result is that there are no at least two nodes writing to the same address.
[0081] During the process of traversing the computation graph, obtain the interface parameters whose in-degree changes after traversal. If the in-degree of the interface parameter with an output attribute whose in-degree changes after traversal is greater than 0, it means that the in-degree of the interface parameter is greater than the number of forward dependencies of the interface parameter, and the obtained address duplicate write detection result is that there are at least two nodes writing to the same address. If the in-degree of the interface parameter with an output attribute whose in-degree changes after traversal is 0, it means that the in-degree of the interface parameter is equal to the number of forward dependencies of the interface parameter, and the obtained address duplicate write detection result is that there are no at least two nodes writing to the same address.
[0082] For example, if Node 1 is connected to Node 2 and Node 3, and their interface parameters are the same, i.e., (Output1, Input2, Input3), the in-degree (number of forward dependencies) of this interface parameter is 1 because the forward dependency of this interface parameter is Node 1. During the process of traversing the computational graph, when Node 1 is traversed, the in-degree of the backward dependency object, i.e., this interface parameter (Output1, Input2, Input3), will be decremented by 1. If the in-degree of the interface parameter of the output attribute whose in-degree has changed after traversal is 0, the detected result of address duplicate writing is that there are no at least two nodes writing to the same address; if the in-degree of this interface parameter (Output1, Input2, Input3) is 2 before traversal, during the process of traversing the computational graph, when Node 1 is traversed, the in-degree of the backward dependency object, i.e., this interface parameter (Output1, Input2, Input3), will be decremented by 1. If the in-degree of the interface parameter of the output attribute whose in-degree has changed after traversal is 1 and greater than 0, it means that the in-degree of this interface parameter is greater than the number of forward dependencies of the interface parameter, that is, there are outputs of other nodes besides Node 1 for this interface parameter, which also means writing to the same address.
[0083] Step S106: Perform head node detection on the set of nodes after topological sorting to obtain the head node detection result.
[0084] Among them, the head node refers to the starting point or input point of the computational graph. It is the first node in the computational graph to receive input data and is usually the starting point of the entire computational process. Nodes after the head node will perform a series of calculations and processes based on the input data.
[0085] Optionally, performing head node detection on the set of nodes after topological sorting to obtain the head node detection result includes: during the process of traversing the computational graph, if the object is a node, put the interface parameters of the node into the set of nodes of the computational graph to complete the layer-by-layer sorting of each node in the computational graph; traverse the interface parameters of the nodes in the set of nodes, and determine the node to which the interface parameter with only input attributes belongs as the head node to obtain the head node detection result.
[0086] During the process of traversing the computational graph, if the object is a node, put the parameter values of the interface parameters of the node into the corresponding nodes in the set of nodes (NodeList) of the computational graph to complete the layer-by-layer sorting of each node in the computational graph.
[0087] Traverse the interface parameters of the nodes in the set of nodes, and determine the node to which the interface parameter with only input attributes belongs as the head node; when traversing to the interface parameter that includes both input attributes and output attributes, stop traversing to obtain the head node detection result, and this head node detection result includes the detected head node information.
[0088] It can be understood that after the nodes are topologically sorted, the head nodes will be placed on the first layer of the computational graph. If the interface parameters include input attributes and output attributes, it means that the second layer of the computational graph is reached at this time, that is, all the head nodes in the computational graph have been found at this time, and the traversal can be stopped.
[0089] Step S108, verify the computational graph based on the loop detection result, the address duplicate write detection result, and the head node detection result.
[0090] Optionally, as Figure 2 shown, the electronic device completes the topological sorting of the computational graph based on the loop detection result, the address duplicate write detection result, and the head node detection result; respectively performs Kernel preprocessing, parameter validity verification, memory space allocation, establishing the topological relationship between nodes in the computational graph, detecting whether there are unvisited nodes, Target verification, and executing Kernel initialization on the topologically sorted computational graph to complete the verification of the computational graph (Verify graph).
[0091] Optionally, after the electronic device verifies the computational graph, it executes the verified computational graph (Processgraph). Executing the verified computational graph includes: determining the first node according to the topological relationship of each node and executing the first node; the first node is the head node when it is executed for the first time; determining the backward dependent nodes of the first node, decrementing the in-degree of the backward dependent nodes by 1, and taking the backward dependent nodes as the new first node, and returning to execute the step of executing the first node until all nodes in the computational graph are executed.
[0092] In the above computational graph processing method, during the process of topologically sorting the nodes, the computational graph is traversed, and loop detection is performed on the computational graph based on the in-degree of the object to obtain the loop detection result, where the in-degree represents the number of forward dependencies of the object; during the process of traversing the computational graph, address duplicate write detection is performed on the nodes based on the in-degree of the parameter to obtain the address duplicate write detection result; head node detection is performed on the set of topologically sorted nodes to obtain the head node detection result; then, based on the loop detection result, the address duplicate write detection result, and the head node detection result, loop detection, address duplicate write detection, and head node detection can be completed simultaneously during one verification process of the computational graph, avoiding the problem of high time complexity caused by traversing and verifying each node and each edge of the computational graph, thereby reducing the time complexity and improving the efficiency of computational graph processing.
[0093] The above computational graph processing method is not only applicable to OPenVX, but also applicable to frameworks that use the computational graph model for scheduling and execution.
[0094] In one embodiment, traverse the computational graph, perform loop detection on the computational graph based on the in-degree of an object, and obtain the loop detection result, including: starting from the first object in the computational graph, traverse the backward-dependent objects of the first object according to the object linked list, decrement the in-degree of the backward-dependent object by 1. If the in-degree of the backward-dependent object is obtained as 0, use the backward-dependent object as the new first object, and return to execute the step of traversing the backward-dependent objects of the first object according to the object linked list; the in-degree of the first object is 0, and the object linked list is generated based on the forward-dependent objects and backward-dependent objects of the attribute parameters of each node; obtain the loop detection result based on the number of traversed first objects.
[0095] Among them, the forward-dependent object is the input object on which the computational graph depends during the forward propagation process. The backward-dependent object is the subsequent object during the backward propagation process of the computational graph. The object linked list is a chained table formed by the connection relationships between various objects in the computational graph.
[0096] Optionally, the electronic device traverses the computational graph, obtains the forward-dependent object and backward-dependent object of an object; generates an object linked list of the computational graph based on the forward-dependent object and backward-dependent object of each object.
[0097] When the object is an interface parameter and the interface parameter is an input attribute, increment the in-degree of the next node by 1, and connect the index of the next-level object of the interface parameter in the object array; when the object is an interface parameter and the interface parameter is an output attribute, increment the in-degree of the interface parameter by 1, and connect the successor object of the forward-dependent node of the interface parameter; obtain the object with an in-degree of 0 in the computational graph, and connect the objects with an in-degree of 0 using a linked list to obtain the object linked list.
[0098] Optionally, the electronic device traverses the interface parameters of each node, obtains the in-degree and backward-dependent object of each object, uses the parameter value (Reference) of the interface parameter as the key value, and the index of the interface parameter in the object array and the attribute of the interface parameter as the value, and stores the key value and value in the hash table.
[0099] When the interface parameter is an input attribute, increment the in-degree of the next node by 1, and connect the indexes of the next-level objects of the interface parameter in the object array using a linked list. For ease of explanation, define a structure type subsequent object list, where the successor object in the subsequent object list represents the index of the next-level object of the object in the object array, and the next member points to the next successor object.
[0100] Such as Figure 3 and Figure 4As shown, take the case of traversing to output parameter 1 / input parameter 3. The parameter values of output parameter 1, input parameter 2, and input parameter 3 are the same, so they are the same object. The in-degree of the backward dependency (node 3) of this interface parameter is incremented by 1, and the successor object of this interface parameter points to the indices of nodes 2 and 3 in the object array. The next member in the successor object list points to the successor object of output parameter 1 / input parameter 2 (i.e., connecting all the backward dependencies of this interface parameter), and the input attribute of the interface parameter is recorded.
[0101] When this interface parameter is an output attribute, increment the in-degree of this interface parameter by 1, and connect the successor objects of the forward dependency nodes of this interface parameter using a linked list.
[0102] Such as Figure 5 And Figure 6 As shown, take the case of traversing to output parameter 3 / input parameter 5. The parameter values of output parameter, input parameter 6, and input parameter 5 are the same, so they are the same object. The in-degree of this interface parameter is incremented by 1, and the successor object of the forward dependency (node 3) of this interface parameter points to the index of this interface parameter in the object array, and the output attribute of this interface parameter is recorded.
[0103] The electronic device searches for objects with an in-degree of 0 in the computational graph and connects the objects with an in-degree of 0 using a linked list to obtain an object linked list, as Figure 7 shown. In Figure 7 , since input parameter 1 has no forward dependency object, its in-degree is 0 and it will be placed in the linked list with an in-degree of 0 first; when input parameter 1 is traversed, the in-degree of node 1 will be reduced to 0 (because node 1 depends on input parameter 1 forwardly, and input parameter 1 has been traversed, which means node 1 has no forward dependency at this time), so node 1 will be placed in the linked list with an in-degree of 0. Obtain the next-level objects of node 1 (i.e., the (output parameter 1, input parameter 2, input parameter 3) interface parameter is obtained by using the above pointer in the code implementation), and reduce the in-degree of this interface parameter from 1 to 0 and place it in the linked list with an in-degree of 0; traverse the lower-level object interfaces 2 and 3 of this interface parameter, and similarly their in-degrees are also 1. If node 2 is traversed first, since the in-degree of node 4 is 2, the in-degree of node 4 is reduced to 1 at this time. When node 3 is traversed, the in-degree of node 4 will be reduced from 1 to 0, and then node 4 will be placed in the linked list with an in-degree of 0. Traverse in turn until all nodes in the computational graph are traversed.
[0104] Optionally, based on the number of the first objects traversed, a loop detection result is obtained, including: if the number of the first objects traversed is the same as the number of objects in the computational graph, the loop detection result is that there is no circular dependency in the computational graph; if the number of the first objects traversed is different from the number of objects in the computational graph, the loop detection result is that there is a circular dependency in the computational graph.
[0105] It can be understood that if the number of the first objects traversed is different from the number of objects in the computational graph, it means that the in-degree of an object in the computational graph cannot be reduced to 0, and the loop detection result is that there is a circular dependency in the computational graph.
[0106] In this embodiment, starting from the first object in the computational graph, the backward dependent objects of the first object are traversed according to the object linked list, and the in-degree of the backward dependent objects is decremented by 1. If the in-degree of the backward dependent object is 0, the backward dependent object is used as the new first object, and the step of traversing the backward dependent objects of the first object according to the object linked list is returned for execution; the in-degree of the first object is 0, and the object linked list is generated based on the forward dependent objects and backward dependent objects of the attribute parameters of each node. Based on the number of the first objects traversed, a more accurate loop detection result can be obtained. Further, if the number of the first objects traversed is the same as the number of objects in the computational graph, the loop detection result is that there is no circular dependency in the computational graph; if the number of the first objects traversed is different from the number of objects in the computational graph, the loop detection result is that there is a circular dependency in the computational graph, and it can accurately detect whether there is a circular dependency in the computational graph. Compared with the traditional method of traversing and verifying each node and each edge of the computational graph, the time complexity is reduced from O(m^2 * n^2) to O(m + n), and the efficiency is increased by 90%.
[0107] In one embodiment, after verifying the computational graph based on the loop detection result, the address duplicate write detection result, and the head node detection result, it further includes: determining a first node according to the topological relationship of each node, and executing the first node; the first node is the head node when it is executed for the first time; determining the backward dependent nodes of the first node, decrementing the in-degree of the backward dependent nodes by 1, and using the backward dependent nodes as the new first node, and returning to execute the step of executing the first node until all nodes in the computational graph are executed.
[0108] Optionally, the electronic device determines the backward dependent nodes of the first node through node scheduling (Schedule nodes), decrements the in-degree of the backward dependent nodes by 1, and uses the backward dependent nodes as the new first node, and returns to execute the step of executing the first node.
[0109] Optionally, after the electronic device executes all nodes in the computational graph, it resets the in-degree of each node with a time complexity of O(m + n).
[0110] Such as Figure 8As shown, the electronic device obtains the first node in the computational graph, executes the first node; determines the backward dependent nodes of the first node, and determines whether all the nodes in the computational graph have been executed. If not, the backward dependent nodes are used as the new first node, and the step of executing the first node is returned until all the nodes in the computational graph are executed.
[0111] Optionally, the above method further includes: traversing each node in the computationally sorted computational graph to obtain the interface parameters of each node; if the interface parameters of the output attribute of the second node are the same as the interface parameters of the input attribute of the third node, then increment the in-degree of the third node by 1; if the in-degree of the third node is greater than 1, then establish a dependency relationship between the third node and the second node until all the nodes in the computationally sorted computational graph are traversed to obtain the topological relationship of each node.
[0112] As Figure 9 shown, starting from the head node in the computationally sorted computational graph, traverse the output parameters of each second node; traverse the input parameters of the remaining third nodes; if the parameter values of the output attribute of the second node are the same as the parameter values of the input attribute of the third node, then the third node is the backward dependent node of the second node, and record the third node and increment the in-degree value by 1; determine whether the in-degree of the third node is greater than 1. If it is greater than 1, it indicates a non-single-link computational graph, and the subsequent execution of the computational graph is performed serially, and a dependency relationship between the second node and the third node is established. If the in-degree of the third node is less than or equal to 1, then mark autoSerialize as false; determine whether all the nodes have been traversed. If not, continue to traverse the output parameters of the second node. If so, end. By performing single-link detection on the computational graph when constructing the topological relationship of the computational nodes, if the computational graph is detected to be a single-link computational graph, the execution of the computational graph is changed to serial execution, reducing the hardware scheduling overhead.
[0113] Referring to Figure 10 and Figure 11 , Figure 10 is a schematic diagram of each node before topological sorting, Figure 11 is a schematic diagram of each node after topological sorting. For each node after topological sorting, the nodes less than or equal to the current layer do not need to be traversed, and only the subsequent nodes need to be traversed and the above operations are repeated, which can reduce the time complexity and improve the computational graph processing efficiency.
[0114] In this embodiment, since the topological relationship between each node has been established during the computational graph verification process, when scheduling nodes, the next node to be executed can be calculated in O(1) time complexity, and the in-degree of all nodes is reset with a time complexity of O(m + n). During the computational graph verification process, the topological relationship of the computational graph is constructed and the in-degree and out-degree of computational nodes are recorded with a time complexity less than O(m^2 * n^2), and the topological relationship of the computational graph only needs to be executed once; it is realized that during the process of running the computational graph, the scheduling of computational nodes is completed with O(1) time complexity. Compared with the traditional method of traversing each node and each edge of the computational graph to find the scheduling, the time complexity is reduced from O(m^2 * n^2) to O(1), and the efficiency is increased by 93%.
[0115] In one embodiment, a computational graph processing method is further provided, which is applied to an electronic device. The computational graph includes nodes and interface parameters of the nodes, and each node and each interface parameter are respectively used as an object; the computational graph processing method includes the following steps:
[0116] Step A1, during the topological sorting of nodes, starting from the first object in the computational graph, traverse the backward dependent objects of the first object according to the object linked list, subtract 1 from the in-degree of the backward dependent objects. If the in-degree of the backward dependent object is obtained as 0, use the backward dependent object as the new first object, and return to execute the step of traversing the backward dependent objects of the first object according to the object linked list; the in-degree of the first object is 0, and the object linked list is generated based on the forward dependent objects and backward dependent objects of the attribute parameters of each node; the in-degree represents the number of forward dependencies of the object.
[0117] Step A2, if the number of traversed first objects is the same as the number of objects in the computational graph, the result of loop detection is that the computational graph does not have circular dependencies; if the number of traversed first objects is different from the number of objects in the computational graph, the result of loop detection is that the computational graph has circular dependencies.
[0118] Step A3, during the process of traversing the computational graph, if there is an interface parameter with an output attribute and the in-degree of the interface parameter is greater than the number of forward dependencies of the interface parameter, the obtained address duplicate writing detection result is that there are at least two nodes writing to the same address; if there is an interface parameter with an output attribute and the in-degree of the interface parameter is equal to the number of forward dependencies of the interface parameter, the obtained address duplicate writing detection result is that there are no at least two nodes writing to the same address.
[0119] Step A4, in the process of traversing the calculation graph, if the object is a node, the interface parameters of the node are placed in the node set of the calculation graph to complete the layer-by-layer sorting of each node in the calculation graph; the interface parameters of the nodes in the node set are traversed, and the node to which the interface parameters with only input attributes belong is determined as the head node to obtain the head node detection result.
[0120] Step A5: Verify the computation graph based on the loop detection result, the address duplicate writing detection result and the head node detection result.
[0121] Step A6, traverse each node in the topologically sorted computation graph to obtain the interface parameters of each node; if the interface parameters of the output attributes of the second node are the same as the interface parameters of the input attributes of the third node, add 1 to the in-degree of the third node; if the in-degree of the third node is greater than 1, establish a dependency relationship between the third node and the second node, until all nodes in the topologically sorted computation graph are traversed to obtain the topological relationship of each node.
[0122] Step A7, determine the first node according to the topological relationship of each node, and execute the first node; the first node is the head node when it is executed for the first time; determine the backward dependent node of the first node, reduce the in-degree of the backward dependent node by 1, and use the backward dependent node as the new first node, return to execute the step of executing the first node, until all nodes in the calculation graph are executed.
[0123] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0124] Based on the same inventive concept, the embodiment of the present application also provides a computational graph processing device for implementing the computational graph processing method involved above. The implementation solution provided by the device to solve the problem is similar to the implementation solution recorded in the above method, so the specific limitations in one or more computational graph processing device embodiments provided below can refer to the limitations on the computational graph processing method above, and will not be repeated here.
[0125] In an exemplary embodiment, Figure 12As shown, a computational graph processing device is provided. The computational graph includes nodes and interface parameters of the nodes, and each node and each interface parameter are respectively regarded as an object. The device includes: a loop detection module 1202, an address duplicate write detection module 1204, a head node detection module 1206, and a computational graph verification module 1208, where:
[0126] The loop detection module 1202 is configured to traverse the computational graph during the topological sorting of the nodes, perform loop detection on the computational graph based on the in-degree of the object, and obtain a loop detection result; the in-degree represents the number of forward dependencies of the object.
[0127] The address duplicate write detection module 1204 is configured to perform address duplicate write detection on the nodes based on the in-degree of the parameter during the traversal of the computational graph, and obtain an address duplicate write detection result.
[0128] The head node detection module 1206 is configured to perform head node detection on the set of nodes after topological sorting, and obtain a head node detection result.
[0129] The computational graph verification module 1208 is configured to verify the computational graph based on the loop detection result, the address duplicate write detection result, and the head node detection result.
[0130] In the above-mentioned computational graph processing device, during the topological sorting of the nodes, the computational graph is traversed, loop detection is performed on the computational graph based on the in-degree of the object, and a loop detection result is obtained. The in-degree represents the number of forward dependencies of the object; during the traversal of the computational graph, address duplicate write detection is performed on the nodes based on the in-degree of the parameter, and an address duplicate write detection result is obtained; head node detection is performed on the set of nodes after topological sorting, and a head node detection result is obtained; then, based on the loop detection result, the address duplicate write detection result, and the head node detection result, loop detection, address duplicate write detection, and head node detection can be completed simultaneously during one verification process of the computational graph, avoiding the problem of high time complexity caused by traversing and verifying each node and each edge of the computational graph one by one, thereby reducing the time complexity and improving the efficiency of computational graph processing.
[0131] In one embodiment, the above-mentioned loop detection module 1202 is further configured to start from the first object in the computational graph, traverse the backward dependent objects of the first object according to the object linked list, decrement the in-degree of the backward dependent objects by 1, and if the in-degree of the obtained backward dependent object is 0, regard the backward dependent object as the new first object, and return to execute the step of traversing the backward dependent objects of the first object according to the object linked list; the in-degree of the first object is 0, and the object linked list is generated based on the forward dependent objects and backward dependent objects of the attribute parameters of each node; based on the number of the first objects traversed, a loop detection result is obtained.
[0132] In one embodiment, the above-mentioned loop detection module 1202 is further configured to: if the number of first objects traversed is the same as the number of objects in the computation graph, the loop detection result is that there is no circular dependency in the computation graph; if the number of first objects traversed is different from the number of objects in the computation graph, the loop detection result is that there is a circular dependency in the computation graph.
[0133] In one embodiment, the above-mentioned address duplicate write detection module 1204 is further configured to, during the process of traversing the computation graph, if there is an interface parameter with an output attribute and the in-degree of the interface parameter is greater than the number of forward dependencies of the interface parameter, the obtained address duplicate write detection result is that there are at least two nodes writing to the same address; if there is an interface parameter with an output attribute and the in-degree of the interface parameter is equal to the number of forward dependencies of the interface parameter, the obtained address duplicate write detection result is that there are not at least two nodes writing to the same address.
[0134] In one embodiment, the above-mentioned head node detection module 1206 is further configured to, during the process of traversing the computation graph, if the object is a node, put the interface parameter of the node into the node set of the computation graph to complete the layer-by-layer sorting of each node in the computation graph; traverse the interface parameters of the nodes in the node set, and determine the node to which the interface parameter with only input attributes belongs as the head node to obtain the head node detection result.
[0135] In one embodiment, the above-mentioned device further includes a computation graph execution module; the computation graph execution module is configured to determine a first node according to the topological relationship of each node and execute the first node; the first node is the head node when it is executed for the first time; determine the backward dependency nodes of the first node, subtract 1 from the in-degree of the backward dependency nodes, and use the backward dependency nodes as the new first node, and return to execute the step of executing the first node until all nodes in the computation graph are executed.
[0136] In one embodiment, the above-mentioned device further includes a node topological relationship acquisition module; the node topological relationship acquisition module is configured to traverse each node in the topologically sorted computation graph to obtain the interface parameters of each node; if the interface parameter of the output attribute of the second node is the same as the interface parameter of the input attribute of the third node, add 1 to the in-degree of the third node; if the in-degree of the third node is greater than 1, establish a dependency relationship between the third node and the second node until all nodes in the topologically sorted computation graph are traversed to obtain the topological relationship of each node.
[0137] Each module in the above-mentioned computation graph processing device can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor in the electronic device in hardware form or be independent of the processor, or can be stored in the memory in the electronic device in software form so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0138] In an exemplary embodiment, an electronic device is provided. The electronic device may be a terminal, and its internal structural diagram may be as shown in Figure 13 . The electronic device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the electronic device is used to exchange information between the processor and external devices. The communication interface of the electronic device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be achieved through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a method for processing computational graphs. The display unit of the electronic device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the electronic device, or an external keyboard, touchpad, or mouse, etc.
[0139] Those skilled in the art can understand that Figure 13 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the electronic device to which the solution of this application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0140] In an embodiment, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps in the above method embodiments are implemented.
[0141] In an embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0142] In an embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0143] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0144] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0145] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.
[0146] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application shall be subject to the appended claims.
Claims
1. A computational graph processing method, characterized in that: The computation graph includes nodes and interface parameters of the nodes, each node and each interface parameter is respectively regarded as an object; the method includes: In the process of topologically sorting the nodes, the computation graph is traversed, and loop detection is performed on the computation graph based on the in-degree of the object to obtain a loop detection result; the in-degree represents the number of forward dependencies of the object; In the process of traversing the computation graph, performing address duplicate writing detection on the node based on the in-degree of the parameter to obtain an address duplicate writing detection result; Perform head node detection on the node set after topological sorting to obtain the head node detection result; The calculation graph is verified based on the loop detection result, the address duplicate writing detection result and the head node detection result.
2. The method according to claim 1, characterized in that The traversing the computation graph and performing loop detection on the computation graph based on the in-degree of the object to obtain a loop detection result includes: Starting from the first object in the computation graph, traversing backward dependent objects of the first object according to the object linked list, reducing the in-degree of the backward dependent object by 1, and if the in-degree of the backward dependent object is 0, taking the backward dependent object as the new first object, and returning to execute the step of traversing backward dependent objects of the first object according to the object linked list; the in-degree of the first object is 0, and the object linked list is generated based on the forward dependent objects and backward dependent objects of the attribute parameters of each node; A loop detection result is obtained based on the number of the traversed first objects.
3. The method according to claim 2, characterized in that The obtaining of a loop detection result based on the number of the traversed first objects includes: If the number of the first objects traversed is consistent with the number of objects in the computation graph, the loop detection result is that there is no loop dependency in the computation graph; If the number of the first objects traversed is inconsistent with the number of objects in the computation graph, the loop detection result is that the computation graph has a loop dependency.
4. The method according to claim 1, characterized in that: In the process of traversing the computation graph, performing address duplicate writing detection on the node based on the in-degree of the parameter to obtain the address duplicate writing detection result includes: In the process of traversing the computation graph, if there is an interface parameter of an output attribute, and the in-degree of the interface parameter is greater than the number of forward dependencies of the interface parameter, the address duplicate write detection result obtained is that there are at least two nodes writing to the same address; If there is an interface parameter of the output attribute, and the in-degree of the interface parameter is equal to the number of forward dependencies of the interface parameter, the address duplicate writing detection result obtained is that there are no at least two nodes writing to the same address.
5. The method according to claim 1, characterized in that The step of performing head node detection on the topologically sorted node set to obtain a head node detection result includes: In the process of traversing the computation graph, if the object is a node, the interface parameter of the node is placed in the node set of the computation graph to complete the layer-by-layer sorting of the nodes in the computation graph; The interface parameters of the nodes in the node set are traversed, and the node to which the interface parameters having only input attributes belong is determined as the head node, so as to obtain the head node detection result.
6. The method according to any one of claims 1 to 5, characterized in that: After verifying the computation graph based on the loop detection result, the address duplicate writing detection result and the head node detection result, the further step further includes: According to the topological relationship of each node, determine the first node, and execute the first node; the first node is the head node when it is executed for the first time; Determine a backward dependent node of the first node, reduce the in-degree of the backward dependent node by 1, use the backward dependent node as the new first node, and return to execute the step of executing the first node until all nodes in the computation graph are executed.
7. The method according to claim 6, characterized in that The method further comprises: Traverse each node in the topologically sorted computational graph and obtain the interface parameters of each node; If the interface parameter of the output attribute of the second node is the same as the interface parameter of the input attribute of the third node, then the in-degree of the third node is increased by 1; If the in-degree of the third node is greater than 1, a dependency relationship between the third node and the second node is established until all nodes in the computational graph after the topological sorting are traversed to obtain the topological relationship of each node.
8. A computational graph processing device, characterized in that: The computation graph includes nodes and interface parameters of the nodes, each node and each interface parameter is respectively regarded as an object; the device includes: A loop detection module, used for traversing the computation graph during topological sorting of the nodes, performing loop detection on the computation graph based on the in-degree of the object, and obtaining a loop detection result; the in-degree represents the number of forward dependencies of the object; An address duplicate writing detection module is used to perform address duplicate writing detection on the nodes based on the in-degree of the parameter in the process of traversing the calculation graph to obtain an address duplicate writing detection result; The head node detection module is used to perform head node detection on the node set after topological sorting to obtain the head node detection result; A calculation graph verification module is used to verify the calculation graph based on the loop detection result, the address duplicate writing detection result and the head node detection result.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Directed acyclic graph (DAG) ring formation detection method based on topological sorting
CN120430427A
A topological sorting-based directed acyclic graph (DAG) loop detection method
CN120430427B