Method and system for constructing a compiler intermediate representation based on a TensorFlow graph

By building a common intermediate representation and applying a recursive tracking algorithm, the problem of implicit representation of control flow in TensorFlow graph is solved, and efficient static analysis and optimization of the intermediate representation of the compiler is achieved.

CN114945898BActive Publication Date: 2025-07-11HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980102408.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-11-22
Publication Date
2025-07-11
Estimated Expiration
2039-11-22

AI Technical Summary

Technical Problem

The lack of explicitly marked control flows in TensorFlow graphs makes it difficult to efficiently apply static analysis and compiler optimization, especially granularity for loops and branches.

Method used

By building a common intermediate representation, the control flow in the TensorFlow graph is extracted and the recursive tracking algorithm is applied, which is converted into a compiler intermediate representation for static analysis and optimization.

Benefits of technology

It realizes effective static analysis and optimization of the compiler's intermediate representation in linear time, including dominant boundary analysis, common subexpression elimination and loop-invariant code extraction and other operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114945898B_ABST
    Figure CN114945898B_ABST
Patent Text Reader

Abstract

A method and system for constructing a common intermediate representation based on a TensorFlow graph are provided. The common intermediate representation can be converted into multiple compiler intermediate representations (IRs), enabling efficient application of compiler optimizations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for constructing a common intermediate representation according to a TensorFlow graph, which can be converted into multiple compiler intermediate representations. Background Art

[0002] TensorFlow is a machine learning platform that provides interfaces for constructing machine learning models and facilities for training and inferring these models. TensorFlow represents a machine learning model as a directed graph (referred to as a TensorFlow graph), which contains nodes and edges. The nodes are operators such as multiplication, addition, division, etc., and the edges represent data or control dependencies. There are no explicit markers for control flows such as loops and branches in the TensorFlow graph; instead, these control flows are implicitly represented by a set of control flow primitives, where these control flow primitives include transforms, merges, next iterations, etc. These control flow primitives are represented as nodes in the TensorFlow graph and produce high-level control flow behaviors such as loops and branches through separate and dynamic evaluations. Therefore, static analysis and compiler optimizations cannot be efficiently applied at the granularity of loops and branches on the TensorFlow graph. Summary of the Invention

[0003] The present invention generally relates to a method and system for constructing a common intermediate representation according to a TensorFlow graph, and the constructed common IR can be converted into an equivalent compiler intermediate representation, where compiler static analysis and optimizations, including but not limited to dominance frontier analysis, common subexpression elimination, loop invariant code motion, and dead code elimination, can be applied to the equivalent compiler intermediate representation in linear time.

[0004] According to one aspect of the present invention, a method is provided, the method including: receiving a TensorFlow graph including one or more control flows; processing the TensorFlow graph to extract the control flows from the TensorFlow graph and storing the extracted control flows in an extracted control flow data structure; applying a recursive tracing algorithm to the output of the TensorFlow graph to construct a common intermediate representation representing hierarchical control flows.

[0005] According to the above aspects, applying the recursive tracing algorithm includes: using the output of the TensorFlow graph as the current node; in response to determining that the current node is not a terminal node and the current node has not been traversed: adding the current node to a traversed data structure that indicates that the current node has been traversed; in response to determining that the current node is included in one of the plurality of control flows in the extracted control flow data structure: applying the tracing algorithm to the input of the control flow that includes the current node; adding the control flow that includes the current node to a traced data structure; in response to determining that the current node is not included in one of the plurality of control flows in the extracted control flow data structure: recursively applying the tracing algorithm to the input of the current node; adding the current node to the traced data structure; outputting the traced data structure, wherein the traced data structure is a list that includes TensorFlow graph nodes and common intermediate representation control flows.

[0006] According to any of the above aspects, the method further includes: constructing at least one compiler intermediate representation based on the common intermediate representation.

[0007] According to any of the above aspects, the control flow includes a plurality of TensorFlow graph branches, wherein data uses one of the branch paths that terminates at a merge node in two branch paths, and the control flow includes a plurality of loop condition nodes, wherein each loop condition node corresponds to a TensorFlow graph loop.

[0008] According to any of the above aspects, processing the TensorFlow graph to extract the control flow from the TensorFlow graph further includes: storing the merge nodes of each of the plurality of TensorFlow graph branches and the plurality of loop condition nodes in a collected node data structure.

[0009] According to the above aspects, the method further includes: (a) determining whether the collected node data structure is empty; (b) in response to determining that the collected node data structure is not empty: selecting nodes from the collected node data structure in an order that ensures all nodes are selected once before any node is selected a second time; (c) in response to determining that the selected node is a merge node: (i) extracting the corresponding TensorFlow graph branch; (ii) in response to determining that the TensorFlow graph branch is valid: constructing a common intermediate representation branch corresponding to the TensorFlow graph branch; storing the common intermediate representation branch in the extracted control flow data structure; removing the merge node from the collected node data structure; returning to (a); (iii) in response to determining that the TensorFlow graph branch is invalid, returning to (a); (d) in response to determining that the selected node is a loop condition node: (i) extracting the corresponding TensorFlow graph loop predicate and loop body; (ii) in response to determining that the TensorFlow graph loop predicate and loop body are valid: constructing a common intermediate representation loop corresponding to the TensorFlow graph loop; storing the common intermediate representation loop in the extracted control flow data structure; removing the loop condition node from the collected node data structure; returning to (a); (iii) in response to determining that the TensorFlow graph loop predicate and loop body are invalid: returning to (a).

[0010] According to the above aspects, extracting the corresponding TensorFlow graph branch includes: identifying the input nodes of the merge node by traversing the input edges of the merge node; recursively applying a tracing algorithm to the input nodes of the merge node until a branch exit condition is reached, and outputting two branch paths representing the TensorFlow graph branch.

[0011] According to the above aspects, the branch exit condition is reached when the tracing algorithm reaches an entry node, an exit node, another merge node, or a loop condition node.

[0012] According to the above aspects, the application tracing algorithm includes: using the input of the merging node as the current node; in response to determining that the current node is not a terminal node, the current node has not been traversed, and the branch exit condition is not satisfied: adding the current node to the traversed data structure, where the traversed data structure indicates that the current node has been traversed; in response to determining that the current node is included in one of the plurality of control flows in the extracted control flow data structure: applying the tracing algorithm to the input of the control flow containing the current node; adding the control flow containing the current node to the traced data structure; in response to determining that the current node is not included in one of the plurality of control flows in the extracted control flow data structure: applying the tracing algorithm to the input of the current node; adding the current node to the traced data structure; outputting the traced data structure, where the traced data structure is a list containing TensorFlow graph nodes and common intermediate representation control flows.

[0013] According to the above aspects, determining whether the two branch paths representing the TensorFlow graph branch are valid includes: determining that the two branch paths satisfy all of the following conditions: the two branch paths do not contain any merging nodes; the two branch paths do not contain any extracted branches that are part of a TensorFlow graph loop; the two branch paths output by the tracing algorithm terminate at the same merging node or at least one branch path terminates at a terminal node.

[0014] According to any of the above aspects, the method includes: removing all forwarding nodes included in the two branch paths; in response to determining that a first branch path among the two branch paths terminates at a conversion node included in a second branch path among the two branch paths: dividing the second branch path into a first slice and a second slice, where the first slice includes the conversion node and the nodes before the conversion node, the second slice includes the nodes after the conversion node, and replacing the second branch path with the second slice and discarding the first slice; storing the constructed common intermediate representation branch corresponding to the TensorFlow graph branch in the extracted control flow data structure, the common intermediate representation branch including a branch predicate and the two branch paths, where the branch predicate is the first input of the conversion node; in response to determining that the first branch path among the two branch paths does not terminate at a conversion node or terminates at a conversion node not included in the second branch path among the two branch paths: storing the constructed common intermediate representation branch corresponding to the TensorFlow graph branch in the extracted control flow data structure, the common intermediate representation branch including a branch predicate and the two branch paths.

[0015] According to any of the above aspects, extracting a corresponding TensorFlow graph loop includes: recursively applying a tracing algorithm to the input nodes of the loop condition node until a loop exit condition is reached; when the loop exit condition is reached, outputting the loop predicate of the TensorFlow graph loop; applying the recursive tracing algorithm to the output nodes of the loop condition node until the loop exit condition is reached; when the loop exit condition is reached, outputting the loop body of the TensorFlow graph loop.

[0016] According to any of the above aspects, when the tracing algorithm reaches an entry node, an exit node, or another loop condition node, the loop exit condition is reached.

[0017] According to any of the above aspects, recursively applying the tracing algorithm includes: using the input of the loop condition node as the current node; in response to determining that the current node is not a terminal node, the current node has not been traversed, and the loop exit condition is not satisfied: adding the current node to the traversed data structure, the traversed data structure indicating that the current node has been traversed; in response to determining that the current node is included in one of the plurality of control flows in the extracted control flow data structure: applying the tracing algorithm to the input of the control flow containing the current node; adding the control flow containing the current node to the traced data structure; in response to determining that the current node is not included in one of the plurality of control flows in the extracted control flow data structure: applying the tracing algorithm to the input of the current node; adding the current node to the traced data structure; outputting the traced data structure, wherein the traced data structure is a list containing TensorFlow graph nodes and common intermediate representation control flows.

[0018] According to any of the above aspects, applying the tracing algorithm includes: using the output of the loop condition node as the current node; in response to determining that the current node is not a terminal node, the current node has not been traversed, and the loop exit condition is not satisfied: adding the current node to the traversed data structure, the traversed data structure indicating that the current node has been traversed; in response to determining that the current node is included in one of the plurality of control flows in the extracted control flow data structure: applying the tracing algorithm to the input of the control flow containing the current node; adding the control flow containing the current node to the traced data structure; in response to determining that the current node is not included in one of the plurality of control flows in the extracted control flow data structure: applying the tracing algorithm to the input of the current node; adding the current node to the traced data structure; outputting the traced data structure, wherein the traced data structure is a list containing TensorFlow graph nodes and common intermediate representation control flows.

[0019] According to any of the above aspects, determining whether the TensorFlow graph loop is valid based on the loop predicate and the loop body includes: determining that the loop predicate and the loop body satisfy all of the following conditions: the TensorFlow graph loop predicate and loop body do not contain any loop condition nodes; the TensorFlow graph loop predicate and loop body do not contain merge nodes that are not part of the extracted branches.

[0020] According to any of the above aspects, the method includes: removing all forwarding nodes included in the loop predicate and the loop body; for each corresponding nested common intermediate representation branch in the loop predicate and the loop body: in response to determining that the corresponding nested common intermediate representation branch has an empty branch predicate: transferring all nodes in the common intermediate representation branch to the loop predicate or the loop body containing the branch; removing the common intermediate representation branch from the loop predicate or the loop body; storing the constructed common intermediate representation loop corresponding to the TensorFlow graph loop in the extracted control flow data structure, wherein the common intermediate representation loop is an object containing the loop predicate and the loop body.

[0021] According to one aspect of the present invention, there is provided a non-transitory computer-readable medium storing computer-readable instructions which, when executed by a processor of a processing system, cause the processing system to perform the following operations: receiving a TensorFlow graph containing one or more control flows; processing the TensorFlow graph to extract control flows from the TensorFlow graph; storing the extracted control flows in an extracted control flow data structure; applying a recursive tracing algorithm to the output of the TensorFlow graph using the extracted control flow data structure to construct a common intermediate representation representing hierarchical control flows.

[0022] According to the above aspect, the computer-readable medium stores further computer-readable instructions which, when executed by a processor of a processing system, cause the processing system to construct at least one compiler intermediate representation according to the constructed common intermediate representation.

[0023] According to one aspect of the present invention, there is provided a processing system including: a processor; a memory storing computer-readable instructions which, when executed by the processor, cause the processing system to perform the following operations: receiving a TensorFlow graph containing one or more control flows; processing the TensorFlow graph to extract control flows from the TensorFlow graph; storing the extracted control flows in an extracted control flow data structure; applying a recursive tracing algorithm to the output of the TensorFlow graph using the extracted control flow data structure to construct a common intermediate representation representing hierarchical control flows. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Some implementations of the present invention are described in conjunction with the following drawings.

[0025] Figure 1The block diagram of a system for generating a compiler intermediate representation according to a TensorFlow graph provided by an exemplary embodiment of the present invention is shown;

[0026] Figure 2 The flowchart of a method for constructing a common intermediate representation according to a TensorFlow graph provided by an exemplary embodiment of the present invention is shown;

[0027] Figure 3 The flowchart of a method for extracting control flow from a TensorFlow graph provided by an exemplary embodiment of the present invention is shown;

[0028] Figure 4 Shown is provided by an exemplary embodiment of the present invention Figure 1 The flowchart of the tracing algorithm executed by the common intermediate representation generator and the control flow extractor in

[0029] Figure 5 Shown is Figure 3 The exemplary workflow for extracting branches from a TensorFlow graph provided by the method in

[0030] Figure 6A and Figure 6B Shown is Figure 3 The exemplary workflow of loops in a TensorFlow graph provided by the method in

[0031] Figure 7 Shown is the block diagram of an exemplary computer that executes the Figure 2 method in Detailed Description

[0032] Throughout the drawings, the same reference numerals represent similar but not necessarily identical elements. The drawings are not necessarily to scale, and the dimensions of some components may be enlarged to more clearly illustrate the examples shown. Additionally, the drawings provide examples and / or implementations consistent with the description; however, the description is not limited to the examples and / or implementations provided in the drawings.

[0033] Throughout the drawings, the term "intermediate representation" is abbreviated as "IR".

[0034] In the present invention, unless otherwise clearly specified in the context, the terms "a" and "the" include plural meanings. Additionally, the terms "comprising", "including", or "having" used in the present invention are used to illustrate the presence of the described elements, but do not exclude the presence or addition of other elements.

[0035] In the present invention, a TensorFlow graph is a directed graph with nodes and edges, where the nodes represent computations or operations, and the edges represent data or control dependencies. Each node in the directed graph receives inputs and performs operations on the input data to generate outputs, which can be used as inputs for other operations.

[0036] In the present invention, an intermediate representation (IR) refers to a data structure that is organized to represent a source program constructed using a programming language or an intermediate representation builder interface.

[0037] In the present invention, a compiler intermediate representation (IR) refers to the intermediate representation design used by early compiler infrastructures such as the low level virtual machine (LLVM) and the GNU Compiler Collection (GCC). Examples of compiler intermediate representations include any or some combination of the following: Static Single Assignment (SSA), Administrative Normal Form (ANF), Continuation Passing Style (CPS), etc. These compiler intermediate representations can be pairwise converted to each other using known transformation algorithms. In other words, each of the above compiler intermediate representations can be converted to any other of the above compiler intermediate representations, and vice versa.

[0038] In the present invention, recursion refers to a method of solving a problem that depends on solutions to smaller instances of the same problem. A recursive algorithm refers to an algorithm that calls itself with smaller or simpler inputs until a base condition is reached.

[0039] In the present invention, a terminal node refers to a TensorFlow graph node that has no inputs. Generally, terminal nodes are TensorFlow graph nodes of the types constant and placeholder.

[0040] In the present invention, a forwarding node refers to a TensorFlow control flow primitive that forwards one of its inputs to its output node. Generally, forwarding nodes are TensorFlow graph nodes of the types switch, enter, exit, and next iteration, other than merge nodes.

[0041] In the present invention, the collected node data structure is obtained by method 300 (see Figure 3All instances of

[0042] In the present invention, the extracted control flow data structure is shared by all instances of method 300 (see Figure 3 ) and the tracing algorithm 400 (see Figure 4 ), which will be described in further detail below. The extracted control flow data structure is unordered and contains constructed common intermediate representation branches and loops. The extracted control flow data structure is associated with operations including: adding a new common intermediate representation branch or loop; enumerating existing common intermediate representation branches and loops.

[0043] In the present invention, the traversed data structure is shared by instances of the tracing algorithm 400 (see Figure 4 ) from the same recursive call. The traversed data structure is unordered and contains TensorFlow graph nodes traversed by instances of the tracing algorithm 400 (such as Figure 4 ). The traversed data structure is associated with operations including: adding a TensorFlow graph node; checking whether a TensorFlow graph node exists in the data structure.

[0044] In the present invention, the traced data structure is shared by instances of the tracing algorithm 400 (see Figure 4 ) from the same recursive call. The traced data structure is ordered and contains TensorFlow graph nodes recorded by instances of the tracing algorithm. The data structure is associated with operations including: appending a TensorFlow graph node to the end of the data structure; enumerating existing TensorFlow graph nodes; slicing the data structure at a given TensorFlow graph node. The slicing operation slices the data structure at a given TensorFlow graph node and obtains two slices: a first slice containing the given TensorFlow graph node and the nodes before it; a second slice containing the nodes after the given TensorFlow graph node.

[0045] The present invention generally relates to a method and system for constructing a common intermediate representation based on a TensorFlow graph, which can be converted into any compiler intermediate representation (IR).

[0046] Figure 1 FIG. shows a block diagram of a system for constructing one or more compiler intermediate representations based on a TensorFlow graph. The system 100 includes a number of software modules or subsystems, and the software modules or subsystems include: a control flow extractor 102; a common intermediate representation generator 104 operatively coupled to the control flow extractor 102; and a compiler intermediate representation generator 106 operatively coupled to the common intermediate representation generator 104. The control flow extractor 102 receives the TensorFlow graph, processes the TensorFlow graph to extract the control flow from the TensorFlow graph, and outputs an extracted control flow data structure. The common intermediate representation generator 104 receives the TensorFlow graph and the extracted control flow data structure from the control flow extractor 102, and uses the extracted control flow data structure to apply a recursive tracing algorithm to the output of the TensorFlow graph to construct a common intermediate representation that explicitly represents the hierarchical control flow. The compiler intermediate representation generator 106 receives the constructed common intermediate representation and constructs one or more compiler intermediate representations based on the common intermediate representation, which will be described in further detail below.

[0047] The common intermediate representation constructed by the common intermediate representation generator 104 can be in a form that can be converted into any one of a variety of different compiler intermediate representations.

[0048] Reference Figure 2 , shows an exemplary embodiment of a method 200 for constructing one or more compiler intermediate representations based on a TensorFlow graph.

[0049] The method 200 starts at step 202, in which a TensorFlow graph is received. One or more of the nodes and edges of the TensorFlow graph are used to indicate the control flow, including loops and branches. Then, the method 200 proceeds to step 204, in which the control flow in the TensorFlow graph is extracted by processing the TensorFlow graph using a method 300 ( Figure 3 ) which will be described in further detail below. Then, the method 200 proceeds to step 206, in which a recursive tracing algorithm 400 ( Figure 3) to construct a common intermediate representation and remove all forwarding nodes from the resulting common intermediate representation. Then, the method 200 proceeds to step 208, in which one or more compiler intermediate representations are constructed based on the common intermediate representation.

[0050] Now refer to Figure 3 , Figure 3 which describes the method 300 for extracting control flow from a TensorFlow graph provided by an embodiment of the present invention. The method 300 starts from step 304. In step 304, loop condition nodes and merge nodes are collected from the TensorFlow graph. By using the TensorFlow Python API, C / C++ API, or Protocol Buffer API to iterate through the nodes in the TensorFlow graph and check the operation type field of each node, loop condition nodes and merge nodes are collected from the TensorFlow graph; then, any loop condition nodes or merge nodes are stored in a collected node data structure for recording. After collecting the loop condition nodes and the merge nodes in step 304, the method 300 proceeds to step 306.

[0051] In step 306, it is determined whether the collected node data structure is empty. If the collected node data structure is empty, the method 300 ends. If the collected node data structure is not empty, the method 300 proceeds to step 308.

[0052] In step 308, nodes are selected from the collected node data structure in an order that ensures all nodes are selected once before any node is selected a second time. Then, the method 300 proceeds to step 310.

[0053] In step 310, it is determined whether the selected node is a merge node and whether the selected node is a loop condition node. If the selected node is a merge node, the method 300 proceeds to step 312. If the selected node is a loop condition node, the method 300 proceeds to steps 318 and 320.

[0054] In step 312, when the current node is an entry node, an exit node, a merge node (not the selected merge node), or a loop condition node and the exit condition is satisfied, the tracking algorithm 400 (see Figure 4 ) is recursively applied to the input nodes of the merge node to obtain two branch paths. Then, the method proceeds to step 314.

[0055] In step 314, determine whether the branch paths are valid. Determine whether the two branch paths are valid under the following conditions: the two branch paths do not contain any merge nodes; the two branch paths do not contain any extracted branches, where the extracted branches are part of a loop predicate or a loop body; the two branch paths output by the tracing algorithm terminate at the same merge node, or at least one branch path terminates at a terminal node. If it is determined in step 314 that the two branch paths are valid, method 300 proceeds to step 316. If it is determined in step 310 that the two branch paths are invalid, method 300 returns to step 306.

[0056] In step 316, remove all forwarding nodes from the valid branch paths. If one branch path terminates at a transition node contained in another branch path, split the other branch path into two slices: a first slice and a second slice, where the first slice contains the transition node and the nodes before it, and the second slice contains the nodes after the transition node; then, replace the other branch path with the second slice and discard the first slice. Construct a common intermediate representation branch based on the valid branch paths and branch predicates (if applicable). If both branches terminate at the same transition node, the first input of the transition node is the branch predicate. Then, add the constructed common intermediate representation branch to the extracted control flow data structure. Then, method 300 proceeds to step 326.

[0057] In step 318, when the current node is an entry node, an exit node, or a loop condition node (not the selected loop condition node) and the exit condition is satisfied, apply the recursive tracing algorithm 400 (see Figure 4 ) to the output node of the loop condition node to obtain the loop body. Then, method 300 proceeds to step 322.

[0058] In step 320, when the current node is an entry node, an exit node, or a loop condition node (not the selected loop condition node) and the exit condition is satisfied, apply the recursive tracing algorithm 400 (see Figure 4 ) to the input node of the loop condition node to obtain the loop predicate. Then, method 300 proceeds to step 322.

[0059] After steps 318 and 320 are completed, step 322 begins. In step 322, it is determined whether the loop predicate and the loop body are valid. The loop predicate and the loop body are determined to be valid under the following conditions: the loop predicate and the loop body do not contain any loop condition nodes; the loop predicate and the loop body do not contain merge nodes that are not part of the extracted branches. If it is determined in step 322 that the loop predicate and the loop body are valid, method 300 proceeds to step 324. If it is determined in step 322 that the loop predicate and the loop body are invalid, method 300 returns to step 306.

[0060] In step 324, all forwarding nodes are removed from the valid loop predicate and loop body; for all nested common intermediate representation branches in the loop predicate and the loop body: if the common intermediate representation branch has an empty branch predicate, all nodes in the branch are transferred to the loop predicate or the loop body that contains the branch; the common intermediate representation branch is removed from the loop predicate or the loop body. A common intermediate representation loop is constructed based on the valid loop predicate and loop body. Then, the constructed common intermediate representation loop is added to the extracted control flow data structure. Then, method 300 proceeds to step 326.

[0061] In step 326, the selected nodes are removed from the collected node data structure. Then, method 300 returns to step 306.

[0062] Now refer Figure 4 , Figure 4 An exemplary embodiment of the tracing algorithm 400 executed by the common intermediate representation generator and the control flow extractor in Figure 1 is shown. The tracing algorithm starts from step 402 and then proceeds to step 404. In 404, the nodes and the exit condition of the TensorFlow graph are received. The received nodes are designated as the current nodes. Then, the tracing algorithm 400 proceeds to step 406.

[0063] In step 406, it is determined whether the current node is a terminal node or a traversed node, or whether the exit condition has been met. If it is determined in step 406 that the current node is a terminal node or a traversed node, or the exit condition has been met, the tracing algorithm 400 proceeds to step 410, and then the tracing algorithm 400 ends. Otherwise, if it is determined in step 406 that the current node is not a terminal node, not a traversed node, and the exit condition has not been met, the tracing algorithm 400 proceeds to step 408.

[0064] In step 408, add the current node to the traversed data structure. Then, the tracing algorithm 400 proceeds to step 412. In step 412, determine whether the current node is included in any control flow in the extracted control flow data structure.

[0065] If it is determined in step 412 that the current node is included in any control flow in the extracted control flow data structure, the tracing algorithm 400 proceeds to step 414. If it is determined in step 412 that the current node is not included in any control flow in the extracted control flow data structure, the tracing algorithm 400 proceeds to step 420.

[0066] In step 414, recursively apply instances of the tracing algorithm 400 to each input of the control flow that contains the current node. After applying all instances of the tracing algorithm 400 to the inputs of the control flow that contains the current node, the tracing algorithm 400 ends. Then, the tracing algorithm 400 proceeds to step 416.

[0067] In step 416, add the control flow in the extracted control flow data structure that contains the current node to the traced data structure, and then the tracing algorithm 400 ends.

[0068] In step 420, recursively apply instances of the tracing algorithm 400 to each input of the current node. After recursively applying all instances of the tracing algorithm 400 to the inputs of the current node, the tracing algorithm 400 ends. Then, the tracing algorithm 400 proceeds to step 422.

[0069] In step 422, add the current node to the traced data structure, and then the tracing algorithm 400 ends.

[0070] Reference Figure 5 , Figure 5 describes a method 300 for extracting control flow from an exemplary TensorFlow graph 500 that includes branches. Figure 5 The exemplary TensorFlow graph 500 shown includes 2 TensorFlow control flow primitives of the types conversion and merge; additionally, Figure 5 includes 5 TensorFlow operations of the types greater, Relu, max pooling, and average pooling.

[0071] Apply the method 300 to the exemplary TensorFlow graph 500. In step 304 of the method 300, store the merge node 504 in the collected node data structure. Then, the method 300 proceeds to step 306.

[0072] As long as the merge node 504 has not been processed and has not been removed from the collected node data structure, the method 300 proceeds from step 306 to step 308.

[0073] In step 308 of the method 300, the merge node 504 is selected from the collected node data structure. Then, the method 300 proceeds to step 312.

[0074] At step 312 of the method 300, an instance of the tracing algorithm 400 is recursively applied to each input node of the merge node 504. The max pooling node 508 is the first input of the merge node 504, and the instance of the tracing algorithm 400 applied to the max pooling node 508 traces to the constant nodes 512, 514, and 516, where a terminal node is reached. A first branch path is obtained according to the traced data structure, which is generated at the end of the instance of the tracing algorithm 400. The average pooling node 510 is the second input of the merge node 504, and the instance of the tracing algorithm 400 applied to the average pooling node 510 traces to the conversion node 502, where the tracing ends when a traversed node is reached. A second branch path is obtained according to the traced data structure, which is generated at the end of the instance of the tracing algorithm 400. The method 300 proceeds to step 314.

[0075] In step 314 of the method 300, it is determined that the two branch paths are valid. The method 300 proceeds to step 316.

[0076] In step 316 of the method 300, the forwarding node is removed; in this case, there is no forwarding node. If one branch path terminates at a conversion node contained in the other branch path, the other branch path is split into two slices: a first slice and a second slice, where the first slice contains the conversion node and the nodes before it, and the second slice contains the nodes after the conversion node; then, the other branch path is replaced with the second slice and the first slice is discarded. In this case, the branch path obtained by tracing the max pooling node 508 after the conversion node 502 is sliced and replaced with the second slice. The constructed form of the common intermediate representation branch is branch 600, where the larger node 506 (i.e., the first input of the conversion node 504 that is the terminal node of the two branch paths) serves as the branch predicate. The branch 600 is added to the extracted control flow data structure. Then, the method 300 proceeds to step 326.

[0077] In step 326 of the method 300, the merge node 504 is removed from the collected node data structure. The method 300 returns to step 306, in which the collected node data structure is empty. Then, the method 300 ends.

[0078] Reference Figure 6A , Figure 6A describes a method 300 for extracting control flow from an exemplary TensorFlow graph 700 that contains loops. Figure 6A The exemplary TensorFlow graph 700 shown contains 10 TensorFlow control flow primitives of types enter, next iteration, exit, loop condition, transition, and merge; additionally, Figure 6A contains 8 TensorFlow operations of types constant, multiplication, addition, and subtraction.

[0079] The method 300 is applied to the exemplary TensorFlow graph 700. In step 304 of the method 300, the merge nodes (706, 704) and the loop condition node 702 are stored in the collected node data structure. Then, the method 300 proceeds to step 306.

[0080] As long as any of the merge nodes (706, 704) and the loop condition node 702 have not been processed and removed from the collected node data structure, the method 300 proceeds from step 306 to step 308.

[0081] In step 308 of the method 300, a node is selected from the collected node data structure. The following description illustrates the scenario when nodes are selected and processed in the order of the merge node 706, the merge node 704, and the loop condition node 702. Any order of selecting nodes can ensure that all nodes are selected once before any node is selected a second time; however, if the loop condition node 702 is selected before the merge node 706 or the merge node 704, the resulting loop predicate and loop body are invalid because the loop predicate and the loop body contain nested merge nodes (706, 704); then, the method 300 returns to step 306. Therefore, ensure that the nested merge nodes (706, 704) are always processed before the loop condition node 702.

[0082] The merge node 706 is selected. Then, the method 300 proceeds to step 310, in which it is determined that the selected node is a merge node. Then, the method 300 proceeds to step 312.

[0083] In step 312 of the method 300, an instance of the tracing algorithm 400 is recursively applied to each input node of the merge node 706. The next iteration node 708 is the first input to the merge node 706, and the instance of the tracing algorithm 400 applied to the next iteration node 708 traces to the merge node 706 and the constant node 712, where the exit conditions are satisfied by reaching the selected merge node and by reaching a terminal node, respectively. A branch path is obtained according to the traced data structure, which is generated at the end of the instance of the tracing algorithm 400. The entry node 710 is the second input to the merge node 706, and the instance of the tracing algorithm 400 applied to the entry node 710 ends immediately when the exit condition is satisfied upon reaching the entry node. A second branch path is obtained according to the traced data structure, which is generated at the end of the instance of the tracing algorithm 400. The method 300 proceeds to step 314.

[0084] In step 314 of the method 300, it is determined that the two branch paths are valid. The method 300 proceeds to step 316.

[0085] In step 316 of the method 300, forwarding nodes such as the next iteration node 708 are removed. The constructed form of the common intermediate representation is a branch 802, where the branch predicate is empty because the branch path does not terminate at a conversion node. The branch 802 is added to the extracted control flow data structure. Then, the method 300 proceeds to step 326.

[0086] In step 326 of the method 300, the merge node 706 is removed from the collected node data structure. Then, the method 300 returns to step 306.

[0087] The merge node 704 is selected. Then, the method 300 proceeds to step 310, in which it is determined that the selected node is a merge node. Then, the method 300 proceeds to step 312.

[0088] In step 312 of the method 300, an instance of the tracing algorithm 400 is recursively applied to each input node of the merge node 704. The next iteration node 714 is the first input of the merge node 704, and the instance of the tracing algorithm 400 applied to the next iteration node 714 traces to the merge node 704, the loop condition node 702, the exit node 718, and the constant node 720, where the exit condition is satisfied by reaching the selected merge node, loop condition node, and exit node respectively, and by reaching the terminal node. The branch path is obtained according to the traced data structure, which is generated at the end of the instance of the tracing algorithm 400. The entry node 716 is the second input of the merge node 704, and the instance of the tracing algorithm 400 applied to the entry node 710 ends immediately when the exit condition is satisfied upon reaching the entry node. The second branch path is obtained according to the traced data structure, which is generated at the end of the instance of the tracing algorithm 400. The method 300 proceeds to step 314.

[0089] In step 314 of the method 300, it is determined that the two branch paths are valid. The method 300 proceeds to step 316.

[0090] In step 316 of the method 300, forwarding nodes such as the next iteration node 714 are removed. The constructed form of the common intermediate representation is the branch 804, where the branch predicate is empty because the branch path does not terminate at a transformation node. The branch 804 is added to the extracted control flow data structure. Then, the method 300 proceeds to step 326.

[0091] In step 326 of the method 300, the merge node 704 is removed from the collected node data structure. Then, the method 300 returns to step 306.

[0092] The loop condition node 702 is selected. Then, the method 300 proceeds to step 310, in which it is determined that the selected node is a loop condition node. Then, the method 300 proceeds to steps 318 and 320.

[0093] In step 318 of the method 300, an instance of the tracing algorithm 400 is recursively applied to each output node of the loop condition node 702. The transformation node 722 is the only output of the loop condition node 702, and the instance of the tracing algorithm 400 applied to the transformation node 722 traces to the entry node 716, the transformation node 722, and the loop condition node 702, where the exit condition is satisfied by reaching the entry node, the traversed nodes, and the selected loop condition node, respectively. The instance of the tracing algorithm 400 also encounters the extracted branch 804 and adds the extracted branch 804 to the traced data structure. The loop body is obtained according to the traced data structure, which is generated at the end of the instance of the tracing algorithm 400. The method 300 proceeds to step 322.

[0094] In step 320 of the method 300, an instance of the tracing algorithm 400 is recursively applied to the input nodes of the loop condition node 702. The subtraction node 724 is the input of the loop condition node 702, and the instance of the tracing algorithm 400 applied to the subtraction node 724 traces to the entry node 726 and the entry node 710, where the exit condition is satisfied by reaching the entry node. The instance of the tracing algorithm 400 also encounters the extracted branch 802 and adds the extracted branch 802 to the traced data structure. The loop predicate is obtained according to the traced data structure, which is generated at the end of the instance of the tracing algorithm 400. The method 300 proceeds to step 322.

[0095] In step 322 of the method 300, it is determined that the loop predicate and the loop body are valid. The loop predicate and the loop body currently take the form of loop 800. The method 300 proceeds to step 324.

[0096] In step 324 of the method 300, all forwarding nodes are removed from the loop predicate and the loop body; in this case, there are no forwarding nodes. For all nested common intermediate representation branches (802, 804) in the loop predicate and the loop body: If the common intermediate representation branch has an empty branch predicate, all nodes in the branch are transferred to the loop predicate or the loop body containing the branch; the common intermediate representation branch is removed from the loop predicate or the loop body; in this case, both of the nested branches (802, 804) have empty branch predicates. The constructed form of the common intermediate representation loop is loop 900. The loop 900 is added to the extracted control flow data structure. Then, the method 300 proceeds to step 326.

[0097] In step 326 of the method 300, the loop condition node 702 is removed from the collected node data structure. The method 300 returns to step 306, in which the collected node data structure is empty. Then, the method 300 ends.

[0098] Reference Figure 6B , Figure 6B is Figure 6A a continuation of, in which the exemplary common intermediate representation loop 900 is obtained by applying the method 300 to the exemplary TensorFlow graph 700. The process of converting the common intermediate representation loop 900 into a valid SSA intermediate representation is described below.

[0099] Define the common intermediate representation of the SSA tracking algorithm. Given a common intermediate representation node, the tracking algorithm will recursively traverse the inputs of the given node. The tracking algorithm also takes as input a basic block, which contains a list of instructions and input and output edges connecting to other basic blocks, as shown in the exemplary SSA loop 1000. If the given node is not a merge node, an instruction equivalent to the operation of the given node is added to the beginning of the given basic block, and the given basic block is provided as input to the recursive tracking algorithm, where the recursive tracking algorithm is applied to the inputs of the given node. If the given node is a merge node, an SSAPhi node equivalent to the given merge node is added to the beginning of the given basic block, and 2 new basic blocks are created as inputs to the given basic block; the new basic blocks are provided as input to the corresponding recursive tracking algorithm, where the corresponding recursive tracking algorithm is applied to the inputs of the given merge node. If the given node has been traversed previously, the basic block containing the given node is added as input to the given basic block, and then the tracking algorithm ends. If the given node is a terminal node, the tracking algorithm ends.

[0100] The conversion of the exemplary common intermediate representation loop 900 to the equivalent SSA intermediate representation 1000 is described below. By applying the common intermediate representation of the SSA tracking algorithm to the subtraction node 906 (i.e., the output of the loop predicate 902), the basic block 1006 containing the equivalent SSA expression of the subtraction node 906 and the merge node 908 is obtained; then, by recursively tracking the addition node 910 (i.e., the first input of the merge node 908), the basic block 1002 is obtained; then, by recursively tracking the constant node 912 (the second input of the merge node 908), the basic block 1004 is obtained. By applying the common intermediate representation of the SSA tracking algorithm to the merge node 914 (i.e., the output of the loop body 904), the basic block 1012 is obtained; then, by recursively tracking the multiplication node 916 (i.e., the first input of the merge node 914), the basic block 1008 is obtained; then, by recursively tracking the constant node 918 (i.e., the second input of the merge node 914), the basic block 1010 is obtained. Finally, a control edge is added between the basic block 1006 (i.e., the output block of the loop predicate identified by looking up the basic block containing the output node of the loop predicate 902) and the basic block 1008 (i.e., the input block of the loop body identified by looking up the basic block containing the terminal node in the loop body 904); an output edge indicating loop exit is added to the basic block 1006; a back edge indicating the next iteration is added between the basic block 1012 (i.e., the output of the loop body identified by looking up the basic block containing the output node of the loop body 904) and the basic block 1002 (i.e., the input block of the loop predicate identified by looking up the basic block containing the terminal node in the loop predicate 906).

[0101] Figure 7 FIG. shows a block diagram of an exemplary computing device 1100 that may be used to implement embodiments of the methods disclosed herein. The computing device 1100 includes a plurality of components, including one or more processors 1102 that control the overall operation of the computing device 1100. The one or more processors 1102 are coupled to and interact with other components of the computing device 1100, which include: one or more non-transitory memory / storage units 1104; one or more volatile memories 1106; one or more input / output (I / O) interfaces 1108 that may enable connection to one or more optional input devices 1110 and / or output devices 1112; and one or more network interfaces 1114.

[0102] The one or more processors 1102 may include a microprocessor, a central processing unit (CPU), a hardware accelerator, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), dedicated logic circuitry, or a combination thereof.

[0103] The one or more non-transitory memories / storage units 1104 may include a mass storage unit (e.g., a solid-state disk, a hard disk drive, a disk drive, and / or an optical disk drive) and one or more non-volatile memories (e.g., flash memory and / or read-only memory (ROM)). The one or more volatile memories 1106 may include random access memory (RAM). The one or more non-transitory memories / storage units 1104 may store software (e.g., software implementing the methods (200, 300, 400) described in the present invention) to be executed by the one or more processors 1102, as well as other software instructions (e.g., software instructions for implementing an operating system and other applications / functions). The software implementing the methods (200, 300, 400) includes computer-readable instructions executable by the one or more processors 1102. The coding of the software implementing the methods (200, 300, 400) is within the scope of those of ordinary skill in the art provided in the present invention. The methods (200, 300, 400) may include more or fewer steps than Figure 2 , Figure 3 and Figure 4 the steps shown and / or the above steps, and the steps may be executed in a different order. In some examples, the instructions stored in the non-transitory memories / storage units 1104 may be temporarily loaded into the one or more volatile memories 1106 for execution by the processor 1102.

[0104] The one or more network interfaces 1114 are used for wired or wireless communication with a network (e.g., an intranet, the Internet, a P2P network, a WAN, and / or a LAN) or other nodes. The one or more network interfaces 1114 may include a wired link (e.g., an Ethernet cable) and / or a wireless link (e.g., one or more antennas) for intra-network and / or inter-network communication.

[0105] In some other examples, one or more data sets and / or modules may be provided by an external memory (e.g., an external drive that communicates with the computing device 1100 wired or wirelessly), or may be provided by a transient or non-transitory computer-readable medium. Examples of non-transitory computer-readable media include RAM, ROM, erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, CD-ROM, or other portable memory.

[0106] The computing device 1100 also includes a bus 1116 to provide communication between the components of the computing device 1100, including the one or more processors 1102, the one or more optional I / O interfaces 1108, the one or more network interfaces 1114, the non-transitory memory / storage unit 1104, and / or the one or more volatile memories 1106. The bus 1116 can be any suitable bus architecture, e.g., including a memory bus, a peripheral bus, or a video bus.

[0107] In Figure 7 , one or more optional input devices 1110 (e.g., a keyboard, a mouse, a microphone, a touch screen integrated in or covering a display device, and / or keys) and one or more optional output devices 1112 (e.g., a display device, a speaker, and / or a printer) are shown as external devices of the computing device 1100. In other examples, one or more of the one or more input devices 1110 and / or one or more of the one or more output devices 1112 may be internal components of the computing device 1100. The one or more input devices 1110 may include: a display device having a display screen; a user interface (UI) navigation input device (e.g., a touch screen, a mouse, a touchpad, a voice interface, or other interface device) for allowing a user to interact with items displayed by the display device.

[0108] Although Figure 7 a physical computing device 1100 is shown, it should be understood that embodiments of the methods and systems disclosed herein may be implemented on one or more virtual machines instantiated by a cloud service provider or a distributed computing system. Alternatively, the methods and systems disclosed herein may be implemented as services provided by a cloud computing provider.

[0109] In the foregoing description, numerous specific details are set forth to provide an understanding of the subject matter disclosed herein. However, implementations may be practiced without these specific details. Other implementations may include modifications and variations to the above details. The appended claims cover all such modifications and variations.

[0110] Advantageously, the methods and systems of the present invention convert a TensorFlow graph into an equivalent compiler intermediate representation. Compiler static analysis and optimizations, including but not limited to dominance frontier analysis, common subexpression elimination, loop invariant code motion, and dead code elimination, can be applied to the equivalent compiler intermediate representation in linear time. Additionally, existing analyses, libraries, algorithms, and techniques applied to the compiler intermediate representation can also be equivalently applied to the program represented by the TensorFlow graph, thereby reducing engineering costs.

Claims

1. A method for constructing a compiler intermediate representation according to a TensorFlow graph, characterized in that The method includes: Receiving a TensorFlow graph that includes one or more control flows; Processing the TensorFlow graph to extract the control flows from the TensorFlow graph and storing the extracted control flows in an extracted control flow data structure; Using the extracted control flow data structure to recursively apply a tracing algorithm to the output of the TensorFlow graph to build a common intermediate representation that represents a hierarchical control flow; The method further includes: Removing all forwarding nodes included in the TensorFlow graph loop predicate and the loop body; For each corresponding nested common intermediate representation branch in the loop predicate and the loop body: in response to determining that the corresponding nested common intermediate representation branch has an empty branch predicate, transferring all nodes in the common intermediate representation branch to the loop predicate or the loop body that contains the branch, and removing the common intermediate representation branch from the loop predicate or the loop body; Storing the common intermediate representation loop corresponding to the TensorFlow graph loop in the extracted control flow data structure, where the common intermediate representation loop includes the loop predicate and the loop body; Constructing at least one compiler intermediate representation according to the common intermediate representation.

2. The method according to claim 1, characterized in that Recursively applying the tracing algorithm includes: Using the output of the TensorFlow graph as the current node; In response to determining that the current node is not a terminal node and the current node is not a traversed node: Adding the current node to a traversed data structure that indicates that the current node has been traversed; In response to determining that the current node is included in one of the multiple control flows in the extracted control flow data structure: Recursively applying the tracing algorithm to the input of the control flow that contains the current node; Adding the control flow that contains the current node to a traced data structure; In response to determining that the current node is not included in one of the multiple control flows in the extracted control flow data structure: Recursively applying the tracing algorithm to the input of the current node; Adding the current node to the traced data structure; Outputting the traced data structure, where the traced data structure is a list that includes TensorFlow graph nodes and common intermediate representation control flow objects.

3. The method according to claim 1, wherein The control flow includes multiple TensorFlow graph branches, where data uses one of the branch paths that terminates at a merge node in two branch paths, and the control flow includes multiple loop condition nodes, where each loop condition node corresponds to a TensorFlow graph loop.

4. The method according to claim 3, characterized in that Processing the TensorFlow graph to extract the control flow from the TensorFlow graph further includes: storing the merge node of each of the multiple TensorFlow graph branches and the multiple loop condition nodes in a collected node data structure.

5. The method according to claim 4, characterized in that, The method further includes: (a) Determining whether the collected node data structure is empty; (b)In response to determining that the collected node data structure is not empty: Select nodes from the collected node data structure in an order that ensures all nodes are selected once before any node is selected a second time; In response to determining that the selected node is a merge node: (i)Extract the corresponding TensorFlow graph branch; (ii)In response to determining that the TensorFlow graph branch is valid: Construct a common intermediate representation branch corresponding to the TensorFlow graph branch; Store the common intermediate representation branch in the extracted control flow data structure; Remove the merge node from the collected node data structure; Return to (a); (iii)In response to determining that the TensorFlow graph branch is invalid: Return to (a); In response to determining that the selected node is a loop condition node: (i)Extract the corresponding TensorFlow graph loop predicate and loop body; (ii)In response to determining that the TensorFlow graph loop predicate and loop body are valid: Construct a common intermediate representation loop corresponding to the TensorFlow graph loop; Store the common intermediate representation loop in the extracted control flow data structure; Remove the loop condition node from the collected node data structure; Return to (a); (iii)In response to determining that the TensorFlow graph loop predicate and loop body are invalid: Return to (a).

6. The method according to claim 5, wherein Extracting the corresponding TensorFlow graph branch includes: Identifying the input nodes of the merge node by traversing the input edges of the merge node; Recursively applying a tracing algorithm to the input nodes of the merge node until a branch exit condition is reached, and outputting two branch paths representing the TensorFlow graph branch.

7. The method according to claim 6, characterized in that, The branch exit condition is reached when the tracing algorithm reaches an entry node, an exit node, another merge node, or a loop condition node.

8. The method according to claim 6, characterized in that, Recursively applying the tracing algorithm includes: Using the input of the merge node as the current node; In response to determining that the current node is not a terminal node, the current node has not been traversed, and the branch exit condition is not met: Adding the current node to the traversed data structure, which indicates that the current node has been traversed; In response to determining that the current node is included in one of the multiple control flows in the extracted control flow data structure: Recursively applying the tracing algorithm to the input of the control flow containing the current node; Adding the control flow containing the current node to the traced data structure; In response to determining that the current node is not included in one of the multiple control flows in the extracted control flow data structure: Recursively applying the tracing algorithm to the input of the current node; Adding the current node to the traced data structure; Outputting the traced data structure, where the traced data structure is a list containing TensorFlow graph nodes and common intermediate representation control flow objects.

9. The method according to claim 6, characterized in that, Determining whether the two branch paths representing the TensorFlow graph branch are valid includes: determining that the two branch paths satisfy all of the following conditions: The two branch paths do not contain any merge nodes; The two branch paths do not contain any extracted branches that are part of a TensorFlow graph loop; The two branch paths output by the tracing algorithm terminate at the same merge node, or at least one branch path terminates at a terminal node.

10. The method according to claim 6, characterized in that, The method further includes: Removing all forwarding nodes contained in the two branch paths; In response to determining that the first branch path of the two branch paths terminates at a transformation node contained in the second branch path of the two branch paths: dividing the second branch path into a first slice and a second slice, where the first slice contains the transformation node and the nodes before the transformation node, the second slice contains the nodes after the transformation node, and replacing the second branch path with the second slice and discarding the first slice; Storing the common intermediate representation branch in the extracted control flow data structure, the common intermediate representation branch containing a branch predicate and the two branch paths, where if both the first branch and the second branch terminate at the transformation node, the branch predicate is the first input of the transformation node.

11. The method according to claim 5, wherein Extracting the corresponding TensorFlow graph loop includes: Recursively applying the tracing algorithm to the input nodes of the loop condition node until a loop exit condition is reached; When the loop exit condition is reached, outputting the loop predicate of the TensorFlow graph loop; Applying the recursive tracing algorithm to the output nodes of the loop condition node until the loop exit condition is reached; When the loop exit condition is reached, outputting the loop body of the TensorFlow graph loop.

12. The method according to claim 11, wherein The loop exit condition is reached when the tracing algorithm reaches an enter node, an exit node, or another loop condition node.

13. The method according to claim 11, wherein Recursively applying the tracing algorithm includes: Using the input of the loop condition node as the current node; In response to determining that the current node is not a terminal node, the current node has not been traversed, and the loop exit condition is not satisfied: Adding the current node to the traversed data structure, the traversed data structure indicating that the current node has been traversed; In response to determining that the current node is included in one of the plurality of control flows in the extracted control flow data structure: Applying the tracing algorithm to the input of the control flow containing the current node; Adding the control flow containing the current node to the traced data structure; In response to determining that the current node is not included in one of the plurality of control flows in the extracted control flow data structure: Applying the tracing algorithm to the input of the current node; Adding the current node to the traced data structure; Output the traced data structure, where the traced data structure is a list containing TensorFlow graph nodes and common intermediate representation control flow objects.

14. The method according to claim 11, wherein The application tracing algorithm includes: Using the output of the loop condition node as the current node; In response to determining that the current node is not a terminal node, the current node has not been traversed, and the loop exit condition is not satisfied: Adding the current node to the traversed data structure, where the traversed data structure indicates that the current node has been traversed; In response to determining that the current node is included in one of the plurality of control flows in the extracted control flow data structure: Applying the tracing algorithm to the inputs of the control flow containing the current node; Adding the control flow containing the current node to the traced data structure; In response to determining that the current node is not included in one of the plurality of control flows in the extracted control flow data structure: Applying the tracing algorithm to the inputs of the current node; Adding the current node to the traced data structure; Output the traced data structure, where the traced data structure is a list containing TensorFlow graph nodes and common intermediate representation control flow objects.

15. The method according to claim 5, wherein Determining whether the TensorFlow graph loop is valid based on the loop predicate and the loop body includes: determining that the loop predicate and the loop body satisfy all of the following conditions: The TensorFlow graph loop does not contain any loop condition nodes; The TensorFlow graph loop does not contain a merge node of the TensorFlow graph branch.

16. A non-transitory computer-readable medium, characterized in that, The non-transitory computer-readable medium stores computer-readable instructions that, when executed by a processor of a processing system, cause the processing system to perform the method according to any one of claims 1 to 15.

17. A computer program, characterized in that, The computing program includes computer-readable instructions that, when executed by a processor of a processing system, cause the processing system to perform the method according to any one of claims 1 to 15.

18. A processing system, characterized in that, The processing system includes: A processor; A memory storing computer-readable instructions that, when executed by the processor, cause the processing system to perform the method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Tool to create reconfigurable interconnect framework

    CN108268940A

  • Fine-grain compute communication execution for deep learning frameworks

    CN108805798A