Loop and library fusion
By analyzing computation graphs to identify fusible operations and compiling them into single fusion operations, the system generates more efficient code, addressing the inefficiencies of existing computation graph compilation systems.
Patent Information
- Application Number
- DE102018100239
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-01-06
- Filing Date
- 2018-01-08
- Publication Date
- 2025-05-08
- Estimated Expiration
- 2038-01-08
AI Technical Summary
Existing computation graph compilation systems do not efficiently generate optimized code by merging fusible operations into single fusion operations, leading to suboptimal performance and memory usage.
A system and method that analyze an unoptimized computation graph using pattern matching to identify fusible operations, transform the graph by replacing these operations with fusion nodes, and compile the optimized graph into efficient code that can be executed by computing devices.
The approach generates faster and more memory-efficient compiled code by merging operations into single fusion calls, improving performance and reducing code size compared to conventional compilation methods.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND
[0001] This patent description relates to the construction and compilation of computational graphs.
[0002] A computational graph defines sequences of operations by the types of operations, the data input to and output from each operation, and computational dependencies. A compiler translates a computational graph of operations to produce compiled code.
[0003] US20040139429 A1 teaches a system and method for generating a fused instruction, wherein a first instruction and a second instruction, both of which are simple instructions, e.g., performing only one operation, and are dependent on each other, are fused together to generate the fused instruction, wherein the fused instruction has an opcode representing the operation performed by the first instruction and the operation performed by the second instruction.
[0004] US 2003 / 0208723 A1 teaches that an automated processor design tool uses a description of user-defined processor instruction set extensions in a standardized language to develop a configurable definition of a target instruction set, a hardware description language description of the circuitry required to implement the instruction set, and development tools such as compilers, assemblers, debuggers, and simulators that can be used to develop and verify applications for the processor. SUMMARY
[0005] The invention is defined by the independent claims. Dependent claims specify advantageous embodiments.
[0006] This patent specification describes technologies related to computational graph systems in general, and more particularly to systems and methods for representing computations as graph operations that can be translated into efficient compiled code.
[0007] The computational graph comprises nodes, directed connecting edges, and directed parameter edges. Each node represents a respective operation. Each directed connecting edge connects a respective first node to a respective second node representing an operation that receives as input an output of an operation represented by the respective first node. Each directed parameter edge connects to a respective node and represents a flow of one or more parameters of a neural network as input to the operation represented by the respective node.
[0008] In general, an innovative aspect of the subject matter described in this specification may be embodied in a system having one or more computers and one or more memory devices storing instructions operable when executed by the one or more computers to cause the one or more computers to perform operations implementing an example method.An example method includes: obtaining a non-optimized computational graph having a plurality of nodes representing operations and directed edges representing data dependencies; analyzing the non-optimized computational graph using pattern matching to determine fusible operations that can be fused together into a single fusion operation; transforming the non-optimized computational graph into an optimized computational graph by replacing the nodes representing the fusible operations in the non-optimized computational graph with a fusion node representing the single fusion operation; and providing the fusion node of the optimized computational graph to a compiler, which the compiler can translate as a call that performs the fused operations to produce efficient code.
[0009] Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each configured to perform the acts of the methods. For a system of one or more computers to be configured to perform specific operations or acts means that the system has software, firmware, hardware, or a combination of them installed that, in operation, causes the system to perform the operations or acts. For one or more computer programs to be configured to perform specific operations or acts means that the one or more programs comprise instructions that, when executed by a data processing device, cause the device to perform the operations or acts.
[0010] These and other embodiments may optionally include one or more of the following features. The efficient code may be delivered to computing devices for execution. Execution may include executing the operations of the computational graph with the single fusion call that performs all fused operations.Analyzing the unoptimized computational graph using pattern matching to determine fused operations that can be fused together into a single fusion operation includes: comparing portions of the unoptimized computational graph with patterns of operations that each correspond to a single fusion operation; determining that a pattern corresponds to a portion of the unoptimized computational graph; and determining that the corresponding portion of the unoptimized computational graph can be replaced in the computational graph with the single fusion operation that corresponds to the corresponding pattern. The single fusion operation can be an external code library operation. The single fusion operation can be a loop operation.Analyzing the non-optimized computational graph using pattern matching to determine fusible operations that can be fused together into a single fusion operation may include: searching the non-optimized computational graph for an input operation that requires computations to generate the input; and determining that the input operation in the computational graph can be replaced by a single fusion operation that corresponds to the computations required to generate the input. The fusible operations may be regular operations. The fusible operations may be regular operations that are fused together into non-regular operations.Analyzing the unoptimized computational graph using pattern matching to determine fusible operations that can be fused together into a single operation may comprise: finding a sequence of operations in a computational graph using a sequencing algorithm; and determining that the sequence of operations can be fused together into a single fusion operation using composition.
[0011] The subject matter described in this specification may be implemented in specific embodiments to achieve one or more of the following advantages. An example implementation produces efficient compiled code by combining operations into a single fused operation that a compiler can translate as a single call, such as a loop or library call. For the purposes of this specification, efficient compiled code means code that is faster and potentially uses less memory than code compiled using a conventional compiler.
[0012] The single call into which the compiler translates the single fused operation performs all operations of the fused operation in a single code-generating compilation phase. This translation allows the compiler to generate code that is faster than code generated by a conventional compiler that translates one operation at a time. By using fused operations, which can use less memory than non-fused operations, the compiler also generates code that is more memory-efficient than code generated by a conventional compiler.
[0013] By merging operations, a sample implementation can also provide programs that are smaller than a conventional compiler.
[0014] The details of one or more embodiments of the subject matter of this patent specification are set forth in the accompanying drawings and the description below. Further features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1 shows an example compilation system. Fig. Figure 2 is a flow diagram of an example process for generating efficient code from computations within a computational graph. Fig. 3 represents a graphical chain of two regular unary operations. Fig. 4 represents operations of non-unary operations that can be merged together. Fig. Figure 5a shows a subgraph of a computational graph representing a transpose and a point operation. Fig. 5b represents an optimized version of the Fig. 5a represents the subgraphs shown. Fig. Figure 6a shows a pattern representing a backward filter convolution. Fig. Figure 6b shows a pattern representing a backward input convolution.
[0015] The same reference symbols and designations in the different drawings indicate the same elements. DETAILED DESCRIPTION
[0016] This patent specification generally describes compilation systems that compile operations represented by a computational graph. More specifically, this patent specification describes techniques for generating efficient code from computations within a computational graph by fusing computations together.
[0017] A computational dataflow graph expresses computations, e.g., of a machine learning model, with nodes representing operations and directed edges representing data dependencies between operations. An incoming edge to a node represents a flow of input to the node—that is, an input argument to the operation represented by the node. When all arguments required for an operation are available to the operation node, the node is activated and can be executed.
[0018] An outgoing edge from a node represents a flow of an output of the operation represented by the node to be used as input to an operation represented by another node. Consequently, a directed edge connecting a first node in the graph to a second node in the graph indicates that an output produced by the operation represented by the first node is used as input to the operation represented by the second node.
[0019] In some implementations, the operations represented in the computational graph are linear algebraic operations, such as matrix multiplication, neural network operations, or operations for another type of machine learning model. A neural network is a machine learning model that uses one or more layers of nonlinear units to predict an output for a received input. Some neural networks are deep neural networks, which include one or more hidden layers in addition to an output layer. The output layer of each hidden layer is used as input to another layer in the network—that is, another hidden layer, the output layer, or both. Some layers of the network generate an output from a received input according to a current value of a respective set of parameters, while other layers of the network may have no parameters.
[0020] The operations represented by the computational graph may be operations required for the neural network to compute an inference, i.e., to process an input through the layers of the neural network to produce a neural network output for the input. Additionally or alternatively, the operations represented by the computational graph may be operations required to train the neural network by performing a neural network training procedure to adjust the values of the neural network parameters, e.g., to determine trained values of parameters from initial values of the parameters using backpropagation. In some cases, e.g., during training of the neural network, the operations represented by the computational graph may include operations performed by multiple copies of the neural network.
[0021] To illustrate, a layer of a neural network that receives input from a previous layer may use a parameter matrix to perform matrix multiplication between the parameter matrix and the input. In some cases, matrix multiplication may be represented as multiple nodes in the computational graph. For example, matrix multiplication may be divided into multiple multiplication and addition operations, and each operation may be represented by a different node in the computational graph. The operation represented by each node may produce a respective output that flows along a directed edge to a subsequent node. After the operation represented by a terminal node produces a result of the matrix multiplication, the result flows along a directed edge to another node.The result is equivalent to an output of the neural network layer that performs the matrix multiplication.
[0022] In some other cases, matrix multiplication is represented as a node in the graph. The operations represented by the node can receive as inputs an input tensor on a first directed edge and a weight tensor, e.g., a parameter matrix, on a second directed edge. The node can process the input and weight tensors, e.g., perform matrix multiplication to output an output tensor on a third directed edge, which is equivalent to an output of the neural network layer.
[0023] Other neural network operations that can be represented by nodes in the computational graph include other mathematical operations, such as subtraction, division, and gradient calculations; ordering operations, such as concatenation, splicing, splitting, or sequencing; and neural network building block operations, such as SoftMax, Sigmoid, Rectified Linear Unit (ReLU), or convolutions.
[0024] In an example system, one or more sets of nodes in the computational graph may represent operations that control the flow of data through a computational graph. For example, the one or more sets of nodes may represent conditional, recursive, and / or iterative control flow statements, including: if statements, while loops, do-while loops, for loops, for-each loops, or nested control flow statements that include a combination of these statements.
[0025] The one or more sets of nodes in the computational graph may represent some operations that can translate into high-performance library operations, possibly including machine-specific high-performance implementations of linear algebraic operations, such as matrix multiplication operations, or neural network operations, such as back convolution. In some implementations, the suitability of operations in a graph for translation into high-performance library operations, and possibly fusion, as explained here, may depend on the underlying hardware architecture and machine configuration.
[0026] In one example compilation system, the compilation system merges multiple operations into a single fusion operation, which can be translated at code generation time into a single call that performs all of the merged operations. This fusion process produces code that is faster and potentially uses less memory for devices such as central processing units (CPUs) or graphics processing units (GPUs).
[0027] Fig. 1 illustrates an example compilation system 100. The compilation system 100 is an example of a system implemented as computer programs on one or more computers at one or more locations in which the systems, components, and techniques described below may be implemented.
[0028] The compilation system 100 receives an unoptimized computational graph as input 102. As described above, the computational graph represents operations as one or more sets of nodes and data dependencies between operations as edges.
[0029] A graph analyzer 106 of the compilation system 100 analyzes the input 102 of the non-optimized computational graph using a pattern matcher 104, which, for example, improves efficiency by comparing a specific pattern in the non-optimized computational graph. The compilation system compares patterns from the pattern matcher 104 with patterns of operations in the computational graph. The graph analyzer 106 then provides the analyzed graph to a graph fusion generator 108. For each corresponding pattern, the graph fusion generator 108 combines or fuses multiple operations from the non-optimized computational graph 102 according to the pattern into a single fusion operation to generate an optimized computational graph with fusion nodes. The graph fusion generator 108 then provides the optimized computational graph with fusion nodes to a code generator 114. The code generator 114 translates each fusion node as a call, e.g.,a loop or library call that performs all fused operations to generate efficient compiled code that can be delivered to multiple devices (116, 118, 120, 122) for execution. Since performing the fused operation in the optimized graph is more efficient than performing the corresponding multiple operations in the unoptimized graph that are replaced by the single fusion operation, the code generated by the optimized graph results in increased efficiency.
[0030] Any devices that perform the operations represented by the efficient compiled code, e.g., devices 116, 118, 120, 122, may include memory, e.g., random access memory (RAM), for storing instructions and data, and a processor for executing stored instructions. In general, each device is a hardware resource that executes the compiled code independently of other devices. For example, each device may have its own processing unit. The devices may be graphics processing units (GPUs), central processing units (CPUs), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), or other operation-specific processors. For illustrative purposes, a machine may house one or more devices, e.g., multiple CPUs, GPUs, FPGAs, ASICs, or operation-specific processors.
[0031] Fig. Figure 2 is a flow diagram of an example process 200 for generating efficient code from calculations within a computational graph. For convenience, the process 200 is described as being performed by a system of one or more computers located at one or more locations and suitably programmed in accordance with this patent specification. An example compilation system 100 of Fig. 1, suitably programmed, can, for example, perform process 200.
[0032] The system receives an unoptimized computational graph with nodes representing operations and directed edges representing data dependencies 202.
[0033] The system then analyzes the computational graph using pattern matching to determine fusible operations that can be merged together into a single fusion operation 204.
[0034] The system transforms the non-optimized computational graph into an optimized computational graph by replacing fusionable operations in the non-optimized computational graph with a node representing the single fusion operation 206.
[0035] The system then generates efficient compiled code by translating the fusion node of the optimized computational graph as a call that performs all fused operations 208. The efficient compiled code can then be delivered to computing devices such as graphical processing units and central processing units for execution. Loop fusion
[0036] Loop operations are implemented by iterating through elements of an input array, potentially multiple times, to compute an output array. Loop operations of a computational graph take the form of regular and non-regular operations. A regular operation, such as addition, exponent, or transpose, typically reads one element from each input array for each element of the output array. A non-regular operation, such as point or convolution, requires reading more than one element of the input array to produce a single element of the output array.
[0037] Regular operations can be decomposed into two types of functions: a function that is applied to the input data and a function that is applied to the output index.
[0038] For example, a regular unary operation can be expressed as: A[i0,..., i n-1 ] = f op (B[findex (i0, ..., i n-1 )]), where {i0, ..., i n-1 ) is a multidimensional index, f op is the data function, e.g., exponentiation, of the operation applied to the data item from the input, and f index is the index function of the operation that applies an index of the output array to an index of the input array. For each iteration of a loop, for example, the example compilation system implicitly constructs the index function f index the operation to determine which element of the input is to be read. The dimensionality of the output of f index may be different from the input, such as in a broadcast operation.
[0039] By expressing regular operations using separate index and data functions, the sample compilation system can easily merge operations using the composition of these functions.
[0040] Fig. 3 represents a graphical chain of two regular unary operations, B = op g (C) and A = op f (B) dar 305. As shown, a first regular one-digit operation may activate the output array B 303b by performing op g at the input device C 303c. A second regular one-digit operation can generate the output device A 303a by performing op f at the input arrangement B 303b. The operation op f has a data function f op and an index function f index on and the operation op g has a data function g op and an index function g index on.
[0041] The output arrangement A 303a can be expressed as a function of C 303c: A[i0,i1]=fop(gop(C[gindex(findex(i0,i1))])).
[0042] This expression merges op f and op g. The expression is in the form of a composition that can be used to merge a sequence of regular operations.
[0043] To merge operations, each operation should be decomposed into data functions and index functions. Some examples of regular operations that can be merged into a data function f op and an index function f index are: Element-wise unary exponentiation operations (1) Elementwise unary exponentiation operations fop(X)=exp(x) findex(i0,...,in−1)={i0,...,in−1} (2) Transposition fop(x)=x findex(i0,i1)={i1,i0} (3) Cutting (takes start and end indices) fop(x)=(x) findex(i0,...,in−1)={i0+start0,...,in−1+startn−1}
[0044] As an example, assume that the code in Table 1 is the code to be compiled. Table 1 C = ... B = op0(C) A = op1(B) D = op2(B)
[0045] The compilation system analyzes the computational graph representing the code using pattern matching. To merge unary operations, the compilation system scans the compilation graph and gathers as many unary operations as possible for the merge. If an operation selected for the merge lies outside the merge operation, the operation must be computed twice—once within the merge operation and once outside the merge operation.
[0046] For example, the compilation system can use known patterns of operations or algorithms to find chains of operations that can be merged together. In one example, patterns can refer to operations provided by a programming language or compiler. In another example, patterns can be provided with high-performance libraries, where the patterns refer to operations included in the high-performance library. In the example in Table 1, applying a merge of a regular unary operation results in: A=(op1∘op0)(C) D=(op2∘op0)(C)
[0047] For non-unary regular operations, e.g., A = op(B,C), the operations can also be expressed as data functions and index functions. For example, A = op(B,C) can be represented as:
[0048] A[i0,..., i n-1 ] = f op (B[findex0 (i0, ..., i n-1 )], C[f index1 (i0, ..., i n-1 )]), where f index0 and f index1 may be the same, e.g., in element-wise addition, where the index function is the identity, or may be different, e.g., as in a concatenation and binary operation with broadcast. The same rules of composition apply to non-unary regular operations. For example, the code in Table 2 may need to be compiled. Table 2 C = on g (D) A = on f (B, C)
[0049] If opg index and data functions g index and g op has and op f f index and f op , then the fusion operation can be expressed as: A[i0,...,in−1]=fop(B[findex0(i0,...,in−1)],gop(D[gindex1(findex1(i0,...,in−1))])).
[0050] Fig. Figure 4 shows operations of non-unary operations from Table 3 that can be merged. Merging non-unary operators is not the subject of the claims. Non-unary operations form a graph rather than a chain. Table 3 D = On P (B) E = ON Q (C) F = On R (D, E) G = On S (E) A = On T (F, G)
[0051] As disclosed above, to find operations that can be merged together, the compilation system analyzes the computational graph representing the code using pattern matching. To merge non-unary operations, the compilation system attempts to merge as many regular operations as possible, subject to a limit on the number of inputs to the merge operation. Too many inputs can increase memory usage and hinder performance.
[0052] For example, the compilation system can use known patterns of operations or algorithms to find operations that can be merged together. In the example in Table 3, applying a merge of regular non-unary operations results in: A[i0,...,in−1]=top(rop(pop(B[b_index])), qop(C[c_index_0])), sop(qop(C[c_index_1]))) where: b_index=pindex(rindex(tindex(i0,...,in−1))) c_index_0=qindex(rindex(tindex(i0,...,in−1))) c_index_1=qindex(sindex(tindex(i0,...,in−1)))
[0053] In this example, X index the index function for op x and X op is the data function for op x . The compilation system constructs this fused representation by traversing all paths in the graph from A 410 to the inputs B 401a and C 401b. Traversing upwards forms the index functions from A 410 to Op T 405e to OPR 405c and Ops 405d and then to Op p 405a and Op Q 405b and finally to inputs B 401a and C 401b. The downward traversal forms the data functions.
[0054] Regular operations can also be merged into some non-regular operations for better code efficiency. For example, consider the following implementation of column reduction in Table 4, which is a non-regular operation:
[0055] The column reduction algorithm divides an input matrix into tiles, each of which is reduced by a thread. Each thread accumulates the partial reduction results into the output vector. This reduction is not a regular operation because each output element is computed by multiple threads instead of one. However, the operations that generate the input elements can be merged into the column reduction if they are regular operations. For example, the input to the reducer operation in line 10 can be a subtraction between two elements, a left-side subtraction operation, Ihs[y][x], and a right-side subtraction operation, rhs[y][x].
[0056] Table 5 illustrates the merging of the operations that generate the input element into the column reduction. Line 10 shows the actual input element calculation. By merging the calculation with the column reduction, the compilation system generates code that does not require a separate kernel for subtraction or additional space to hold the subtraction result.
[0057] As disclosed above, to find input operations that can be merged together, the compilation system analyzes the computational graph representing the code using pattern matching. For example, the compilation system can merge element-wise operations that are a subset of regular operations into non-regular operations. Element-wise operations read the input element at the same index as the output element. Therefore, merging element-wise operations does not change the memory access pattern of the non-regular operation. Library merger
[0058] Some hardware vendors ship high-performance libraries with their hardware. These libraries may contain high-performance implementations of operations. However, the libraries are often non-free and / or written using hardware knowledge that is not public. A sample compilation system uses these vendor-provided libraries for specific patterns of calculations within a computational graph.
[0059] The example compilation system searches for operations in a computational graph that can be merged together by analyzing the computational graph representing the code using pattern matching. For example, the compilation system can improve efficiency by searching for specific patterns in subgraphs of operations that correspond to known library operations. These subgraphs can then be replaced in the computational graph with fusion nodes representing library operations.
[0060] Fig. Figure 5a illustrates a subgraph of a computational graph representing a transpose and a point operation. This subgraph computes an output array C by performing a point operation on an input array 501a and a transpose operation 505 that transposes an input array B 501b. The subgraph has the form: C = Point(A, Transpose(B)) as a pattern corresponding to a known library call. This pattern may correspond to a library call from a library provided by an external hardware vendor.
[0061] Fig. 5b represents an optimized version of the Fig. 5a. The compilation system may use a pattern 525 of a library call comprising input arrays 502a, 502b, a transpose operation 515, and a point operation to find patterns in the subgraph of Fig. 5a. After the compilation system has applied the pattern 525 of the library call to the subgraph of Fig. 5a, the compilation system can merge the subgraph into a single library call fusion operation. The compilation system can then replace subgraph 5a in the computational graph with a fusion node 530 representing the single fusion library call. During code generation, the compilation system translates the fusion node into the library call that performs all fused operations to produce the efficient compiled code.
[0062] Fig. Figure 6a illustrates a pattern representing a backward filter convolution. The pattern adapts a convolution operation 605 to activations A 601a and gradients G 601b, followed by a transpose operation 610. When a computational graph includes this pattern, the compilation system may merge the subgraph representing the backward filter convolution into a single fusion operation and replace the subgraph in the computational graph with a single fusion node representing the single backward filter convolution fusion operation.
[0063] Fig.Figure 6b illustrates a pattern representing a backward input convolution. The pattern adapts a convolution operation 640 of gradient G 602a and a mirrored filter F 602b generated by the inverse operation 630. When a computational graph includes this pattern, the compilation system may merge the subgraph representing the backward input convolution into a single fusion operation and replace the subgraph in the computational graph with a single fusion node representing the single backward input convolution fusion operation.
[0064] Once a compilation system replaces subgraphs within computational graphs with fusion nodes, the compilation system can translate the fusion nodes into calls that perform all fused operations. This process produces code that is more efficient than code compiled with one operation at a time.
[0065] Embodiments of the subject matter and functional operations described in this specification may be implemented in digital electronic circuitry, in tangible computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory storage medium for execution by or for controlling the operation of a data processing device.The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a direct or serial access memory device, or a combination of one or more of these. Alternatively or additionally, the program instructions may be encoded on a synthetically generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal generated to encode information for transmission to a suitable receiving device for execution by a data processing device.
[0066] The term "data processing device" refers to data processing hardware and includes all types of devices, apparatus, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. The device may also be or further comprise special-purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). The device may optionally comprise, in addition to hardware, code that creates an execution environment for computer programs, such as code that forms processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of these.
[0067] A computer program, which may also be referred to or described as a program, software or software application, app, module, a software module, script, or code, may be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but does not have to, correspond to a file in a file system. A program may be stored in a section of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in several coordinated files, e.g.,Files that store one or more modules, subroutines, or sections of code. A computer program can be deployed to run on one computer or on multiple computers located at one location or distributed across multiple locations and connected by a data communications network.
[0068] The processes and logic sequences described in this patent specification may be performed by one or more programmable computers executing one or more computer programs to perform functions by processing input data and generating output. The processes and logic sequences may also be performed by special-purpose logic circuitry, such as an FPGA or ASIC, or by a combination of special-purpose logic circuitry and one or more programmed computers.
[0069] Computers suitable for executing a computer program may be based on general-purpose or special-purpose microprocessors, or both, or some other type of central processing unit. Generally, a central processing unit receives instructions and data from read-only memory or random-access memory, or both. The essential elements of a computer are a central processing unit for carrying out instructions and one or more memory devices for storing instructions and data. The central processing unit and memory may be supplemented by or integrated with special-purpose logic circuitry. Generally, a computer also includes, or is operatively coupled to, one or more mass storage devices for storing data, for receiving data from, or transferring data to, such as magnetic, magneto-optical, or optical disks.However, a computer need not include such devices. Moreover, a computer may be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive, to name a few.
[0070] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0071] To provide for interaction with a user, embodiments of the subject matter described in this specification may be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, and a keyboard and pointing device, e.g., a mouse or trackball, through which the user can provide input to the computer. Other types of devices may also be used to provide for interaction with a user; feedback provided to the user may, for example, be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including auditory, speech, or tactile input.Additionally, a computer may interact with a user by sending documents to and receiving documents from a device used by the user; for example, by sending web pages to a web browser on a user's device in response to requests received from the web browser. A computer may also interact with a user by sending text messages or other forms of messaging to a personal device; for example, a smartphone running a messaging application and receiving reply messages from the user in return.
[0072] Embodiments of the subject matter described in this specification may be implemented in a computing system that includes a backend component, e.g., a data server, or that includes a middleware component, e.g., an application server, or that includes a frontend component, e.g., a client computer with a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such backend, middleware, or frontend components. The components of the system may be interconnected by any form or medium for digital data communication, e.g., a communications network. Examples of communications networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.
[0073] The computing system may include clients and servers. A client and a server are generally remote from each other and typically interact through a communications network. The client and server relationship is established by computer programs running on the respective computers, which have a client-server relationship with each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data for and receiving user input from a user interacting with the device acting as a client. Data generated at the user device, e.g., a result of the user interaction, may be received at the server from the device.
[0074] In addition to the embodiments of the appended claims and the embodiments described above, the following numbered embodiments are also innovative:
[0075] Embodiment 1 is a method comprising: obtaining a non-optimized computational graph having a plurality of nodes representing operations and directed edges representing dependencies; analyzing the non-optimized computational graph using pattern matching to determine fusible operations that can be fused together into a single fusion operation; transforming the non-optimized computational graph into an optimized computational graph by replacing the nodes representing the fusible operations in the non-optimized computational graph with a fusion node representing the single fusion operation; and providing the fusion node of the optimized computational graph to a compiler, which the compiler can translate as a call that performs the fused operations to produce efficient code.
[0076] Embodiment 2 is the method of Embodiment 1, further comprising: delivering the efficient code to computing devices for execution.
[0077] Embodiment 2 is the method of Embodiment 2, wherein the execution comprises: executing the operations of the computational graph with the single fusion call that performs all fused operations.
[0078] Embodiment 4 is the method of any of Embodiments 1 to 3, wherein analyzing the non-optimized computational graph using pattern matching to determine fusible operations that can be fused together into a single fusion operation comprises: comparing portions of the non-optimized computational graph with patterns of operations that each correspond to a single fusion operation; determining that a pattern corresponds to a portion of the non-optimized computational graph; and determining that the corresponding portion of the non-optimized computational graph can be replaced in the computational graph with the single fusion operation that corresponds to the corresponding pattern.
[0079] Embodiment 5 is the method of any of Embodiments 1 to 4, wherein the single fusion operation is an external code library operation.
[0080] Embodiment 6 is the method of any one of Embodiments 1 to 5, wherein the single fusion operation is a loop operation.
[0081] Embodiment 7 is the method of any of Embodiments 1 to 6, wherein analyzing the unoptimized computational graph using pattern matching to determine fusible operations that can be fused together into a single fusion operation comprises: searching the unoptimized computational graph for an input operation that requires computations to generate the input; and determining that the input operation in the computational graph can be replaced by a single fusion operation that corresponds to the computations required to generate the input.
[0082] Embodiment 8 is the method of any of Embodiments 1 to 7, wherein the fusible operations are regular operations.
[0083] Embodiment 9 is the method of any of Embodiments 1 to 8, wherein the fusible operations are regular operations that are fused into non-regular operations.
[0084] Embodiment 10 is the method of any of Embodiments 1 to 9, wherein analyzing the non-optimized computational graph using pattern matching to determine fusible operations that can be fused together into a single fusion operation comprises: finding a sequence of operations in a computational graph using a sequencing algorithm; and determining that the sequence of operations can be fused together into a single fusion operation using composition.
[0085] Embodiment 11 is a system of one or more computers and one or more storage devices storing instructions operable when executed by the one or more computers to cause the one or more computers to perform the operations of any of Embodiments 1-10.
[0086] Embodiment 12 is one or more non-transitory computer-readable storage media having instructions stored thereon that are executable by a processing device and, when so executed, cause the processing device to perform the operations of any of Embodiments 1-10.
[0087] Although this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features described in this specification in connection with separate embodiments may be implemented in combination in a single embodiment. Conversely, various features described in connection with a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination.Moreover, although features may be described above as operating in certain combinations and may even be initially claimed as such, in some cases one or more features from a claimed combination may be carved out of the combination and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0088] Likewise, although operations are illustrated in a particular order in the drawings, this should not be understood to require that such operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed to achieve desired results. Under certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood to require such separation in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or may be packaged into multiple software products.
[0089] Specific embodiments of the subject matter have been described. In some cases, multitasking and parallel processing may be advantageous.
Claims
[1] Computer-implemented method in a compilation system, comprising: Obtaining (202) a non-optimized computational graph having a plurality of nodes representing operations and directed edges representing data dependencies; Analyzing (204) the non-optimized computational graph using a plurality of known patterns and pattern matching to determine fusionable operations, where a pattern specifies a concatenation of several unary regular operations, where in concatenation the result of a regular operation is an input of a concatenated subsequent operation, where a regular operation is an array operation that reads an element from each input array to calculate one element of the output array using a data function, when recognizing the pattern of a concatenation of unary regular operations: Decomposing each of the unary regular operations into an index function and a data function, where the index operation maps an index of the output array into an index of the input array; Concatenating the index functions and concatenating the data functions of the respective fusionable operations into a single fusion operation; Transforming (206) the non-optimized computational graph into an optimized computational graph by replacing the nodes representing the fusionable operations in the non-optimized computational graph with a single fusion node representing the single fusion operation; and Providing the single fusion node of the optimized computational graph to a compiler, wherein the compiler translates the single fusion node as a call that performs the single fusion operation to generate efficient code in the code generation phase of compilation (208). [2] The method of claim 1, further comprising: Delivering efficient code to computing devices for execution. [3] The method of claim 2, wherein the execution comprises: Execute the operations of the computational graph with the call that performs the single fusion operation. [4] The method of any one of claims 1 to 3, wherein analyzing the non-optimized computational graph using pattern matching to determine fusible operations that can be fused together into a single fusion operation comprises: Comparing sections of the unoptimized computational graph with patterns of operations, each corresponding to a single fusion operation; Determining that a pattern corresponds to a portion of the unoptimized computational graph; and Determine that the corresponding section of the unoptimized computational graph can be replaced in the computational graph by the single fusion operation that matches the corresponding pattern. [5] The method of any one of claims 1 to 4, wherein the single fusion operation is an external code library operation. [6] A method according to any one of claims 1 to 4, wherein the single fusion operation is a loop operation. [7] The method of any one of claims 1 to 6, wherein analyzing the non-optimized computational graph using pattern matching to determine fusible operations that can be fused together into a single fusion operation comprises: Searching the unoptimized computational graph for a node representing a first operation that takes as input an output generated by a chain of operations represented by a chain of nodes in the unoptimized computational graph and requires computations to generate the input; and Determining that the chain of operations in the computational graph can be replaced in the optimized computational graph by a single fusion operation that corresponds to the chain of operations required to generate the input for the first operation. [8] The method of any one of claims 1 to 7, wherein the fusionable operations are regular operations. [9] A method according to any one of claims 1 to 8, wherein the fusionable operations are regular operations that are fused into non-regular operations. [10] The method of any one of claims 1 to 9, wherein analyzing the non-optimized computational graph using pattern matching to determine fusible operations that can be fused together into a single fusion operation comprises: Finding a sequence of operations in a computational graph using a sequencing algorithm; and Determining that the operation sequence can be merged into a single fusion operation using composition. [11] Compilation system that includes: one or more computers; and one or more storage devices that store instructions operable when executed by one or more computers to cause the one or more computers to perform operations that include: Obtaining (202) a non-optimized computational graph having a plurality of nodes representing operations and directed edges representing data dependencies; Analyzing (204) the non-optimized computational graph using a plurality of known patterns and pattern matching to determine fusionable operations, where a pattern specifies a concatenation of several unary regular operations, where in concatenation the result of a regular operation is an input of a concatenated subsequent operation, where a regular operation is an array operation that reads an element from each input array to calculate one element of the output array using a data function, when recognizing the pattern of a concatenation of unary regular operations: Decomposing each of the unary regular operations into an index function and a data function, where the index operation maps an index of the output array into an index of the input array; Concatenating the index functions and concatenating the data functions of the respective fusionable operations into a single fusion operation; Transforming (206) the non-optimized computational graph into an optimized computational graph by replacing the nodes representing the fusionable operations in the non-optimized computational graph with a single fusion node representing the single fusion operation; and Providing the single fusion node of the optimized computational graph to a compiler, wherein the compiler translates the single fusion node as a call that performs the single fusion operation to generate efficient code in the code generation phase of compilation (208). [12] The compilation system of claim 11, wherein the operations further comprise: Delivering efficient code to computing devices for execution. [13] A compilation system according to claim 12, wherein the execution comprises: Execute the operations of the computational graph with the call that performs the single fusion operation. [14] A compilation system according to any one of claims 11 to 13, wherein analyzing the non-optimized computational graph using pattern matching to determine fusible operations that can be fused together into a single fusion operation comprises: Comparing sections of the unoptimized computational graph with patterns of operations, each corresponding to a single fusion operation; Determining that a pattern corresponds to a portion of the unoptimized computational graph; and Determine that the corresponding section of the unoptimized computational graph can be replaced in the computational graph by the single fusion operation that matches the corresponding pattern. [15] A compilation system according to any one of claims 11 to 14, wherein the single fusion operation is an external code library operation. [16] A compilation system according to any one of claims 11 to 14, wherein the single fusion operation is a loop operation. [17] A compilation system according to any one of claims 11 to 16, wherein analyzing the unoptimized computational graph using pattern matching to determine fusible operations that can be fused together into a single fusion operation comprises: Searching the unoptimized computational graph for a node representing a first operation that takes as input an output generated by a chain of operations represented by a chain of nodes in the unoptimized computational graph and requires computations to generate the input; and Determining that the chain of operations in the computational graph can be replaced in the optimized computational graph by a single fusion operation that corresponds to the chain of operations required to generate the input for the first operation. [18] A compilation system according to any one of claims 11 to 17, wherein the fusionable operations are regular operations. [19] A compilation system according to any one of claims 11 to 18, wherein analyzing the unoptimized computational graph using pattern matching to determine fusible operations that can be fused together into a single fusion operation comprises: Finding a sequence of operations in a computational graph using a sequencing algorithm; and Determining that the operation sequence can be fused together into a single fusion operation using a composition. [20] One or more non-transitory computer-readable storage media having stored thereon instructions executable by a processing device and, when so executed, causing the processing device to perform operations in a compilation system comprising: Obtaining (202) a non-optimized computational graph having a plurality of nodes representing operations and directed edges representing data dependencies; Analyzing (204) the non-optimized computational graph using a plurality of known patterns and pattern matching to determine fusionable operations, where a pattern specifies a concatenation of several unary regular operations, where in concatenation the result of a regular operation is an input of a concatenated subsequent operation, where a regular operation is an array operation that reads an element from each input array to calculate one element of the output array using a data function, when recognizing the pattern of a concatenation of unary regular operations: Decomposing each of the unary regular operations into an index function and a data function, where the index operation maps an index of the output array into an index of the input array; Concatenating the index functions and concatenating the data functions of the respective fusionable operations into a single fusion operation; Transforming (206) the non-optimized computational graph into an optimized computational graph by replacing the nodes representing the fusionable operations in the non-optimized computational graph with a single fusion node representing the single fusion operation; and Providing the single fusion node of the optimized computational graph to a compiler, wherein the compiler translates the single fusion node as a call that performs the single fusion operation to generate efficient code in the code generation phase of compilation (208).
Citation Information
Patent Citations
Automated processor generation system for designing a configurable processor and method for the same
US20030208723A1
System and method for fusing instructions
US20040139429A1