Method and apparatus for converting a non-string parallel control flow graph into a data flow
Through a non-serial parallel converter detecting and generating PICK instructions that consume operands, the control flow diagram is converted into data flow code, which solves the problem that complex control flow programs cannot be converted efficiently in parallel computing systems, and realizes efficient parallel execution of the data flow architecture.
Patent Information
- Application Number
- CN201811344350.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-12-20
- Filing Date
- 2018-11-13
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2038-11-13
AI Technical Summary
The prior art is difficult to efficiently convert complex control flow programs into data flow programs, resulting in the inability to fully utilize the parallelism of the data flow architecture in parallel computing systems.
The non-serial parallel node in the control flow graph is detected by a non-serial parallel converter, and a PICK instruction including the consumed operand is generated, which converts it into data flow code, and uses the node analyzer and instruction generator to realize the conversion of the non-serial parallel node.
It realizes the conversion of complex control flow programs into data flow programs, which can be efficiently executed in parallel computing systems, and improves the parallel processing capabilities of the computing system.
Smart Images

Figure CN109947427B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to compilers in computing systems, and more particularly, to methods and devices for improving compiler efficiency in computing systems by converting a non-serial-parallel control flow graph into a data flow. Background Art
[0002] Many computing systems operate according to a control flow architecture. In a control flow architecture, the execution of instructions of a program is driven by a program counter that stepwise traverses the instructions of the program. In other words, the execution order of the instructions of the program is defined by the structure of the program itself. In some cases, when attempting to implement parallel processing, the control flow architecture may operate inappropriately. For example, a program may claim to execute an instruction even though the inputs (e.g., operands) of the instruction have not yet been updated by a parallel operation instruction.
[0003] Some computing systems utilize a data flow architecture. A data flow architecture is not driven by the instruction execution order defined by a program. Instead, a data flow architecture executes instructions based on the availability of the inputs (e.g., operands) of the instructions. For example, if an instruction has three operands, a computing system utilizing a data flow architecture will execute the instruction once the three operands are provided to the instruction by other instruction(s) on which the instruction depends. Thus, a data flow architecture can execute in a highly parallel environment without worrying about an instruction being executed before the data dependencies of the instruction are updated / satisfied. For example, a data flow architecture can be used in a large-scale computing system that uses a large number of processing elements to highly parallelize processing. Brief Description of the Drawings
[0004] Figure 1 is a block diagram of an example system for converting control flow code into data flow code.
[0005] Figure 2 is Figure 1 a block diagram of an example implementation of a non-serial-parallel converter of
[0006] Figure 3 illustrates an example control flow graph of serial-parallel.
[0007] Figure 4 illustrates an example control flow graph including nodes of non-serial-parallel.
[0008] Figure 5 is a flowchart representing example machine-readable instructions for implementing Figure 3 and / or Figure 4 of a non-serial-parallel converter.
[0009] Figure 6 is Figure 5 instructions that can execute Figure 3and / or Figure 4 Block diagram of an example processing device of a non-serial-parallel converter.
[0010] The accompanying drawings are not drawn to scale. As used in this patent, reciting that any component (e.g., layer, film, region, or plate) is positioned in any manner on (e.g., positioned on, located on, disposed on, or formed on, etc.) another component indicates that the referenced component is in contact with the other component, or that the referenced component is above the other component and one or more intermediate components are between the referenced component and the other component. Reciting that any component is in contact with another component means that there are no intermediate components between the two components. Detailed Description
[0011] Series-parallel graphs are widely used in application engineering and algorithm research theory. If a graph is a series-parallel graph, some graph problems that are generally NP-Complete problems can be solved in linear time. Highly efficient algorithms have been developed to generate highly parallel data flow codes for control flow graphs that are series-parallel graphs.
[0012] A directed graph G is series-parallel at both ends and has endpoints s and t if the directed graph G can be generated by the following sequence of operations:
[0013] 1. Create a new graph consisting of a single edge pointing from s to t.
[0014] 2. Given two series-parallel graphs X and Y at both ends, having endpoints sX, tX, sY, and tY, form a new graph G = P(X, Y) by identifying s = sX = sY and t = tX = tY. This is called the parallel combination of X and Y.
[0015] 3. Given two series-parallel graphs X and Y at both ends, having endpoints sX, tX, sY, and tY, form a new graph G = S(X, Y) by identifying s = sX, tX = sY, and t = tY. This is called the serial combination of X and Y.
[0016] An undirected graph is series-parallel at both ends with endpoints s and t if, for some orientation of the edges of the undirected graph, the undirected graph forms a directed series-parallel graph at both ends with endpoints s and t. A directed graph or an undirected graph is series-parallel if it is series-parallel at both ends with respect to two vertices s and t.
[0017] Although data flow architectures can be advantageously used in computing systems, not all programs are designed to and / or capable of being directly written in data flow code. For example, existing programs may be implemented for control flow architectures. The methods and devices disclosed herein facilitate the conversion of control flow programs to data flow programs. Specifically, the methods and devices disclosed herein identify non-series-parallel control flows in programs (e.g., Figure 4Example non-serial-parallel flow) and convert the non-serial-parallel control flow into data flow code. Thus, in some disclosed examples, traditional programs with complex control flows can run on highly parallel data flow architectures.
[0018] Figure 1 FIG. 4 is a block diagram of an example system 100 that includes example input code 102 compiled by an example compiler 103 to generate example data flow assembly code 112. For example, system 100 can be implemented within a compiler to convert control flow code into data flow code.
[0019] The input code 102 of the illustrated example is a serialized instruction software program that is transformed and compiled into assembly code (e.g., for execution on a computing platform). The example input code 102 is written in the C programming language. Alternatively, the input code can be written in any other language, such as C++, Fortran, Java, etc.
[0020] The example compiler 103 includes an example intermediate representation (IR) transformer 304, an example serialized instruction transformer 106, an example control flow to data flow converter 108, and an example non-serial-parallel converter 110. The illustrated example compiler 103 is an LLVM compiler that has been modified to include the non-serial-parallel converter 110. Alternatively, any other type of pre-existing or newly created compiler can be used.
[0021] The illustrated example IR transformer 104 transforms the example input code 102 into intermediate representation code. The example IR transformer 104 transforms the input code 102 into LLVM intermediate representation. Alternatively, the IR transformer 104 can transform the input code 102 into any type of intermediate representation recognized by the compiler used in system 100 (e.g., standard portable intermediate representation, Java bytecode, common intermediate representation, etc.).
[0022] The example serialized instruction transformer 106 transforms the intermediate representation from the IR transformer 104 into serialized control flow machine instructions for the target machine. According to the illustrated example, the serialized instruction transformer 104 transforms the intermediate representation code one part at a time to facilitate verification and conversion to data flow instructions.
[0023] The example control flow to data flow converter 108 converts control flow serialized instructions into data flow assembly code 112. The example control flow to data flow converter 108 uses well-known techniques to convert serial-parallel control flow code into data flow code.
[0024] The example control flow to data flow converter 108 utilizes two basic data flow instructions, SWITCH and PICK.
[0025] The PICK instruction selects a value from two registers based on a predicate. PICK has the following format: %r = PICK %r0, %r1, %r2, where %r is the symbol of a virtual register in LLVM. Such an instruction can be written in the C programming language as %r = (%r0)? (%r2) : (%r1);
[0026] The SWITCH instruction sends an operand to one of two identified registers based on a predicate. SWITCH has the following format: %r1, %r2 = SWITCH %r0, %r3, where this format sends the value in %r3 to %r1 or %r2 based on the predicate in %r0. In C, this format is similar to: if (%r0) %r2 = %r3 else %r1 = %r3.
[0027] For example, Figure 3 FIG. illustrates an example control flow graph 300 of serial-parallel. In the illustrated example, the first node 302 assigns the value X to X11 or X12 based on the value of C1. (For example, if C1 is false, then X11 = X, and if C1 is true, then X12 = X.) Similarly, when C2 is false, the second node 304 assigns the value X11 to X21, and when C2 is true, the second node 304 assigns the value X11 to X22. When C3 is false, the third node 306 assigns the value X12 to X31, and when C3 is true, the third node 306 assigns the value X12 to X32. Finally, the fourth node 308 is a Phi operation on X21, X22, X31, and X32. The Phi function is a pseudo function in static single assignment form. The Phi function will assign the value (for example, one of X21, X22, X31, or X32) that the control reaches the fourth node 308 to X4.
[0028] According to the example control flow graph 300, the second node 304 is serial with the first node 302 and parallel with the third node 306. Therefore, since the control flow graph 300 is serial-parallel, the control flow graph 300 can be easily converted to a data flow using well-known techniques.
[0029] For example, when converting the control flow code represented by Figure 3 to data flow code, the example control flow to data flow converter will generate the following PICK tree: X2 = PICK C2, X21, X22; X3 = PICK C3, X31, X32; and X4 = PICK C1, X2, X3.
[0030] In fact, many control flow graphs are serial-parallel or can be easily reduced to serial-parallel using well-known techniques. However, some control flow graphs are not easily convertible. For example, goto ("go to"), indirect jumps, or optimizations such as inlining function calls and loop unrolling may create very complex control flow graphs that cannot be reduced to serial-parallel graphs. For example, Figure 4 FIG. Figure 4 illustrates an example control flow graph 400 generated using the goto command. The example control flow graph 400 includes a first node 402, a second node 404, a third node 406, a fourth node 408, a fifth node 410, and a sixth node 412. Figure 4 The control flow graph 400 of FIG. Figure 4 is not serial-parallel because the example first node 402 and the example second node 404 are not added to the graph using a serial or parallel combination.
[0031] The example control flow-to-data flow converter 108 also includes an example non-serial-parallel converter 110 for detecting non-serial-parallel code within the control flow serialization code from the serialization instruction converter 106 and converting the non-serial-parallel code into data flow code. For example, when applied to Figure 4 the instructions illustrated in FIG. Figure 4 , the non-serial-parallel converter 110 detects that the instructions include non-serial-parallel code that cannot be converted into parallel optimized data flow code according to existing methods. For example, when calculating the out-going value of X from the example fourth node 408, the evaluation of only the assertion value C2 is not sufficient. The evaluation of C3 is also required. Additionally, the order of evaluation is important because if the branch to C2 is taken, C3 will have no value in the data flow machine. A complex logical combination of C2 and C3 or the implementation of the evaluation of both branches is required to calculate X4 in the fourth node 408, and the same is true for X5 in the fifth node 410, both of which waste power.
[0032] To facilitate the conversion of non-serial-parallel code into data flow code that can be used in, for example, a highly parallel environment, the example non-serial-parallel converter 110 utilizes a consume operand, any value of which is consumed by the system. An example consume operand is %ign, which can be referred to as an ignore operand. The value assigned to the consume operand is not stored. According to the illustrated example, the non-serial-parallel converter 110 includes a consume operand in the PICK instruction to account for branches in the node tree that may be irrelevant to determining a particular value.
[0033] For example, the non-serial-parallel converter 110 may generate the following PICK tree for determining the value of X4 in the example fourth node 408: X2 = PICK C2, X21 %ign; X3 = PICK C3, X31, %ign; and X4 = PICK C1, X2, X3.
[0034] Figure 2 is Figure 1 a block diagram of an example implementation of the non-serial-parallel converter 100. Figure 2 The example non-serial-parallel converter 110 includes an example non-serial-parallel detector 202, an example node analyzer 204, and an example instruction generator 206.
[0035] The example non-serial-parallel detector 202 analyzes the control-flow serialized code processed by the example control-flow-to-data-flow converter 108 to identify non-serial-parallel code. The example non-serial detector 202 uses a conceptual control-dependency graph (CDG) to detect non-serial-parallel graphs. In a traditional control-flow graph, if node B is on a path from node A to an exit (sink) node of the control-flow graph and there is another path from node A to the exit that does not pass through node B, then node B is control-dependent on node A. In the corresponding control-dependency graph, there is an edge from A to B. Node A is called the parent node of node B. The example non-serial-parallel detector 202 constructs a control-dependency graph for the control-flow graph of the input program. When the CDG has the property that each node has one and only one control-dependency parent node, the graph is serial-parallel. When at least one node does not have one and only one control-dependency parent node (e.g., more than one parent node, zero parent nodes, etc.), the example non-serial-parallel detector 202 identifies a non-serial-parallel graph.
[0036] For example, the non-serial-parallel detector 202 can apply a set of rules to identify non-serial-parallel code. In some examples, the non-serial-parallel detector 202 can determine that the code is not serial-parallel (e.g., does not meet the rules for serial-parallel code). In other examples, the non-serial-parallel detector 202 can determine that the code is non-serial-parallel (e.g., meets the rules for non-serial-parallel code). For example, when the code includes a node that is neither serial nor parallel to other nodes, the non-serial-parallel detector 202 can determine that the code is non-serial-parallel. For example, code that does not meet the following rules can be classified as non-serial-parallel:
[0037] Given two two-terminal serial-parallel graphs X and Y with endpoints sX, tX, sY, and tY, a new graph G = P(X, Y) is formed by identifying s = sX = sY and t = tX = tY. This is called the parallel composition of X and Y.
[0038] Given two two-terminal serial-parallel graphs X and Y with endpoints sX, tX, sY, and tY, a new graph G = S(X, Y) is formed by identifying s = sX, tX = sY, and t = tY. This is called the serial composition of X and Y.
[0039] When the example non-serial-parallel detector 202 identifies non-serial-parallel code, the non-serial-parallel detector 202 triggers the analysis performed by the example node analyzer 204.
[0040] The example node analyzer 204 of the illustrated example analyzes the code to identify non-serial-and-parallel nodes. For example, the node analyzer 204 can identify nodes that are not serial and / or parallel. The example node analyzer 204 triggers the example instruction generator 206 to determine PICK instructions for the non-serial-and-parallel nodes.
[0041] The example instruction generator 206 generates PICK instructions for converting non-serial-and-parallel control flow serialization code into data flow code. The example instruction generator 206 generates data flow code by analyzing the non-serial-and-parallel nodes and prior nodes to generate PICK instructions that include consumed operands (e.g., ignored operands such as %ign). Although the illustrated example uses PICK instructions, any instruction of a compiler that selects a value from an operand based on an assertion can be utilized.
[0042] When analyzing a node and the prior nodes of that node, the example instruction generator 206 generates PICK instructions for each prior node, sets the first operand to a value that can propagate to the non-serial-and-parallel node, and sets the second operand to a consumed operand (e.g., because the second possible path from the prior node does not propagate to the non-serial-and-parallel node). For example, when generating a PICK instruction for Figure 4 the fourth node 408 of Figure 4 for the second node 404 of Figure 4 the instruction generator 206 generates the PICK instruction X2 = PICK C2, X21, %ign. In this instance, the second output from the second node 404 does not propagate to the fourth node 408 and is therefore set to a consumed operand to ensure that the PICK instruction is executed and does not propagate a value to X4 when C2 is true, thus remaining consistent with the desired operation programmed in the input code. Thus, the instruction generator 206 can generate the following PICK tree for the fourth node 408: X2 = PICK C2, X21 %ign; X3 = PICK C3, X31, %ign; and X4 = PICK C1, X2, X3.
[0043] In some examples, in cases where the processing architecture on which the code is executed does not propagate consumed operands, the instruction generator 206 can utilize a white value such as zero (zero), and if the execution of a code branch moves away from a node, a SWITCH instruction can be included to consume the zero. For example, for the fourth node 408, the instruction generator can generate the following data flow code:
[0044] X2 = PICK C2, X21, 0
[0045] X3 = PICK C3, X31, 0
[0046] X’4 = PICK C1,X2,X3
[0047] E2 =!C1 && C2 / / Take branch C2 -> B5
[0048] E3 = C1 && C3 / / Take branch C2 -> B5
[0049] E = E1 || E2
[0050] X4, %ign = SWITCH E,X’4
[0051] In this example, && is the short - circuit logical AND and || is the short - circuit logical OR. Since the second operand is only evaluated if the first operand does not determine the output of the operator, the short - circuit nature of those operators saves power. Additionally, white values (e.g., zero) never change when propagated through the PICK tree, which also consumes less power during execution.
[0052] Although Figure 1 illustrates an example way of implementing the non - serial - parallel converter 110, Figure 1 one or more of the elements, processes, and / or devices illustrated in Figure 2 can be combined, split, rearranged, omitted, eliminated, and / or implemented in any other way. Additionally, the example non - serial - parallel detector 202, the example node analyzer 204, the example instruction generator 206, and / or more generally, the example non - serial - parallel converter 110 can be implemented by hardware, software, firmware, and / or any combination of hardware, software, and / or firmware. Thus, for example, any one of the example non - serial - parallel detector 202, the example node analyzer 204, the example instruction generator 206, and / or more generally, the example non - serial - parallel converter 110 can be implemented by one or more analog or digital circuits, logic circuits, (one or more) programmable processors, (one or more) application - specific integrated circuits (ASICs), (one or more) programmable logic devices (PLDs), and / or (one or more) field - programmable logic devices (FPLDs). When reading any of the apparatus claims or system claims of this patent to cover a pure software and / or firmware implementation, at least one of the example non - serial - parallel detector 202, the example node analyzer 204, the example instruction generator 206, and / or more generally, the example non - serial - parallel converter 110 is hereby expressly defined to include a non - transitory computer - readable storage device or storage disk, such as a memory, a digital versatile disc (DVD), a compact disc (CD), a Blu - ray disc, etc., that includes the software and / or firmware. Additionally, in addition to Figure 2In place of those illustrated, example non-serial-to-parallel converter 110 may include one or more elements, processes, and / or devices, and / or may include more than one of any or all of the elements, processes, and devices illustrated.
[0053] Figure 5 illustrates a flow diagram representing example machine-readable instructions for implementing non-serial-to-parallel converter 110. In the example, the machine-readable instructions include a program for execution by a processor such as Figure 6 application processor 601 and / or a processor such as processor 612 shown in the example processor platform 600 discussed below in connection with Figure 6 Although the program can be embodied as software stored on a non-transitory computer-readable storage medium such as a CD-ROM, floppy disk, hard drive, digital versatile disc (DVD), Blu-ray disc, or memory associated with processor 612, all or portions of the program can alternatively be executed by a device other than processor 612 and / or can be embodied as firmware or dedicated hardware. Additionally, although the example program is described with reference to Figure 5 the flow diagram illustrated, many other methods of implementing example non-serial-to-parallel converter 110 can alternatively be used. For example, the order of execution of the blocks can be changed, and / or some of the blocks described can be changed, eliminated, or combined. Additionally or alternatively, any or all of the blocks can be implemented by one or more hardware circuits (e.g., discrete and / or integrated analog and / or digital circuits, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), comparators, operational amplifiers (op-amps), logic circuits, etc.) structured to perform the corresponding operations without executing software or firmware.
[0054] As mentioned above, the Figure 5Example processes, non-transitory computer and / or machine-readable media such as: hard disk drives, flash memories, read-only memories, compact discs, digital versatile discs, caches, random access memories, and / or any other storage device or storage disk that stores information therein for any duration (e.g., over an extended period of time, permanently, during a brief instance, during temporary buffering, and / or during caching of information). As used herein, the term "non-transitory computer-readable storage medium" is expressly defined to include any type of computer-readable storage device and / or storage disk, and to exclude propagated signals and to exclude transmission media. "Comprising" and "including" (and all of their forms and tenses) are used herein as open-ended terms. Thus, whenever a claim lists any content following any form of "comprising" or "including" (e.g., includes, comprises, etc.), it is to be understood that additional elements, items, etc. may exist without falling outside the scope of the corresponding claim. As used herein, when the phrase "at least" is used as a transitional term in the preamble of a claim, it is open-ended in the same manner as the terms "comprising" and "including".
[0055] When the non-serial-parallel detector 202 detects a non-serial-parallel flow graph being processed by the example control-to-data flow converter 108, Figure 5 the program begins (block 502). For example, the non-serial-parallel detector 202 may apply one or more rules to the control flow serialization code to test for non-serial-parallel nodes. In response to detecting a non-serial-parallel flow graph, the example node analyzer 204 detects non-serial-parallel nodes (block 504). For example, when processing Figure 4 the illustrated code, the example node analyzer 204 may detect that the example fourth node 408 is a non-serial-parallel node.
[0056] In response to detecting a non-serial-parallel node, the example node analyzer 204 selects the node in the serialization code that precedes the non-serial-parallel node (block 506). For example, when Figure 4 processing the fourth node 408 as a non-serial-parallel node, the node analyzer 204 selects the example second node 404 as the preceding node. The example instruction generator 206 generates a PICK instruction with an ignore operand for the selected preceding node (block 508). For example, when Figure 4 processing the second node 404 as the preceding node of the fourth node 408, the instruction generator 206 generates X2 = PICK C2,X21, %ign. Although the example instruction generator 206 generates a PICK instruction, any similar instruction for a particular compiler environment that takes a consuming operand may be utilized. For example, any instruction that ensures an assertion is evaluated with a consuming operand (e.g., an ignore operand, a white value, etc.) may be utilized.
[0057] The example node analyzer 204 determines whether there is another previous node (block 510). When there is another previous node, control returns to block 506 to generate a PICK instruction for that previous node. For example, after processing the second node 404, the third node 406 can be processed.
[0058] When there is no other previous node (block 510), the example instruction generator 206 generates instructions to combine the results of the PICK instructions from the previous nodes. According to the illustrated example, the instruction generator 206 generates a PICK instruction to combine the results to the previous nodes. For example, when generating an instruction for the fourth node 408, after generating the PICK instructions for the second node 404 and the third node 406, the example instruction generator 206 generates: X4 = PICK C1,X2,X3.
[0059] The example node analyzer 204 determines whether there is another non-serial-parallel node in the code being analyzed (block 514). When there is another non-serial-parallel node, control returns to block 502 to process the next node in order to generate the data flow assembly code 112. When there is no other non-serial-parallel node, the example instruction generator 206 outputs the generated data flow code (block 516), and Figure 5 the program ends.
[0060] Although the foregoing examples generate instructions for an environment in which consumed operands are propagated, in some environments, such consumed operands are not propagated. In an environment that does not support the propagation of consumed operands, the consumed operands can be replaced with white values such as zero (e.g., X2 = PICK C2,X21,0 and X3 = PICK C3,X31,0). Additionally, the instructions for combining elements (block 512) can include several instructions for selecting appropriate values. For example, the combination instruction can store a value in a temporary register (X’4 = PICK C1,X2,X3), and a switch can be used to select the final value to ensure that the white value is consumed (E2 =!C1&&C2; E3 = C1&&C3; E = E1||E2; and X4,%ign = SWITCH E,X’4).
[0061] Figure 6 is capable of executing Figure 5 instructions to implement Figure 1 is a block diagram of an example processor platform 600 of a non-serial-parallel converter 110 that can execute TM instructions. The processor platform 600 can be, for example, a server, a personal computer, a mobile device (e.g., a cellular phone, a smart phone, a tablet device such as an iPad
[0062] The processor platform 600 of the illustrated example includes a processor 612. The processor 612 of the illustrated example is hardware. For example, the processor 612 can be implemented by one or more integrated circuits, logic circuits, microprocessors, or controllers from any desired family or manufacturer. The hardware processor can be a semiconductor-based (e.g., silicon-based) device. In this example, the processor 612 implements the non-serial parallel detector 202, the node analyzer 204, and the instruction generator 206.
[0063] The processor 612 of the illustrated example includes local memory 613 (e.g., a cache). The processor 612 of the illustrated example communicates with main memory including volatile memory 614 and non-volatile memory 616 via a bus 618. The volatile memory 614 can be implemented by synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), RAMBUS dynamic random access memory (RDRAM), and / or any other type of random access memory device. The non-volatile memory 616 can be implemented by flash memory and / or any other desired type of memory device. Access to the main memory 614, 616 is controlled by a memory controller.
[0064] The processor platform 600 of the illustrated example also includes interface circuitry 620. The interface circuitry 620 can be implemented by any type of interface standard, such as an Ethernet interface, a universal serial bus (USB), and / or a PCI Express interface.
[0065] In the illustrated example, one or more input devices 622 are connected to the interface circuitry 620. The input device(s) 622 permit a user to input data and / or commands into the processor 612. The input device(s) can be implemented by, for example, an audio sensor, a microphone, a camera (still or video), a keyboard, a button, a mouse, a touch screen, a trackpad, a trackball, an isopoint mouse, and / or a voice recognition system.
[0066] One or more output devices 624 are also connected to the interface circuitry 620 of the illustrated example. The output device 624 can be implemented by, for example, a display device (e.g., a light-emitting diode (LED), an organic light-emitting diode (OLED), a liquid crystal display, a cathode ray tube display (CRT), a touch screen, a haptic output device, a printer, and / or a speaker). Thus, the interface circuitry 620 of the illustrated example typically includes a graphics driver card, a graphics driver chip, and / or a graphics driver processor.
[0067] The interface circuit 620 of the illustrated example also includes communication devices such as a transmitter, a receiver, a transceiver, a modem, and / or a network interface card to facilitate the exchange of data with an external machine (e.g., any kind of computing device) via a network 626 (e.g., an Ethernet connection, a digital subscriber line (DSL), a telephone line, a coaxial cable, a cellular phone system, etc.).
[0068] The processor platform 600 of the illustrated example also includes one or more mass storage devices 628 for storing software and / or data. Examples of such mass storage devices 628 include floppy disk drives, hard disk drives, compact disk drives, Blu-ray disk drives, RAID systems, and digital versatile disk (DVD) drives.
[0069] Figure 5 The encoded instructions 632 can be stored in the mass storage device 628, stored in the volatile memory 614, stored in the non-volatile memory 616, and / or stored on a removable tangible computer-readable storage medium such as a CD or a DVD.
[0070] Example methods, devices, systems, and articles for converting a non-serial-parallel control flow graph to a data flow are disclosed herein. Further examples and combinations thereof include the following.
[0071] Example 1 includes a device for converting control flow code to data flow code, the device including: a node analyzer for detecting non-serial-parallel nodes in serialized code; and an instruction generator for: generating instructions for a prior node that include consuming operands; generating a combining instruction for combining the results of the instructions; and outputting the combining instruction and the instructions to generate data flow code.
[0072] Example 2 includes the device as defined in Example 1, wherein the consuming operand is an ignore operand.
[0073] Example 3 includes the device as defined in Example 1 or Example 2, further including a non-serial-parallel detector for detecting that the control flow graph of the serialized code is a non-serial-parallel flow graph.
[0074] Example 4 includes the device as defined in Example 1 or Example 2, wherein the instruction is a PICK instruction.
[0075] Example 5 includes the device as defined in Example 1, wherein the consuming operand is a white value.
[0076] Example 6 includes the device as defined in Example 5, wherein the instruction generator is further for generating an instruction for consuming the white value.
[0077] Example 7 includes a device as defined in Example 1 or Example 2, wherein the node analyzer is used to detect the non-serial-parallel nodes by analyzing the control dependence graph.
[0078] Example 8 includes a device as defined in Example 7, wherein the node analyzer is used to: identify the non-serial-parallel nodes when the control dependence graph has more than one parent node for the non-serial-parallel nodes.
[0079] Example 9 includes a non-transitory computer-readable medium that includes instructions that, when executed, cause a machine to at least: detect non-serial-parallel nodes in serialized code; generate instructions for prior nodes that consume operands; generate combination instructions for combining the results of the instructions; and output the combination instructions and the instructions to generate data flow code.
[0080] Example 10 includes a non-transitory computer-readable medium as defined in Example 9, wherein the consumed operand is an ignored operand.
[0081] Example 11 includes a non-transitory computer-readable medium as defined in Example 9 or Example 10, wherein the instructions, when executed, cause the machine to detect that the control flow graph of the serialized code is a non-serial-parallel flow graph.
[0082] Example 12 includes a non-transitory computer-readable medium as defined in Example 9 or Example 10, wherein the instructions are PICK instructions.
[0083] Example 13 includes a non-transitory computer-readable medium as defined in Example 9, wherein the consumed operand is a white value.
[0084] Example 14 includes a non-transitory computer-readable medium as defined in Example 13, wherein the instructions, when executed, cause the machine to generate instructions for consuming the white value.
[0085] Example 15 includes a non-transitory computer-readable medium as defined in Example 9 or Example 10, wherein the instructions, when executed, cause the machine to detect the non-serial-parallel nodes by analyzing the control dependence graph.
[0086] Example 16 includes a non-transitory computer-readable medium as defined in Example 15, wherein the instructions, when executed, cause the machine to: identify the non-serial-parallel nodes when the control dependence graph has more than one parent node for the non-serial-parallel nodes.
[0087] Example 17 includes a method for compiling code, the method including: detecting non-serial-parallel nodes in serialized code; generating instructions for a prior node that include consuming an operand; generating a combining instruction for combining the results of the instructions; and outputting the combining instruction and the instructions to generate data flow code.
[0088] Example 18 includes the method defined in Example 17, wherein the consuming operand is an ignore operand.
[0089] Example 19 includes the method defined in Example 17 or Example 18, further including: detecting that a control flow graph of the serialized code is a non-serial-parallel flow graph.
[0090] Example 20 includes the method defined in Example 17 or Example 18, wherein the instruction is a PICK instruction.
[0091] Example 21 includes the method defined in Example 17, wherein the consuming operand is a white value.
[0092] Example 22 includes the method defined in Example 21, further including: generating an instruction for consuming the white value.
[0093] Example 23 includes the method defined in Example 17 or Example 18, further including: detecting the non-serial-parallel nodes by analyzing a control dependence graph.
[0094] Example 24 includes the method defined in Example 23, further including: identifying the non-serial-parallel nodes when the control dependence graph has more than one parent node for the non-serial-parallel nodes.
[0095] Example 25 includes a device for converting control flow code to data flow code, the device including: means for detecting non-serial-parallel nodes in serialized code; and means for generating instructions to perform the following operations: generating instructions for a prior node that include consuming an operand; generating an instruction for combining the results of the instructions; and outputting the combining instruction and the instructions to generate data flow code.
[0096] Example 26 includes the device defined in Example 25, wherein the consuming operand is an ignore operand.
[0097] Example 27 includes the method defined in Example 25 or Example 26, further including: means for detecting that a control flow graph of the serialized code is a non-serial-parallel flow graph.
[0098] Example 28 includes the device defined in Example 25 or Example 26, wherein the instruction is a PICK instruction.
[0099] Example 29 includes the apparatus defined in Example 25, wherein the consumed operand is a white value.
[0100] Example 30 includes the apparatus defined in Example 29, wherein the means for generating instructions is further for generating means for consuming the white value.
[0101] Example 31 includes the apparatus defined in Example 25 or Example 26, wherein the means for detecting is for detecting the non-serial-parallel nodes by analyzing a control dependence graph.
[0102] Example 32 includes the apparatus defined in Example 31, wherein the means for detecting is for: identifying the non-serial-parallel nodes when the control dependence graph has more than one parent node for the non-serial-parallel nodes.
[0103] From the foregoing, it will be appreciated that example methods, apparatuses, and articles capable of efficiently converting non-serial-parallel serialized code into data flow code have been disclosed. In some examples, using a consumed operand in the converted instructions allows the original non-serial-parallel code to be executed on a parallel computing system (such as a high-performance computing system operating a data flow architecture).
[0104] Although certain example methods, apparatuses, and articles have been disclosed herein, the scope covered by this patent is not limited thereto. Instead, this patent covers all methods, apparatuses, and articles falling within the scope of the claims of this patent.
Claims
1. An apparatus for converting control flow code into data flow code, the apparatus comprising: A logic circuit; A node analyzer for detecting non-serial-parallel nodes in serialized code; And An instruction generator for: Generating instructions for prior nodes associated with the detected non-serial-parallel nodes, at least one of the instructions including a consume operand identified as an ignore operand, the ignore operand being used to cause any value assigned to the ignore operand to be consumed without being stored; Generating a combine instruction for combining the results of the instructions generated for the prior nodes; and Outputting the combine instruction and the instructions generated for the prior nodes to generate data flow code, wherein at least one of the node analyzer or the instruction generator is implemented by the logic circuit.
2. The apparatus according to claim 1, further comprising a non-serial-parallel detector for detecting that the control flow graph of the serialized code is a non-serial-parallel flow graph.
3. The device according to claim 1, characterized in that, The instruction is a PICK instruction.
4. The device according to claim 1, characterized in that, The consume operand includes a white value.
5. The device according to claim 4, characterized in that, The instruction generator is further for generating an instruction for consuming the white value.
6. The device according to claim 1, characterized in that The node analyzer is for detecting the non-serial-parallel nodes by analyzing a control dependence graph.
7. The device according to claim 6, characterized in that, The node analyzer is for: when the control dependence graph has more than one parent node for the non-serial-parallel node, identifying the non-serial-parallel node.
8. A method for compiling code, the method comprising: Detecting non-serial-parallel nodes in serialized code; Generating instructions for prior nodes associated with the detected non-serial-parallel nodes, at least one of the instructions including a consume operand identified as an ignore operand, the ignore operand being used to cause any value assigned to the ignore operand to be consumed without being stored; Generating a combine instruction for combining the results of the instructions generated for the prior nodes; And Outputting the combine instruction and the instructions generated for the prior nodes to generate data flow code.
9. The method according to claim 8, further comprising: Detecting that the control flow graph of the serialized code is a non-serial-parallel flow graph.
10. The method according to claim 8, wherein The instruction is a PICK instruction.
11. The method according to claim 8, wherein, The consume operand includes a white value.
12. The method according to claim 11, further comprising: Generating an instruction for consuming the white value.
13. The method according to claim 8, further comprising: Detecting the non-serial-parallel nodes by analyzing a control dependence graph.
14. The method according to claim 13, further comprising: When the control dependence graph has more than one parent node for the non-serial-parallel node, identifying the non-serial-parallel node.
15. A computer-readable medium comprising instructions that, when executed, cause a machine to perform the method according to any one of claims 8-14.
Citation Information
Patent Citations
Automatic Conversion of Text-Based Code Having Function Overloading and Dynamic Types into a Graphical Program for Compiled Execution
US20080022264A1
System and method for automatically generating a graphical program to perform an image processing algorithm
US6763515B1