A method and apparatus for generating a data flow diagram

By analyzing dependency information in the code block and generating a data flow graph, the problem that data flow graphs in the prior art cannot be applied to parallel computing is solved, and efficient compilation is achieved for heterogeneous processors.

CN114791808BActive Publication Date: 2025-07-18BEIJING QINGWEI INTELLIGENT INFORMATION TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210116285.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-07
Publication Date
2025-07-18
Estimated Expiration
2042-02-07

AI Technical Summary

Technical Problem

The existing data flow graph generation software cannot effectively consider the hardware environment, cannot be used by the compiler, and lacks details such as operators and operands, resulting in the inability to generate data flow graphs suitable for parallel computing.

Method used

Use the llvm front-end and intermediate representation (IR) to analyze the dependency information in the code block, generate the block data flow graph, and eliminate duplicate variables, establish the connection relationship between process blocks, and generate the data flow graph file. It is suitable for the front-end of the reconstructible compiler.

Benefits of technology

The generated data flow diagram can clearly reflect the parallel execution process of program code, and is suitable for compilation of heterogeneous processors, improving compilation efficiency and parallel computing support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114791808B_ABST
    Figure CN114791808B_ABST
Patent Text Reader

Abstract

The present invention provides a method and apparatus for generating a data flow diagram. The method includes: analyzing the dependency information of data in the code in the scenario where there is a parent loop in the code block; traversing each process block in the code block to generate a block data flow diagram corresponding to each process block; analyzing and adding dependency information to the data flow diagrams corresponding to each process block to generate a code data flow diagram corresponding to the code block; eliminating duplicate variables in the code data flow diagram, simplifying the process blocks and establishing connection relationships between each process block to generate a first data flow diagram file, so as to display the data flow diagram of the loop code through the first data flow diagram file. The above method utilizes the llvm front end, the intermediate representation (IR) and some existing analysis steps thereof. llvm is an open-source compiler framework that facilitates users to add compilation steps (passes) or modify the compilation process according to the architecture of their own processors. The present invention is used for the front end of a reconfigurable compiler, but is not limited to a reconfigurable compiler.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of compilation for heterogeneous processors, and particularly to a method and apparatus for generating a data flow graph during code compilation for loop acceleration. Background Art

[0002] A data flow graph (abbreviated as dfg) is a powerful graphical tool for describing the data processing process in a software system. From the perspective of data transfer and processing, the data flow graph depicts the movement and transformation process of data flow from input to output. Since it can clearly reflect the execution process of program code, it is often an important part of the compiler for heterogeneous processors.

[0003] Currently, there are some software (such as autoflowchart) that can automatically generate data flow graphs of code. However, these software are for programmers to analyze or display code logic, which is a manifestation of business logic. They only analyze the literal meaning at the code level, do not consider the hardware environment, or mostly consider the interface effects (such as layout and wiring) presented to the user by the data flow graph, and do not care about its application background. Due to the lack of details such as operators and operands, they cannot be used by the compiler. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for generating a data flow graph, which utilizes the llvm front end, intermediate representation (IR), and its existing partial analysis processes. Llvm is an open-source compiler framework that facilitates users to add or modify compilation processes (passes) according to the architecture of their own processors. It is used for the front end of a reconfigurable compiler, but is not limited to reconfigurable compilers. Any architecture that requires parallel execution of instructions can be used.

[0005] In the first aspect of the present invention, a method for generating a data flow graph is provided, including: in the scenario where there is parent loop code in a code block, analyzing the dependency information between sub-code blocks in the code block; traversing each process block in the code block to generate a block data flow graph corresponding to each process block; adding the dependency information to the data flow graph corresponding to each process block to generate a code data flow graph corresponding to the code block; eliminating duplicate variables in the code data flow graph, establishing the connection relationship between each process block, and generating a first data flow graph file, so as to display the data flow graph of the code block through the first data flow graph file.

[0006] Optionally, after eliminating duplicate variables in the code data flow graph, simplifying the process blocks, and establishing the connection relationships between the respective process blocks to generate a first data flow graph file, such that the data flow graph of the code block is presented through the first data flow graph file, the method further includes: counting the connection relationships between the respective process blocks in the code block; extracting process blocks that meet a preset condition and simplifying the process blocks in the code block; obtaining the association or nesting relationships between the respective process blocks; and complementing the dependency distances between the data in the respective process blocks to generate a second data flow graph file, where the dependency distance represents the number of loop iterations between multiple accesses to the data in memory during multiple accesses to the code block.

[0007] Optionally, after complementing the dependency distances between the data in the respective process blocks to generate a second data flow graph file, the method further includes: generating code block loading information.

[0008] Optionally, the dependency relationship includes a sequential order and / or an inclusion relationship.

[0009] Optionally, after obtaining the association or nesting relationships between the respective process blocks, the method further includes: deleting or converting useless process blocks in the respective process blocks into loop information.

[0010] Optionally, traversing the respective process blocks in the code block to generate a block data flow graph corresponding to each process block includes: presenting the respective process blocks in a tiled structure in an intermediate representation IR, parsing IR statements one by one, and filling the operators and operands of each IR into corresponding data structures to generate a block data flow graph corresponding to each process block.

[0011] Optionally, in a scenario where there is parent loop code in the code block, before analyzing the dependency information between sub-code blocks in the code block, the method further includes: converting code block files in different formats into the intermediate representation IR format file.

[0012] In a second aspect of the present invention, there is provided a data flow graph generation device, the device including: an analysis unit for analyzing the dependency information between sub-code blocks in the code block in a scenario where there is parent loop code in the code block; a first generation unit for traversing the respective process blocks in the code block to generate a block data flow graph corresponding to each process block; a second generation unit for adding the dependency information to the data flow graph corresponding to each process block to generate a code data flow graph corresponding to the code block; and a third generation unit for eliminating duplicate variables in the code data flow graph, establishing the connection relationships between the respective process blocks, and generating a first data flow graph file, such that the data flow graph of the code block is presented through the first data flow graph file.

[0013] Optionally, the device further includes: a statistics unit, configured to eliminate duplicate variables in the code data flow graph, establish connection relationships between the respective process blocks, and obtain a data flow graph file, so that after displaying the data flow graph of the code block through the first data flow graph file, the connection relationships between the respective process blocks in the code block are statistically analyzed; a simplification unit, configured to extract process blocks that meet preset conditions and simplify the process blocks in the code block; an acquisition unit, configured to acquire the association or nesting relationships between the respective process blocks; a fourth generation unit, configured to complete the dependency distance between the data in the respective process blocks and generate a second data flow graph file, where the dependency distance represents the number of loop times between multiple accesses to the memory in the code block during multiple accesses.

[0014] Optionally, the device further includes: a fifth generation unit, configured to generate the code block loading information after completing the dependency distance between the data in the respective process blocks and generating the second data flow graph file.

[0015] Optionally, the dependency relationship includes a sequence and / or an inclusion relationship.

[0016] Optionally, the device further includes: a deletion and conversion unit, configured to delete or convert useless process blocks in the respective process blocks into loop information after acquiring the association or nesting relationships between the respective process blocks.

[0017] Optionally, the first generation unit includes: a generation module, configured to present the respective process blocks in a tiled structure in the intermediate representation IR, parse the IR statements one by one, fill the operators and operands of each IR into corresponding data structures, and generate the block data flow graph corresponding to the respective process blocks.

[0018] Optionally, the device further includes: a conversion unit, configured to convert code block files in different formats into the intermediate representation IR format file before analyzing the dependency information between sub-code blocks in the code block in the scenario where there is parent loop code in the code block.

[0019] Hereinafter, the characteristics, technical features, advantages and implementation manners of a data flow graph generation method and system will be further described in a clear and understandable manner in conjunction with the accompanying drawings. Description of the Drawings

[0020] Figure 1 is a schematic flowchart for explaining the data flow graph generation method in an embodiment of the present invention;

[0021] Figure 2 is a general flowchart of the data flow graph in an embodiment of the present invention;

[0022] Figure 3 It is a schematic diagram of the data structure relationship in another embodiment of the present invention;

[0023] Figure 4 It is a structure diagram of a for loop (I) in another embodiment of the present invention;

[0024] Figure 5 It is a structure diagram of a while loop in one embodiment of the present invention;

[0025] Figure 6 It is a schematic diagram of the dependency distance calculation in another embodiment of the present invention;

[0026] Figure 7 It is a flowchart of the hierarchical calculation in another embodiment of the present invention;

[0027] Figure 8 It is a simplified for loop structure in another embodiment of the present invention;

[0028] Figure 9 It is a structure diagram of a loop with branches in another embodiment of the present invention;

[0029] Figure 10 It is a schematic diagram of the loading information data structure in one embodiment of the present invention;

[0030] Figure 11 It is an example of a dot graph in another embodiment of the present invention;

[0031] Figure 12 It is a schematic diagram of the batch display effect in another embodiment of the present invention;

[0032] Figure 13 It is a schematic diagram of a data flow graph generation device in another embodiment of the present invention. Specific embodiments

[0033] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0034] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0035] A data flow graph (abbreviated as DFG) is a powerful graphical tool for describing the data processing process in a software system. From the perspective of data transfer and processing, the data flow graph depicts the movement and transformation process of data flow from input to output. Since it can clearly reflect the execution process of program code, it is often an important part of the compiler for heterogeneous processors.

[0036] The task of the compiler is to translate the code logic into instructions corresponding to the target architecture. For traditional (sequentially executing instructions) architectures, after lexical analysis, syntax analysis, and semantic analysis of the code, the compiler generates an intermediate representation (IR) that can be directly mapped to its instruction set. After translation into instructions, only sequential loading and execution are required at runtime. However, with the development of the times, people's requirements for processor performance and power consumption are getting higher and higher. Sometimes, the IR cannot meet the requirements of parallel computing architectures (because it is not sequential loading and execution), so a more macroscopic logical representation method is needed, and the data flow graph can just show the situation of parallel computing. Obviously, the more the composition of the data flow graph is oriented to the architecture, the more conducive it is to translating into machine instructions.

[0037] One aspect of the present invention provides a data flow graph generation method, as Figure 1 shown, the data flow graph generation method includes:

[0038] Step S101, in the scenario where there is parent loop code in the code block, analyze the dependency information of the data in the code block.

[0039] Step S103, traverse each process block in the code block and generate a block data flow graph corresponding to each process block.

[0040] Step S105, analyze and add dependency information to the data flow graph corresponding to each process block to generate a code data flow graph corresponding to the code block.

[0041] Step S107, eliminate the duplicate variables in the code data flow graph, simplify the process blocks, and establish the connection relationships between the process blocks to generate a first data flow graph file, so as to display the data flow graph of the code block through the first data flow graph file.

[0042] In this embodiment, inherit the process runOnLoop at the loop level (so that this process is entered only when there is a loop in the code), and obtain the process block names and the parent loop in the loop. Perform standard (existing) loop range analysis, loop information analysis, and dependence analysis. Then parse the IR one by one and record the information into the corresponding data structures (operators and operands), and establish the associations between these data structures. After that, attribute these data structures to the current process block, and the data flow graph dfg structure of this block is obtained. Then, according to the previous dependence analysis, add the dependence relationships between the corresponding operators (to improve the dfg). When all the process blocks are obtained, establish the logical associations (sequence or inclusion relationship) between the process blocks, and at the same time eliminate the (duplicate) data variables shared by multiple blocks to form the overall dfg. At this time, the dfg data structure can be saved as the original dfg file for comparison and error checking with the optimized dfg.

[0043] Next is the simplification of the process blocks. Remove the parts that the compiler does not care about, and only leave the loop body itself and the necessary logical blocks. The steps include simplifying the memory access calculation; inferring the relationships (sequence or nesting) according to the connections of the process blocks; deleting or converting the useless process blocks into loop information; completing information such as the dependence distance; generating the simplified dfg; and finally generating the data loading information.

[0044] Through the embodiment provided by the present application, in the scenario where there is parent loop code in the code block, analyze the dependence information of the data in the code block; traverse each process block in the code block to generate the block data flow graph corresponding to each process block; analyze and add the dependence information to the data flow graph corresponding to each process block to generate the code data flow graph corresponding to the code block; eliminate the duplicate variables in the code data flow graph, simplify the code block, and establish the connection relationships between the process blocks to generate a first data flow graph file, so as to display the data flow graph of the code block through the first data flow graph file. The above method utilizes the llvm front end, the intermediate representation (IR), and some existing analysis steps thereof. llvm is an open-source compiler framework that facilitates users to add compilation steps (passes) or modify the compilation process according to the architecture of their own processors. It is used for the front end of a reconfigurable compiler, but is not limited to a reconfigurable compiler.

[0045] Optionally, duplicate variables in the code data flow graph are eliminated, the connection relationships between individual process blocks are established, and a data flow graph file is obtained, so that after the data flow graph of the code block is displayed through the first data flow graph file, the above method may further include: counting the connection relationships between individual process blocks in the code block; extracting process blocks that meet preset conditions and simplifying the process blocks in the loop code; obtaining the association or nesting relationships between individual process blocks; completing the dependency distances between the data in individual process blocks, and generating a second data flow graph file, where the dependency distance represents the number of loop iterations between multiple accesses to data in memory during multiple accesses to the code block.

[0046] Optionally, after completing the dependency distances between the data in individual process blocks and generating a second data flow graph file, the above method may further include: generating code block loading information.

[0047] In this embodiment, the code block loading information includes but is not limited to the amount of data to be loaded, the amount of data to be generated, the number of computing units to be used, the computing order level of the computing unit, and whether there are operators of unknown types (see Figure 10 Demand).

[0048] Optionally, the dependency relationship includes a sequence and / or an inclusion relationship.

[0049] Optionally, after obtaining the association or nesting relationships between individual process blocks, the above method may further include: deleting or converting useless process blocks in individual process blocks into loop information.

[0050] In this embodiment, the loop information includes but is not limited to the loop start value, end value, loop step size, loop direction, loop variable name, the nesting depth of this loop, the name of this loop, the name of the parent loop, the name of the child loop, and the initialization list <variable name: value> (see Figure 10 Scalar).

[0051] Optionally, traversing individual process blocks in the code block and generating a block data flow graph corresponding to each process block includes: presenting individual process blocks in a tiled structure in the intermediate representation IR, parsing each IR statement one by one, filling the operator and operands of each IR into the corresponding data structure, and generating a block data flow graph corresponding to each process block.

[0052] Optionally, in the scenario where there is parent loop code in the code block, before analyzing the dependency information between sub-code blocks in the code block, the above method may further include; converting code block files in different formats into intermediate representation IR format files.

[0053] As an optional embodiment, the present application also provides a method for a compiler to automatically generate a data flow graph for accelerating loop code. As Figure 2As shown, the compiler automatically generates a flowchart of the data flow graph method. The specific process is as follows.

[0054] 1. Use the clang / clang++ command to convert C and C++ code into a.ll (intermediate representation IR) format file.

[0055] Such as clang xxx.c -emit-llvm -g -c -S -o xxx.ll -Xclang -disable-O0-optnone

[0056] Where xxx.c is the C code file to be optimized for loop code; xxx.ll is the IR format file; the others are necessary compilation options.

[0057] 2. Inherit the loop-level process runOnLoop (so that this process is only entered when there is a loop in the code), and obtain the process block names and parent loops in the loop. Perform standard (existing) loop range analysis, loop information analysis, and dependency analysis. Then parse the IR one by one and record the information into the corresponding data structures (operators and operands), and establish associations between these data structures. Then attribute these data structures to the current process block to obtain the dfg structure of the block. Then, according to the previous dependency analysis, add the dependency relationships between the corresponding operators (improve the dfg). When all process blocks are obtained, establish the logical associations (sequential or inclusion relationships) between the process blocks, and at the same time eliminate the common (duplicate) data variables among multiple blocks to form the overall dfg. At this time, the dfg data structure can be saved as the original dfg file for comparison and error checking with the optimized dfg.

[0058] After that is the simplification of the process block, removing the parts that the compiler does not care about, and only leaving the loop body itself and the necessary logical blocks. The steps include simplifying memory access calculations; inferring relationships (sequence or nesting) according to the connections of the process blocks; deleting or converting useless process blocks into loop information; completing information such as dependency distances; generating a simplified dfg; and finally generating data loading information.

[0059] In this embodiment, data structures are defined, including operator nodes OpNode, data nodes VarNode, process blocks ProcBlock, dependency information DepInfo, branch information BranchEdge, etc., as Figure 3 shown, the schematic diagram of the data structure relationship.

[0060] Among them, Node is the base class of OpNode, Var Node, and ProcBlock, containing basic information such as name and ID. The operand VarNode includes information such as the type and size of the data, and whether it is an input or output of certain operators; the operator OpNode includes inputs, outputs (operands), calculation levels, predecessor, successor branches, and dependency information; the procedure block ProcBlock includes all operators and operands contained in the block, as well as association information with other procedure blocks (including specific operator nodes).

[0061] In this embodiment, each block within the loop is constructed and mapped to a ProcBlock. A loop consists of multiple blocks and may also include sub-loops. In the runOnloop of llvm's loop pass, when there is loop nesting, the loop blocks appear in the order of child first and then parent. If there are multiple sibling loops, the sibling loops appear in order. In the case of complex loops, a depth-first traversal is performed. A for loop contains four parts: cond, body, inc, and end. As shown in the following figure, there can be branch blocks (if.then, if.else) between for.body and for.inc in the loop.

[0062] When the parent loop appears, it will carry the previously appeared sub-loops. Therefore, to save time, only the outermost loop needs to be analyzed once. As Figure 4 shown, the for loop structure diagram (one).

[0063] For(i = 0; i < 10; ++i) / / for.cond + for.inc

[0064] {

[0065] … / / for_body can have no calculations and only jumps

[0066] For(j = 0; j < 20; j += 2) / / for.cond1 + for.inc1

[0067] {

[0068] A[i][j] = …; / / for.body1

[0069] }

[0070] … / / for.end

[0071] }

[0072] In the case of a while loop, there is no inc module, and the rest is the same.

[0073] AsFigure 5 As shown, the while loop structure diagram. All blocks (including branches and labels) are presented in a flattened structure in the IR. The transformation of blocks is implemented using br instructions, similar to the assembly style.

[0074] In this embodiment, each loop block is traversed, that is, each IR statement is parsed one by one, and the operator and operands of each IR are filled into the corresponding data structure. The form of the IR is as follows:

[0075] for.cond: ; preds = %for.inc, %entry

[0076] %0 = load i32, i32* %i, align 4

[0077] %cmp = icmp slt i32 %0, 10

[0078] br i1 %cmp, label %for.body, label %for.end

[0079] for.body: ; preds = %for.cond

[0080] %1 = load i32, i32* %i, align 4

[0081] %2 = load i32, i32* %i, align 4

[0082] %mul = mul nsw i32 %1, %2

[0083] %3 = load i32, i32* %i, align 4

[0084] Among them, the part before the colon such as for.cond and for.body represents the loop block name; those starting with % are variable names or block names; the part immediately following "=" is the operator;... Since llvm has a mechanism to ensure the uniqueness of variable names, there is no need to worry about ambiguity issues. In this way, it is very easy to map the IR to the custom data structure.

[0085] In this embodiment, dependency information is added to the operator.

[0086] Using the llvm-built-in dependency analysis process executed before, we can obtain memory access dependencies. These dependency information prompt the compiler that the read and write order of the memory operated by a certain operator (such as load, store, etc.) cannot be disrupted. Most of this process is correct, but I later found that there will be errors when calculating the dependency distance.

[0087] The dependence distance simply means the number of loop iterations between multiple accesses to a certain memory location within a loop.

[0088] Dependence information between two OpNodes indicates that the execution order of these two operations cannot be disrupted and they cannot be executed in parallel. The operation being depended on must be executed first, and the dependent operation must be executed later.

[0089] For example

[0090] for(i=10; i<100; ++i)

[0091] {

[0092] a[i]= a[i-2];

[0093] }

[0094] a[i] will be accessed again after two rounds.

[0095] However, the addressing calculation can be very complex. In addition to linear (arithmetic) operations, there may also be two-dimensional or multi-dimensional operations.

[0096] Assuming that we don't care about the specific calculation of the subscript and write it in the form of A[f(i)] = A[g(i)] (only care about one dimension at a time for multi-dimensions), the dependence occurs when f(i) equals g(i). Here, i is the loop independent variable, and f(i) and g(i) are the subscripts. The independent variable has a domain (discrete) of the loop upper and lower bounds [b, t]. Assuming that both f and g are monotonic, then f(i) and g(i) have their respective value ranges, and the extreme values occur at the boundaries of the independent variable. Such as [f(b), f(t)] and [g(b), g(t)] (in the case where both are monotonically increasing). As Figure 6 shown in the schematic diagram of dependence distance calculation.

[0097] Assume that the value range of f is represented by [fmin, fmax], and the value range of g is represented by gmin, gmax. Obviously, when

[0098] fmin < gmax < fmax or fmin < gmin < fmax, there is a dependence. ①

[0099] For example, in Figure 6 the shown case, the first occurrence of dependence is when g(d) = f(b). Here, d is the dependence distance.

[0100] Specific method:

[0101] Write a function that can calculate the value of the expression, and use this function to calculate the addressing value (for the following calculations).

[0102] Replace the independent variables \(i\) of the subscript expressions \(f\) and \(g\) with the upper and lower limits \(b\) and \(t\) respectively to obtain \(f_{min}\), \(f_{max}\); \(g_{min}\), \(g_{max}\) and the monotonicity of \(f\) and \(g\).

[0103] Determine the dependence and the first dependence value, which is \(f(b)\) or \(f(t)\), according to the value range relationship ①. Assume it is \(f1\).

[0104] Calculate the distance using the bisection method.

[0105] For example, first let \(d=(t - b) / 2\) to obtain the value of \(g(d)\), assume it is \(g2\).

[0106] If ((g2 > f1 && g is increasing) || (g2 < f1 && g is decreasing))

[0107] g3 = g(d / 2)

[0108] else if ((g2 > f1 && g is decreasing) || (g2 < f1 && g is increasing))

[0109] g3 = g(d*3 / 2)

[0110] else if (g2 == f1)

[0111] return d;

[0112] Recursive calculation can be performed.

[0113] In this embodiment, calculate the operator level to achieve parallel computing. Since parallel computing hopes to obtain all (or as many as possible) intermediate results simultaneously, we need to know which steps need to be executed simultaneously. With the operator level, this information can be obtained. For example:

[0114] x = a 2 + b 2 + c 2

[0115] We hope to place the square operations of the three numbers \(a\), \(b\), and \(c\) at the first level, place \(t = a\) 2 + b 2 at the second level, and place \(t + c\) 2 at the third level, so as to give full play to the advantages of parallel computing.

[0116] The method for calculating the level of the operator node is as Figure 7 shown, the flow chart of level calculation.

[0117] For an operator node, if it has no inputs or only pure inputs (i.e., the inputs do not come from the outputs of other operators), assign the node level a value of 0; if some of its inputs come from the outputs of the previous-level operations, after recursively calculating the level of the previous level, add 1 as the level of this operator node.

[0118] During this process, the levels of the upper-level nodes of this operator node are calculated incidentally. However, it should also be considered whether the upper-level nodes are in the same or a process block containing the node being calculated. If not, stop the upward recursive calculation.

[0119] In this embodiment, the method for simplifying nodes in a single process block is as follows:

[0120] Keep only one of the same operations with the same input parameters. Because the same variables will appear in multiple blocks, and we hope that a variable only occupies memory once, and there should be no repeated memory allocation.

[0121] Delete operations that the architecture does not care about, such as sext (sign extension), zext (0-type extension), bitcast (type conversion), etc.

[0122] Delete debug-related instructions.

[0123] Using expression calculation, simplify the multiple loads of a multi-dimensional array into one (for the extraction of loaded information). For example, when there is a getelementptr in the original IR, it represents an addressing operation, and it has 3 parameters, namely the starting address of the variable, the offset address of the variable, and the compilation address of the structure member. For example

[0124] %tmp6 = getelementptr %struct.munger_struct* %P, i32 2, i32 1

[0125] Represents the 1st member of the 2nd array index of struct.munger (the numbers start from 0). Sometimes the array index and member number are also calculated values, so the calculation needs to be expanded. Finally, it is presented in the form of an addressing expression in the figure and recorded in a custom.pre file. For example:

[0126] ID: Datum618, detail: vla1[k.2*10*reg2][j.0*reg2][(i.0-1)]: pointerinteger, size: 4, addr: 16, writen: 0;

[0127] Its format is ID: [id name], detail: variable addressing expression (including calculations), pointer type, length of each data in the array, offset address, and whether it is an output.

[0128] Release the dependencies that disappeared due to simplification.

[0129] Reason and establish connections between blocks

[0130] A for loop contains four parts, cond, body, inc, and end. As shown in the figure below, it shows the structure of a nested loop.

[0131] Such as Figure 4 As shown, it is a schematic diagram of the for loop structure. Among them, cond has one or more phi (branch option) nodes, assigns initial values or increments (or cumulative values) to one or more variables, and then has a comparison and a branch jump (br) determined by the comparison result.

[0132] inc has add and br.

[0133] The body can have operations or only one br (jumping to other loops). If there are sub-loops, there are the following situations:

[0134] 1. Only sub-loops: The parent loop has only a br that jumps to the cond of the sub-loop.

[0135] 2. There is ordinary calculation before the sub-loop: The ordinary calculation is the content of the body of the parent loop, and then br jumps to the sub-loop.

[0136] 3. There is ordinary calculation after the sub-loop: The body of the parent loop has only a br that jumps to the cond of the sub-loop. After the sub-loop is executed, it jumps to for.end (belonging to the sub-loop) to execute the calculation after the sub-loop, and then executes for.inc of the parent loop.

[0137] 4. There are multiple sub-loops in the parent loop: The body of the parent loop has only one br that jumps to the cond of the first sub-loop. After one sub-loop ends, it jumps to the cond of the next sub-loop. If there is a non-loop between the two loops, the non-loop is considered as for.end of the previous sub-loop. The execution order is as in 3.

[0138] end can have calculations (when it is not a perfect loop) or not. Different from the body, end is only executed once and is executed at the end of the sub-loop and belongs to the sub-loop.

[0139] During simplification, cond and inc with little calculation are simplified to sub-graphs containing only additional information plus br, and the calculations in the end part are taken as another sub-graph (if there are calculations). Such asFigure 8 As shown, the simplified for loop structure.

[0140] Among them, cond+inc is simplified and appended to the body graph in the form of loop information, rather than a graphical connection. The loop information format is as follows:

[0141] Loop information name (same as body name), loop variable name (such as i), loop start value, loop end value, loop increment value, increment direction (+, -), depth of nesting of this loop, parent loop name.

[0142] There is no inc module in the while loop, so there is no need to simplify it into loop information.

[0143] When there is an if branch, as Figure 9 shown, the panoramic view of the loop structure with branches.

[0144] The part between if_then or if_else and if_end is nested, and the front and back are sequential.

[0145] Transformations to be done:

[0146] 1. Retain all branches except cond and inc.

[0147] 2. Convert cond and inc to br (the jump node of the original cond).

[0148] 3. Change the jump form of the sub-loop to a nested form.

[0149] 4. Also add the variables in the end part to the symbol table.

[0150] In this embodiment, loading information is generated. The purpose of generating the dfg is not only to generate a visual graph structure, but more importantly, to transfer as much information of the parsed code as possible to the backend of the compiler to make its implementation more convenient. Therefore, this loading information is generated.

[0151] Such as Figure 10As shown, it is a structure diagram of loaded information data. Among them, FlowElems represents the overall loaded information of a process block, which is similar to the traditional process structure. It includes three segments: the uninitialized data segment bss, the initialized data segment data, and the code segment code (the content of the code segment is left for the back-end compiler to fill in). Each segment contains the starting address and size information; symbols is the symbol table, which consists of a series of SymbolInfo structures. Each SymbolInfo indicates information such as the name, offset address, length, and whether it is an output of a variable; demand represents the resources that the process block needs to occupy, including the number of operators (PE_num), the number of memory banks load_num that need to be loaded simultaneously, the number of memory banks store_num that need to be stored simultaneously, the maximum number of levels levels, and whether there are unknown operators hasUnknownOp; Scalar represents loop information, corresponding to what was introduced before, but also has the sub-loop name subLoop and the initial value table initTable.

[0152] To record the generated graph for later manual inspection or reusing the data structure, the graph is converted into a.dot file using the standard format. With the previous data structure, generating the.dot file is a matter of course. As Figure 11 shown, it is an example dot graph.

[0153] Among them, the dfg file follows the dot format standard. The elements include operators, data, blocks, loop information, and the connections between them.

[0154] The black rectangular box represents data, and the four cells respectively represent the data name, type, type size, and the number of data.

[0155] The ellipse represents an operator, which includes the operator name and the priority (separated by :). The priority is calculated according to the structure of the graph, but the dependency relationship is not considered currently.

[0156] An operator may have one input and one output, which are connected by edges. The output may in turn become the input of another operator. Through such connections, a graph is naturally formed. The input edge has a number indicating the serial number of the input data as the parameter of the operator.

[0157] The edge of arrow line a represents memory dependence and points to the dependent instruction (operator). Memory dependence has three space-separated parameters. The first represents the node type, which can be's' (SingleInstruction),'m' (MultiInstruction), 'pi' (PiBlock), 'r' (Root), and 'u' (unknown). The second parameter represents the constraint type, which can be 'c' (confuse), 'f' (flow), 'o' (output), 'a' (anti), and 'i' (input). The third parameter represents the dependence distance, which can be a number or a vector, and each number in it corresponds to a one-dimensional subscript in sequence.

[0158] The large boxes c and d containing data and operations represent the boundaries of blocks, such as loops, branches, etc.

[0159] The dotted arrow line b connecting different blocks or operators is to let us know which one to execute next (because blocks at the same level should have an order, and some operators have no inputs or outputs). Nested blocks are executed according to the order of the sub-loops in the loop information of the parent loop.

[0160] The loop condition and loop increment blocks are replaced by a set of loop information. The loop information includes the loop variable name, start, end, step values, increment direction, the depth of this loop, and the sub-loop name.

[0161] In this embodiment, the process (pass) of the present invention is called using the llvm optimization command opt to generate a.dot (graph description) file.

[0162] Such as opt -load dfg.so -scalar-evolution -dfg -analyze pathTo / xxx.ll

[0163] Where -load dfg.so is the command option to call this process; -dfg is the option to execute this process; pathTo / xxx.ll is the path of the previously generated IR file.

[0164] Generate an html file (for convenient batch checking).

[0165] When verifying the correctness of the entire project, a large amount of inspection and verification work is required. Comparing file by file takes a lot of time. Therefore, this method is invented to facilitate batch inspection. The method is to place all the dot files to be inspected in the same path, traverse all the.dot files in this path, and use the xdot tool to convert the.dot files into png or jpg files. Write a script to generate an html file. Each section of the script contains a record to be inspected, including information such as file name, C / C++ source code (link), IR code (link), dfg graph before simplification (png link), and dfg graph after simplification (png format). As Figure 12 shown, the schematic diagram of the batch display effect.

[0166] In this embodiment, the IR of llvm can be converted into a custom dfg data structure. Remove the loop conditions and loop endings of the above dfg data structure from the graph and convert them into loop information. Simplify the multi-step addressing calculation into a one-step addressing operation. Judge the compilability of the code according to the architecture instruction set and constraints. Recalculate the memory dependence distance based on the existing dependence analysis process. Save the custom dfg data structure as a.dot format file. Convert the.dot file into a graphic file and display it batch by batch in a browser for easy error checking.

[0167] In this embodiment, using the existing partial analysis process of the llvm front end and the intermediate representation (IR), a data flow graph is automatically generated for the compiler. Among them, the process of automatically generating the data flow graph for the compiler generates a data structure with a relational graph as the core, which can be viewed manually and can also provide guidance for the subsequent automatic instruction set mapping and data banking of the compiler. It is an important part of the reconfigurable compiler.

[0168] Among them, this embodiment can be used for the front end of the reconfigurable compiler, but is not limited to the reconfigurable compiler. Any architecture that requires parallel execution of instructions can be used.

[0169] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes contributions to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), including several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0170] In this embodiment, a data flow graph generation device is further provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated here. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0171] Figure 13 is a structural block diagram of a data flow graph generation device according to an embodiment of the present invention. As Figure 13 shown, the data flow graph generation device includes:

[0172] An analysis unit 1401, configured to analyze the dependency information of data in a code block in a scenario where there is parent loop code in the code block.

[0173] A first generation unit 1403, configured to traverse each process block in the code block and generate a block data flow graph corresponding to each process block.

[0174] A second generation unit 1405, configured to analyze and add dependency information to the data flow graph corresponding to each process block, and generate a code data flow graph corresponding to the code block.

[0175] A third generation unit 1407, configured to eliminate duplicate variables in the code data flow graph, simplify the process block, and establish the connection relationship between each process block, and generate a first data flow graph file, so as to display the data flow graph of the code block through the first data flow graph file.

[0176] Through the embodiment provided by this application, the analysis unit 1401 analyzes the dependency information of data in the code block in a scenario where there is parent loop code in the code block; the first generation unit 1403 traverses each process block in the code block and generates a block data flow graph corresponding to each process block; the second generation unit 1405 analyzes and adds dependency information to the data flow graph corresponding to each process block, and generates a code data flow graph corresponding to the code block; the third generation unit 1407 eliminates duplicate variables in the code data flow graph, simplifies the process block, and establishes the connection relationship between each process block, and generates a first data flow graph file, so as to display the data flow graph of the code block through the first data flow graph file. The above method utilizes the llvm front end, intermediate representation (IR), and some existing analysis steps thereof. Llvm is an open-source compiler framework that facilitates users to add compilation steps (passes) or modify the compilation process according to the architecture of their own processors. It is used for the front end of a reconfigurable compiler, but is not limited to a reconfigurable compiler.

[0177] Optionally, the above device may further include: a statistics unit, configured to eliminate duplicate variables in the code data flow graph, establish connection relationships between respective process blocks, and obtain a data flow graph file, so that after the data flow graph of the code block is displayed through the first data flow graph file, the connection relationships between respective process blocks in the code block are counted; a simplification unit, configured to extract process blocks that meet preset conditions and simplify the process blocks in the code block; an acquisition unit, configured to acquire the association or nesting relationships between respective process blocks; a fourth generation unit, configured to complete the dependency distance between data in respective process blocks and generate a second data flow graph file, where the dependency distance represents the number of loop times between multiple accesses to data in the memory during multiple accesses to the code block.

[0178] Optionally, the above device may further include: a fifth generation unit, configured to generate code block loading information after completing the dependency distance between data in respective process blocks and generating a second data flow graph file.

[0179] Optionally, the dependency relationship includes a sequence and / or an inclusion relationship.

[0180] Optionally, the above device may further include: a deletion and conversion unit, configured to delete or convert useless process blocks in respective process blocks into loop information after acquiring the association or nesting relationships between respective process blocks.

[0181] Optionally, the above first generation unit 1403 may include: a generation module, configured to present respective process blocks in a tiled structure in the intermediate representation IR, parse IR statements one by one, fill the operators and operands of each IR into corresponding data structures, and generate a block data flow graph corresponding to each process block.

[0182] Optionally, the above device may further include: a conversion unit, configured to convert code block files in different formats into intermediate representation IR format files before analyzing the dependency information between sub-code blocks in the code block in a scenario where there is parent loop code in the code block.

[0183] It should be noted that the above respective modules may be implemented by software or hardware. For the latter, it may be implemented in the following ways, but not limited thereto: the above modules are all located in the same processor; or, the above respective modules are separately located in different processors in any combination form.

[0184] An embodiment of the present invention further provides a storage medium, in which a computer program is stored, where the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0185] Optionally, in this embodiment, the above storage medium may be configured to store a computer program for executing the following steps:

[0186] S1. In the scenario where there is parent loop code in the code block, analyze the data dependency information in the code block;

[0187] S2. Traverse each process block in the code block to generate a block data flow diagram corresponding to each process block;

[0188] S3. Analyze and add dependency information to the data flow diagrams corresponding to each process block to generate a code data flow diagram corresponding to the code block;

[0189] S4. Eliminate duplicate variables in the code data flow diagram, simplify the process blocks and establish the connection relationships between each process block to generate a first data flow diagram file, so as to display the data flow diagram of the code block through the first data flow diagram file.

[0190] Optionally, in this embodiment, the above storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs that can store computer programs.

[0191] An embodiment of the present invention further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0192] Optionally, the above electronic device may further include a transmission device and input / output devices. Among them, the transmission device is connected to the above processor, and the input / output device is connected to the above processor.

[0193] Optionally, in this embodiment, the above processor may be configured to execute the following steps through a computer program:

[0194] S1. In the scenario where there is parent loop code in the code block, analyze the data dependency information in the code block;

[0195] S2. Traverse each process block in the code block to generate a block data flow diagram corresponding to each process block;

[0196] S3. Analyze and add dependency information to the data flow diagrams corresponding to each process block to generate a code data flow diagram corresponding to the code block;

[0197] S4. Eliminate duplicate variables in the code data flow diagram, simplify the process blocks and establish the connection relationships between each process block to generate a first data flow diagram file, so as to display the data flow diagram of the code block through the first data flow diagram file.

[0198] Optionally, specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated herein.

[0199] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a sequence different from that here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.

[0200] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for generating a data flow diagram, characterized in that, The method includes: In the scenario where there is parent loop code in the code block, analyzing the dependency information of the data in the code block; Traversing each process block in the code block to generate a block data flow graph corresponding to each process block; Analyzing and adding the dependency information to the data flow graphs corresponding to each process block to generate a code data flow graph corresponding to the code block; Eliminating duplicate variables in the code data flow graph, simplifying the process blocks, and establishing the connection relationships between the process blocks to generate a first data flow graph file, so as to display the data flow graph of the code block through the first data flow graph file; After the eliminating duplicate variables in the code data flow graph, simplifying the process blocks, and establishing the connection relationships between the process blocks to generate a first data flow graph file, so as to display the data flow graph of the code block through the first data flow graph file, the method further includes; Counting the connection relationships of each process block in the code block; Extracting process blocks that meet preset conditions and simplifying the process blocks in the code block; Obtaining the association or nesting relationships between the process blocks; Completing the dependency distance between the data in each process block to generate a second data flow graph file, where the dependency distance represents the number of loop iterations between multiple accesses to the data in memory during multiple accesses to the code block.

2. The data flow diagram generation method according to claim 1, wherein After the completing the dependency distance between the data in each process block to generate a second data flow graph file, the method further includes: Generating the code block loading information.

3. The method for generating a data flow diagram according to any one of claims 1 to 2, characterized in that, The dependency information includes a sequence and / or an inclusion relationship.

4. The method for generating a data flow diagram according to claim 1, wherein After the obtaining the association or nesting relationships between the process blocks, the method further includes: Deleting or converting useless process blocks in each process block into loop information.

5. The data flow diagram generation method according to claim 1, wherein The traversing each process block in the code block to generate a block data flow graph corresponding to each process block includes: Each process block is presented in a flat structure in the intermediate representation IR, and the IR statements are parsed one by one, and the operators and operands of each IR are filled into the corresponding data structures to generate a block data flow graph corresponding to each process block.

6. The data flow diagram generation method according to claim 1, characterized in that Before the analyzing the dependency information between the sub-code blocks in the code block in the scenario where there is parent loop code in the code block, the method further includes; Converting code block files in different formats into IR format files.

7. A data flow diagram generation device, characterized in that, The apparatus includes: An analysis unit for analyzing the dependency information of the data in the code block in the scenario where there is parent loop code in the code block; A first generation unit for traversing each process block in the code block to generate a block data flow graph corresponding to each process block; A second generation unit for analyzing and adding the dependency information to the data flow graphs corresponding to each process block to generate a code data flow graph corresponding to the code block; A third generation unit for eliminating duplicate variables in the code data flow graph, simplifying the process blocks, and establishing the connection relationships between the process blocks to generate a first data flow graph file, so as to display the data flow graph of the code block through the first data flow graph file; It further includes: A statistical unit, which is used to eliminate duplicate variables in the code data flow graph, establish the connection relationships between the respective process blocks, and obtain a data flow graph file, so that after the data flow graph of the code block is displayed through the first data flow graph file, the connection relationships between the respective process blocks in the code block are statistically analyzed; A simplification unit, which is used to extract process blocks that meet preset conditions and simplify the process blocks in the code block; An acquisition unit, which is used to acquire the association or nesting relationships between the respective process blocks; A fourth generation unit, which is used to complete the dependency distance between the data in the respective process blocks and generate a second data flow graph file, where the dependency distance represents the number of loop iterations between multiple accesses to the data in memory when the code block is accessed multiple times.

8. The data flow diagram generation device according to claim 7, wherein, The apparatus further includes: A fifth generation unit, which is used to generate the code block loading information after completing the dependency distance between the data in the respective process blocks and generating the second data flow graph file.

Citation Information

Patent Citations

  • Technology To Use Control Dependency Graphs To Convert Control Flow Programs Into Data Flow Programs

    CN108984210A