High-Level Synthesis Process Layout Method
By using a planar algorithm in FPGA layout and wiring to build control data flow diagrams, segmenting and scheduling layout constraints, and inserting pipelines for delay balancing, the congestion and delay problems in FPGA layout and wiring are solved, and circuit throughput and design quality are improved.
Patent Information
- Application Number
- CN202211039147.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-29
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-08-29
AI Technical Summary
The prior art is difficult to accurately estimate interconnection delays in FPGA layout and wiring, resulting in layout and wiring congestion, increasing the overall wiring length, reducing circuit throughput, and ignoring the Block boundary interconnection can easily lead to longer wiring and delays.
The control data flow chart is constructed using a planar planning algorithm, and layout constraints are obtained through segmentation and scheduling, pipelines are inserted for delay balancing, resource partitioning and binding are combined with the FPGA architecture, and comprehensive netlists and layout and routing netlists are generated to reduce layout and routing congestion and delays across Block boundaries.
It effectively reduces layout and wiring congestion, reduces delay, improves the overall throughput and design quality of the circuit, and optimizes the running time of the physical synthesizer.
Smart Images

Figure CN115422876B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of circuit simulation, and particularly to a flow layout method for high-level synthesis. Background Art
[0002] High-level Synthesis (HLS) refers to the process of automatically converting a logical structure described in a high-level language into a circuit model described in a low-level abstraction language. HLS tools are efficient and fast, which can reduce the design time of hardware engineers and enable software engineers to complete hardware design.
[0003] However, in related technical solutions, one of the main reasons for the quality gap between HLS design and manual design is that it is difficult to accurately estimate the interconnect delay at the HLS level and obtain a good global physical layout. Especially in the layout and routing of Field Programmable Gate Arrays (FPGAs), the physical synthesizer often uses resources that are relatively close, resulting in congestion in layout and routing, increasing the overall wiring length, and thus reducing the throughput of the circuit. In addition, the resources of FPGAs are arranged in the form of Configurable Logic Blocks (CLBs), and the resource conditions of different blocks may be different. When ignoring the boundary crossing blocks for interconnection, it is easy to bring longer wiring and delay. Summary of the Invention
[0004] In view of this, to at least partially solve one of the above technical problems or defects, an object of an embodiment of the present invention is to provide a flow layout method for high-level synthesis that can effectively reduce congestion in layout and routing and reduce delay.
[0005] The technical solution of the present application provides a flow layout method for high-level synthesis, including the following steps:
[0006] Obtain the circuit description of the target circuit, and construct a control data flow graph corresponding to the target circuit according to the circuit description; the control data flow graph is a directed graph representing the arithmetic operations of the target circuit;
[0007] Segment the control data flow graph by a floorplanning algorithm to obtain layout constraints;
[0008] Schedule the control data flow graph, and bind the scheduling result to obtain a register transfer level description;
[0009] Obtain a target netlist according to the layout constraints from the register transfer level description, and determine the flow layout of the target circuit according to the target netlist.
[0010] In a feasible embodiment of the solution of the present application, the splitting of the control data flow graph by the plane planning algorithm to obtain layout constraints includes:
[0011] Obtain the FPGA architecture of the target circuit in the circuit description;
[0012] Compile the circuit module at the register transfer level according to the function corresponding to the data flow process in the control data flow graph;
[0013] Partition the FPGA architecture according to the circuit module, and determine the cost function of the partitioning result;
[0014] Determine that the cost function is the minimum value or the resources in the partition reach the critical value of the resource constraint by performing splitting iteration on the partitioning result of the FPGA architecture, and output the layout constraints.
[0015] In a feasible embodiment of the solution of the present application, the scheduling of the control data flow graph and binding the scheduling result to obtain the register transfer level description includes:
[0016] Schedule the subgraph of the control data flow graph;
[0017] Insert a pipeline into the connection line between nodes in the control data flow graph to balance the delay;
[0018] Mathematically integrate the scheduling result and the result after delay balancing, and bind according to the result after mathematical integration with the target circuit to obtain the register transfer level description.
[0019] In a feasible embodiment of the solution of the present application, the layout constraints include at least one of timing constraints or physical constraints; the target netlist includes at least one of a synthesis netlist or a placement and routing netlist; the process of obtaining the target netlist according to the layout constraints for the register transfer level description and determining the flow layout of the target circuit includes:
[0020] Construct a first input according to the timing constraint and / or the physical constraint;
[0021] Construct a second input according to the register transfer level description;
[0022] Through the FPGA physical synthesizer, integrate and process the first input and the second input and output a synthesis netlist;
[0023] Through the FPGA physical synthesizer, perform placement and routing processing according to the first input and the second input and output a placement and routing netlist;
[0024] Determine the flow layout of the target circuit according to the comprehensive netlist and the placement and routing netlist.
[0025] In a feasible embodiment of the solution of this application, the cost function is used to characterize the number of wires at the partition boundary of the FPGA architecture; the cost function is:
[0026]
[0027] where C is the cost value, v i and v j represent nodes in the control data flow graph, i = 1, 2, 3, … n, j = 1, 2, 3, … n, n is a positive integer, E represents the set of FIFO channels between nodes, and e ij is the connection line between v i and v j row represents the number of rows, col represents the number of columns, and width represents the data bit width.
[0028] In a feasible embodiment of the solution of this application, the expression of the resource constraint is as follows:
[0029]
[0030] where v d represents the partition space allocated to node v, v area represents the required resources of the node, r v represents the set of nodes accommodated in the current partition r, and (r child ) area represents the number of resources in each partition.
[0031] In a feasible embodiment of the solution of this application, the method of determining that the cost function is the minimum value or the resources in the partition reach the critical value of the resource constraint by performing split iteration on the partition result of the FPGA architecture, and outputting the layout constraint includes:
[0032] Obtain the first coordinates of the nodes in the control data flow graph before split iteration, determine the coordinate transformation relationship according to the split method, and transform the first coordinates according to the coordinate transformation relationship to obtain the second coordinates;
[0033] The split method includes horizontal split or vertical split.
[0034] In a feasible embodiment of the solution of this application, the expression of the coordinate transformation relationship is as follows:
[0035]
[0036]
[0037] Among them, v.row represents the row coordinate in the second coordinate, v.col represents the column coordinate in the second coordinate, and (v.row) prev represents the row coordinate in the first coordinate, and (v.col) prev represents the column coordinate in the first coordinate, and v d represents the partition space allocated to node v. Vertical partition represents horizontal division, and horizontal partition represents vertical division.
[0038] In a feasible embodiment of the solution of the present application, in the step of inserting a pipeline into the connection line between nodes in the control data flow graph for delay balancing, the expression of delay balancing is as follows:
[0039] e ij.balance =(S i -S j -e ij.lat )
[0040] Among them, S i represents the time step of node v i , S j represents the time step of node v j , and S i -S j represents the maximum delay between all paths between node v i and node v j ; e ij.lat represents the additional time delay existing before inserting the pipeline; e ij.balance represents the balanced time delay generated after inserting the pipeline.
[0041] In a feasible embodiment of the solution of the present application, the step of inserting a pipeline into the connection line between nodes in the control data flow graph for delay balancing includes:
[0042] Constructing an objective function of area overhead according to the balanced time delay, and the objective function is:
[0043]
[0044] Among them, e ij.width is the maximum data bit width between node v i and node v j .
[0045] The advantages and beneficial effects of the present invention will be partially given in the following description, and other parts can be understood through the specific implementation manners of the present invention:
[0046] The technical solution of this application proposes a full - process layout method for high - level synthesis - guided FPGA physical layout constraints based on a planar planning algorithm. The method constructs a control data - flow graph through the circuit description of the target circuit, and divides the control data - flow graph through the planar planning algorithm to obtain layout constraints. Based on the layout constraints, scheduling and binding of the corresponding resources of the target circuit are performed to obtain a register - transfer - level description, and further synthesis processing, layout, and routing are carried out to obtain the netlist of the target circuit. High - level synthesis reduces the congestion situation in layout and routing, and can also reduce the increased delay caused by crossing the FPGA block boundary during the routing process. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following - described drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0048] Figure 1 It is the flowchart of the steps of the high - level synthesis process layout method provided in the technical solution of this application;
[0049] Figure 2 It is the schematic diagram of the control data - flow graph in the technical solution of this application;
[0050] Figure 3 It is the schematic diagram of the iterative partitioning process in the technical solution of this application;
[0051] Figure 4 (a) is one of the schematic diagrams of balanced delay in the technical solution of this application;
[0052] Figure 4 (b) is the other schematic diagram of balanced delay in the technical solution of this application;
[0053] Figure 5 It is the schematic diagram of the FIFO pipeline in the technical solution of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] The following details the embodiments of the present invention. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention. For the step numbers in the following embodiments, they are only set for the convenience of explanation and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0055] In current related technical solutions, especially in FPGA placement and routing, physical synthesizers often use resources that are relatively close to each other, which can easily lead to congestion in placement and routing, increase the overall routing length, and thus reduce the throughput of the circuit. In view of the technical deficiencies in the related technical solutions, the technical solution of this application proposes a full-process placement method for guiding FPGA physical placement constraints based on a floorplanning algorithm in a high-level synthesis tool.
[0056] In a first aspect, as Figure 1 shown, the technical solution of this application provides a process placement method for high-level synthesis; the method includes steps S100 - S400:
[0057] S100. Obtain the circuit description of the target circuit, and construct a control data flow graph corresponding to the target circuit according to the circuit description;
[0058] Among them, as Figure 2 shown, the control data flow graph is a directed graph representing the arithmetic operations of the target circuit. In Figure 2 , the solid lines represent data dependency relationships, the dashed lines represent control dependency relationships, and the triangle symbols represent branch operations. In the embodiment, the circuit description includes, but is not limited to, the content describing the signal input, components, and logical operations performed by the components in the target circuit through VHDL language or Verilog language; the target circuit may refer to a hardware circuit in a real scenario. Specifically in the embodiment, first, the input circuit description is obtained, and the control data flow graph is constructed. In the embodiment, the control data flow graph (Control Data Flow Graph, CDFG) is a directed graph G = <V, E>, where V represents the set of all nodes in the control data flow graph, and each node in the control data flow graph represents an arithmetic operation in the target circuit; E represents the set of all directed connection lines in the control data flow graph, and each directed edge connecting two nodes in the control data flow graph represents the data or control dependency relationship existing between the two corresponding arithmetic operations. The control relationship dependency edges of the control data flow graph reflect the control dependency of the circuit description, and the data relationship dependency edges reflect the data dependency of the circuit description. Based on the properties of the control data flow graph, in the process of constructing the control data flow graph in the embodiment, first, the embodiment uses the compiler front-end to generate intermediate code from the high-level language code of the behavioral description in the circuit description according to the source code; then the embodiment uses the compiler back-end to map the variables in the circuit description to nodes, map the control and data dependencies to directed edges, and construct the control flow data flow graph.
[0059] S200. Divide the control data flow graph through a floorplanning algorithm to obtain placement constraints;
[0060] Among them, the planar planning algorithm in the embodiment can be solved and planned by the cutting plane algorithm; the layout in the embodiment can be defined as a set of physical constraints for controlling the placement of logic in the model. Specifically, in the embodiment, first, according to the FPGA architecture corresponding to the target circuit, the number of partitions, resources, and the maximum resource utilization rate corresponding to the target circuit are determined; then, the function corresponding to the data flow process in the control data flow graph is compiled into a Register Transfer Level (RTL) module and placed in the initial partition; based on the cutting plane algorithm, the current partition is horizontally or vertically divided into two, the scheme with the minimum cost function is calculated and selected, and based on the obtained scheme, the layout constraints corresponding to the target circuit in the FPGA architecture are determined.
[0061] In some feasible implementation manners, the step S200 of obtaining the layout constraints by dividing the control data flow graph through the planar planning algorithm may include steps S210 - S240:
[0062] S210. Obtain the FPGA architecture of the target circuit in the circuit description;
[0063] S220. Compile according to the function corresponding to the data flow process in the control data flow graph to obtain a circuit module at the register transfer level;
[0064] S230. Partition the FPGA architecture according to the circuit module, and determine the cost function of the partition result;
[0065] S240. Determine that the cost function is the minimum value or the resources in the partition reach the critical value of the resource constraint by performing split iteration on the partition result of the FPGA architecture, and output the obtained layout constraints;
[0066] Specifically, in the embodiments, there are multiple Blocks in the FPGA architecture, and the resources of different Blocks may vary. Good placement helps reduce routing congestion and improve the quality of the timing results (QoR) achievable in the design; in the embodiments, placement constraints use the Pblock instruction to specify resource partitioning. The Pblock boundary allows the use of the clock region boundary to define the size of the pblock, rather than using ranges such as SLICE, BRAM, DSP, etc., which helps limit clock skew and contributes to the overall clock placement of the design. And based on the control data flow graph, the basis for partitioning the HLS design in the embodiments is the data flow programming style, that is, the HLS design is streaming, and the design structure is described as a directed graph; in the directed graph, nodes represent units that need to perform arithmetic processing, and the connection lines between nodes describe the data transmission paths; in the directed graph, adjacent nodes transmit data through the connection lines, the nodes consume data for calculation, and the generated data is output to the input / output sequence as the input of the next calculation unit.
[0067] Exemplarily, in the example, the HLS design adopts a data flow programming model, where each function corresponds to a data flow process, each function corresponds to an RTL module, and FIFOs are used for communication between the modules. Then, a graph G = <V, E> is constructed, where V represents the set of data flows, and each node represents a function; E represents the set of FIFO channels between the vertices.
[0068] S300. Schedule the control data flow graph, and bind the scheduling result to obtain a register transfer level description.
[0069] Specifically, in the embodiments, corresponding scheduling needs to be performed on the subgraph of the control data flow graph; perform latency balancing on the control data flow graph; furthermore, in some feasible implementation schemes, step S300 may include steps S310 - S330:
[0070] S310. Schedule the subgraph of the control data flow graph.
[0071] S320. Insert pipelines into the connection lines between the nodes in the control data flow graph for latency balancing.
[0072] S330. Mathematically integrate the scheduling result and the result after latency balancing, and bind according to the result after mathematical integration with the target circuit to obtain the register transfer level description.
[0073] Specifically in the embodiment, the sub-graphs of the control data flow graph are scheduled in the default manner of the high-level synthesis tool; then, pipelines are inserted into the cut edges of the control data flow graph to balance the latency; after mathematically integrating the scheduling results obtained in steps S310 - S320, a comprehensive scheduling result is obtained; the resources in the FPGA architecture corresponding to the target circuit are bound to the comprehensive scheduling result to obtain a register transfer level description.
[0074] S400. Obtain a target netlist according to the layout constraints for the register transfer level description, and determine the process layout of the target circuit according to the target netlist.
[0075] Specifically in the embodiment, according to the layout constraints obtained in step S200 and the register transfer level description obtained in step S300; perform integration processing and placement and routing operations on the resources in the FPGA architecture to obtain a corresponding target netlist, thereby determining the control process layout corresponding to the target circuit.
[0076] In some feasible implementation manners, the layout constraints in the embodiment include at least one of timing constraints or physical constraints; the target netlist includes at least one of a synthesis netlist or a placement and routing netlist; thus, step S400 of obtaining a target netlist according to the layout constraints for the register transfer level description and determining the process layout of the target circuit according to the target netlist may include steps S410 - S450:
[0077] S410. Construct a first input according to the timing constraints and / or the physical constraints.
[0078] S420. Construct a second input according to the register transfer level description.
[0079] S430. Through the FPGA physical synthesizer, integrate and process the first input and the second input and output a synthesis netlist.
[0080] S440. Through the FPGA physical synthesizer, perform placement and routing processing on the first input and the second input and output a placement and routing netlist.
[0081] S450. Determine the process layout of the target circuit according to the synthesis netlist and the placement and routing netlist.
[0082] Specifically in the embodiment, first, the layout constraints obtained in step S200 and the timing constraints and physical constraints obtained by the high-level synthesis tool itself are input as a set of constraint conditions for synthesis implementation, which is the first input; the register transfer level description obtained in step S300 is used as the RTL input for synthesis implementation, which is the second input; then the embodiment runs the FPGA physical synthesizer to perform synthesis, placement and routing operations, obtains the synthesized netlist and the netlist after placement and routing, and determines the flow layout of the target circuit according to the obtained netlist.
[0083] In the embodiment, in step S230, the FPGA architecture is partitioned according to the circuit module, and the cost function of the partitioning result is determined, where the physical meaning of the cost function is the sum of the number of wires passing through the partition boundary. Furthermore, the cost function in the embodiment is:
[0084]
[0085] where C is the cost value, v i and v j represent nodes in the control data flow graph, i = 1, 2, 3,... n, j = 1, 2, 3,... n, n is a positive integer, E represents the set of FIFO channels between nodes, and e ij is the connection line between v i and v j ; row represents the number of rows, col represents the number of columns, and width represents the data bit width.
[0086] In the embodiment, in step S240, by performing split iteration on the partitioning result of the FPGA architecture, it is determined that the cost function is the minimum value or the resources in the partition reach the critical value of the resource constraint, and the layout constraint is output; the expression of the resource constraint is as follows:
[0087]
[0088] where v d represents the partition space allocated to node v, v area represents the required resources of the node, r v represents the set of nodes accommodated in the current partition r, and (r child ) area represents the number of resources in each partition.
[0089] In the embodiment, the step S240 of performing split iteration on the partitioning result of the FPGA architecture to determine that the cost function is the minimum value or the resources in the partition reach the critical value of the resource constraint and output the layout constraint may include steps S241 - S242:
[0090] S241. Obtain the first coordinates of the nodes in the control data flow graph before the splitting iteration, determine the coordinate transformation relationship according to the splitting method, and transform the first coordinates according to the coordinate transformation relationship to obtain the second coordinates;
[0091] S242. The splitting method includes horizontal splitting or vertical splitting;
[0092] Specifically, in the embodiment, as Figure 3 shown, the process of partitioning can be regarded as iteratively splitting into two each time until the cost function is minimized or the constraint conditions are no longer satisfied, and finally adding a pipeline FIFO; the first step is to map all functions into RTL modules and place them in a partition, called the initialization partition, where the dependency relationships are 1 pointing to 2, 3, 4, 2, 3, 4 pointing to 5, and the resource occupancy of 2 and 3 is relatively small; the second step is to vertically split into two, with 123 on the upper side and 45 on the lower side; the third step is to horizontally split each partition into two, and the result is that 2 and 3 are in the upper left, 1 is in the upper right, 4 is in the lower left, and 5 is in the lower right; the last step is to add a FIFO pipeline for the traces crossing the block boundary to ensure the throughput of the circuit design.
[0093] Further, in the embodiment, the expression of the coordinate transformation relationship is as follows:
[0094]
[0095]
[0096] where v.row represents the row coordinate in the second coordinates, v.col represents the column coordinate in the second coordinates, (v.row) prev represents the row coordinate in the first coordinates, (v.col) prev represents the column coordinate in the first coordinates, v d represents the partition space allocated to the node v, vertical partition represents horizontal splitting, and horizontal partition represents vertical splitting.
[0097] In the embodiment, in step S320, pipelines are inserted into the connection lines between the nodes in the control data flow graph for delay balancing.
[0098] Specifically, in the embodiment, given a partitioned and pipelined data flow graph G<V, E>, each vertex v ∈ V represents a function in the data flow design, each edge e ∈ E represents a FIFO channel between functions, the width e.width represents the bit width of the edge, the delay e.lat represents the additional delay inserted in the previous pipeline step, and the balance delay e.balance represents the balance delay in the current step. For each edge e ∈ E, the total delay of each path can be expressed as:
[0099]
[0100] where {p1, p2} represents a pair of re-converging paths. Further, in delay balancing for each edge (connection line) e, it can be considered that S i ≥ S j + e ij.lat , and the additional balance delay can be expressed as:
[0101] e ij.balance =(S i - S j - e ij.lat )
[0102] where S i represents the time step of node v i , S j represents the time step of node v j , S i - S j represents the maximum delay between all paths between node v i and node v j ; e ij.lat represents the additional delay inserted in the longest path between vertex v i and v j in the previous pipeline step; e ij.balance represents the additional delay inserted in the longest path between vertex v i and v j in the current pipeline step.
[0103] Exemplarily, as Figure 4 shown, Figure 4 (a) represents the cut-set1 cut set during the balance delay process; Figure 4(b) shows cut - set2 and cut - set3 during the balanced delay process. Among them, edges e13, e37, and e27 are pipelined according to the planar graph partition, and then each edge carries 1 unit of insertion delay. At the same time, it is assumed that the bit - width of e14 is 2, and the bit - width of all other edges is 1. In the delay balancing step, the optimal solution is to increase the delay of each of the edges e47, e57, and e67 by 2 units, and increase the delay of e12 by 1 unit each. Note that e27 and e37 can exist in the same cut - set.
[0104] As Figure 5 shown, after the partition is divided, it is connected based on FIFO to enable pipelining. By using FIFO, the matching interface signals can be directly scheduled without affecting the function, and the parallelism of the circuit function is provided.
[0105] In some feasible embodiments, the step S320 of inserting a pipeline into the connection line between nodes in the control data - flow graph for delay balancing may further include step S321:
[0106] S321. Construct an objective function of area overhead according to the balanced time delay;
[0107] Specifically in the embodiment, the optimization objective of balanced delay is to minimize the total area overhead, and the bit - width overhead of each edge is considered. The objective function in the embodiment is:
[0108]
[0109] where e ij.width is the maximum data bit - width of the pipeline between node v i and node v j ;
[0110] From the above - mentioned specific implementation process, it can be summarized that the technical solution provided by the present invention has the following advantages or advantages compared with the prior art:
[0111] The present invention can be applied to the high - level synthesis design of medium - sized and large - sized data - flow programming models, giving play to the characteristics of high - level synthesis tools that can quickly prototype, and at the same time extending the design process from high - level language to hardware description language and then to physical layout. The full - process design method in which the high - level synthesis tool guides the physical layout can further improve the layout and routing situation of the high - level synthesis design, reduce layout congestion, and ensure the overall throughput of the circuit.
[0112] Exemplarily, for the RISCV CPU design, using the default HLS and the HLS with application layout constraints respectively, it can be clearly determined that the HLS design with application layout constraints performs partition pipelining on the CPU module, FFT module, USB1 module, and USB2 module with the highest resource occupancy rate, and distributes them in adjacent but different block blocks at the same time. Under the condition of meeting the timing constraints, the timing margin remains basically unchanged, but the congestion situation is greatly improved, and the running time of physical synthesis is also reduced.
[0113] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order mentioned in the operation diagrams. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated, where the order of various operations is changed and where sub-operations described as part of a larger operation are executed independently.
[0114] In addition, although the present invention has been described in the context of functional modules, it should be understood that unless otherwise stated to the contrary, one or more of the functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It can also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. More precisely, considering the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the module will be understood within the ordinary skills of an engineer. Therefore, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It can also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0115] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device).
[0116] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0117] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and purposes of the present invention, and the scope of the present invention is defined by the claims and their equivalents.
[0118] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.
Claims
1. High-level synthesis process layout method, characterized in that, Including the following steps: Obtain the circuit description of the target circuit, and construct a control data flow graph corresponding to the target circuit according to the circuit description; the control data flow graph is a directed graph representing the arithmetic operations of the target circuit; Segment the control data flow graph through a floorplanning algorithm to obtain layout constraints; Schedule the control data flow graph, bind the scheduling result to obtain a register transfer level description; obtain a target netlist according to the layout constraints for the register transfer level description, and determine the process layout of the target circuit according to the target netlist; The segmenting the control data flow graph through a floorplanning algorithm to obtain layout constraints includes: Obtain the FPGA architecture of the target circuit in the circuit description; Compile the functions corresponding to the data flow processes in the control data flow graph to obtain circuit modules at the register transfer level; Partition the FPGA architecture according to the circuit modules, and determine the cost function of the partitioning result; determine that the cost function is the minimum value or the resources in the partition reach the critical value of the resource constraint through iterative segmentation of the partitioning result of the FPGA architecture, and output to obtain the layout constraints; The cost function is used to characterize the number of wires at the partitioning boundary of the FPGA architecture; the cost function is: Where C is the cost value, v i and v j represent nodes in the control data flow graph, i = 1, 2, 3, … n, j = 1, 2, 3, … n, n is a positive integer, E represents the set of FIFO channels between nodes, e ij is the connection line between v i and v j The row represents the number of rows, the col represents the number of columns, and the width represents the data bit width.
2. The high-level synthesis process layout method according to claim 1, wherein The scheduling the control data flow graph, binding the scheduling result to obtain a register transfer level description includes: Schedule the subgraphs of the control data flow graph; Insert a pipeline into the connection lines between the nodes in the control data flow graph to balance the delays; Mathematically integrate the scheduling result and the result after delay balancing, and bind the result after mathematical integration to the target circuit to obtain the register transfer level description.
3. The high-level synthesis process layout method according to claim 1, characterized in that The layout constraints include at least one of timing constraints or physical constraints; the target netlist includes at least one of a synthesis netlist or a placement and routing netlist; the obtaining a target netlist according to the layout constraints for the register transfer level description, and determining the process layout of the target circuit according to the target netlist includes: Construct a first input according to the timing constraints and / or the physical constraints; Construct a second input according to the register transfer level description; Through an FPGA physical synthesizer, integrate and process the first input and the second input and output to obtain a synthesis netlist; Through an FPGA physical synthesizer, perform placement and routing processing according to the first input and the second input and output to obtain a placement and routing netlist; Determine the process layout of the target circuit according to the synthesis netlist and the placement and routing netlist.
4. The high-level synthesis process layout method according to claim 1, characterized in that The expression of the resource constraint is as follows: Among them, v d represents the partition space allocated to node v, and v area represents the required resources of the node, and r v represents the set of nodes accommodated by the current partition r, (r child ) area represents the amount of resources in each partition.
5. The high-level synthesis process layout method according to claim 1, characterized in that The determining that the cost function is the minimum value or the resources in the partition reach the critical value of the resource constraint through iterative segmentation of the partitioning result of the FPGA architecture, and output to obtain the layout constraints includes: Obtain the first coordinates of the nodes in the control data flow graph before iterative segmentation, determine the coordinate transformation relationship according to the segmentation method, and transform the first coordinates according to the coordinate transformation relationship to obtain second coordinates; The splitting method includes horizontal splitting or vertical splitting.
6. The high-level synthesis process layout method according to claim 5, characterized in that The expression of the coordinate transformation relationship is as follows: Among them, v.row represents the row coordinate in the second coordinate, v.col represents the column coordinate in the second coordinate, and (v.row) prev represents the row coordinate in the first coordinate, and (v.col) prev represents the column coordinate in the first coordinate, and v d represents the partition space allocated to node v. vertical partition represents splitting in the horizontal direction, and horizontal partition represents splitting in the vertical direction.
7. The high-level synthesis process layout method according to claim 2, characterized in that In the step of inserting a pipeline into the connection line between nodes in the control data flow graph for latency balancing, the expression of latency balancing is as follows: e ij.balance = (S i - S j - e ij.lat ) Among them, S i represents the time step of node v i , S j represents the time step of node v j , S i -S j represents the maximum delay among all paths between node v i and node v j ; e ij.lat represents the additional time delay existing before inserting into the pipeline; e ij.balance represents the balanced time delay generated after inserting into the pipeline.
8. The high-level synthesis process layout method according to claim 7, characterized in that, The step of inserting a pipeline into the connection line between nodes in the control data flow graph for latency balancing includes: Construct an objective function of area overhead according to the balanced time delay, and the objective function is: Among them, e ij.width is the maximum data bit width between the pipeline at node v i and node v j
Citation Information
Patent Citations
High-level synthesis method and system
CN102419789A
Placing and routing method for implementing back bias in FDSOI
CN107086218A