Method for planning partitions for a programmable gate array
By changing the partition boundary direction and optimizing resource matching in FPGA design, the problem of excessive construction time in the existing technology is solved, more efficient resource utilization and layout planning is achieved, and the development cycle is shortened.
Patent Information
- Application Number
- CN202110190480.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-18
- Filing Date
- 2021-02-18
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2041-02-18
AI Technical Summary
During the gradual development of existing FPGA design, as the program logic requirements increase and the construction time is too long, the existing floor plan needs to be completely rebuilt, resulting in the development process being time-consuming and inefficient.
By changing the partitioning planning method of programmable gate arrays, the partition boundary direction is changed, the logical components are converted, and resource matching is optimized using the weighted distance function and gradient matrix to achieve efficient allocation and layout planning of logical components.
Reduces the construction time of programmable gate arrays, especially during reconstruction, significantly shortens the development cycle, and improves resource utilization efficiency and layout planning efficiency.
Smart Images

Figure CN113343623B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a method for planning the design of partitions, each partition being for a programmable gate array and a plurality of program routines, the programmable gate array comprising different categories of logic elements at predetermined locations, the plurality of program routines comprising at least one first program routine and at least one further program routine.
[0002] The present invention also relates to a computer-based system and a computer program product for performing such a method. Background Art
[0003] A so-called field programmable gate array (FPGA) is an integrated circuit in the form of a programmable gate array that contains digital function blocks with logic elements that can be programmed and interconnected by the user. The corresponding program routines then predetermine the FPGA design, i.e., the design of the partitions for the corresponding program routines. The development of an FPGA design is also known as "floorplanning." One aspect of floorplanning involves planning the design of the corresponding partitions, given the number and location of different types of logic elements in the programmable gate array.
[0004] Typically, FPGA designs are developed incrementally. If the program logic's demands on the FPGA's resources increase during development, requiring a larger partition, the existing floor plan is often outdated, requiring a complete rebuild of the FPGA. This makes the development process very time-consuming.
[0005] An example of this step-by-step development of FPGA designs is the abstract FPGA development environment "FPGA Programming Blockset" and the corresponding hardware board from the applicant's dspace. These allow users to develop their own hardware even without detailed knowledge of FPGAs and tool flows. Based on the model, users can create and implement FPGA designs and run them on the corresponding hardware. Regular changes to the FPGA model and subsequent rebuilds are a natural part of the development process (for example, in the case of fast control models).
[0006] A problem with FPGA reconstruction ("rebuilding") is the very long build time ("build time"), which is typically in the range of several hours. As FPGAs become more complex and offer more resources (registers, logic, DSP blocks, memory), build times are increasing. This is only partially offset by better build algorithms (e.g., placement and routing) and faster computers. It is now common to rebuild the entire FPGA with minimal changes.
[0007] FPGA tool manufacturers offer solutions such as a hierarchical design flow to address the issue of long assembly times. In this hierarchical construction process, the entire model is divided into several sections, and these sections are then reserved for specific areas within the FPGA. This allows, for example, the construction of model components to be distributed across multiple computers in parallel and accelerated. Furthermore, already implemented components that remain unchanged during the construction can be reused. Patent US9870440B2 describes such a process.
[0008] An important aspect here is the creation of a floor plan that defines the areas of the submodels in the FPGA, which can then be built independently and in parallel. It should be ensured that sufficient logic element resources are available for each submodel. In addition to the actual logic elements (CLBs) as resources, there are also dedicated logic elements, such as RAM or DSP. However, these are usually present in significantly fewer numbers and are unevenly distributed on the FPGA. This results in a complex floorplanning problem, which, in addition to the resource requirements of the individual submodels, must also take into account the distances between modules and, therefore, the lengths of the communication lines between the submodels. Summary of the Invention
[0009] Therefore, based on the above-mentioned prior art, the object of the present invention is to provide a method for designing partitions of a programmable gate array and a system and a computer program product for executing the method, which can reduce the construction time of the programmable gate array, especially the construction time during reconstruction.
[0010] This object is achieved according to the invention by the features of the independent claim. Advantageous embodiments of the invention are given in the dependent claims.
[0011] The invention relates to a method for planning the design of partitions, each partition being for a programmable gate array and a plurality of program routines, the programmable gate array comprising logic elements of different categories at predetermined locations, the plurality of program routines comprising at least one first program routine and at least one further program routine, the method comprising the following method steps:
[0012] (i) providing for assigning a first partition of the programmable gate array to a first program routine and for assigning at least one further partition of the programmable gate array to the at least one further program routine, wherein the partitions are separated from one another by current partition boundaries,
[0013] (ii) determining the requirements of the first program routine for various classes of logic elements,
[0014] (iii) matching the demand with resources of corresponding classes of logical elements present in the first partition, and
[0015] (iv) If the first program routine's demand for logic elements of this class exceeds the corresponding resources of the first partition, at least one logic element of the corresponding class is selected from the further partition or at least one of the further partitions and transferred to the first partition by changing the course of the partition boundary between the partitions, wherein the partition boundary runs in a curved manner in at least one section before and / or after its course is changed. This type of conversion of logic elements by changing the course of the partition boundary creates new planning possibilities (new floorplans NFP). To utilize these possibilities particularly effectively, a judicious selection of the logic elements to be converted and corresponding options for changing the course of the corresponding partition boundaries are required.
[0016] According to the present invention, the partition boundary is simply moved without maintaining its previous, in particular linear, course. By switching at least one logic element, the partition boundary has, at least in some areas, a wavy course. This curved course is formed by bends and is not a straight line. At least one bend forms a projection of the first partition into the further partition or one of the further partitions. The switched logic element or at least one of the switched logic elements is then located in this projection. Advantageously, by switching at least one logic element, the partition boundary has a less linear course than before the switching.
[0017] The method is in particular a computer-implemented method.
[0018] In particular, the method provides for the selection of the logic element to be switched to depend on a weighted distance function of the distance of the respective logic element from the current course of the partition boundary. Such a weighted distance function is implemented, in particular, by a gradient matrix.
[0019] According to a preferred embodiment of the present invention, it is provided that when switching the at least one corresponding logical element from the further partition to the first partition, preference is given to selecting a logical element which is located near the current course of the partition boundary of the first partition.
[0020] According to a preferred embodiment of the present invention, logic elements are successively transferred from a boundary region of the further partition or of at least one of the further partitions, directly adjacent to the partition boundary, into the first partition until the need for additional logic elements in the first partition is met.
[0021] In this case, provision is preferably made for, in each individual transformation, to preferentially select a boundary region in which at least two of the required logic elements are present with the highest possible spatial density.
[0022] The selection / identification of the logical element to be subsequently switched on the plane of the further subregion is preferably carried out with the aid of a two-dimensional selection function, which is created in the following manner:
[0023] (a) Creating a distribution function for each class of desired logical elements on the plane of the further partition. The distribution function is an additive superposition of delta distributions, each delta distribution having a delta peak at each point on the plane of the further partition at which a logical element of the corresponding class being sought is stored.
[0024] (b) Create a density function for each class of desired logical elements by convolving the distribution function with the distance function.
[0025] The distance function is characterized by having a global maximum at the reference point (0, 0) and decreasing strictly monotonically in all directions as the distance from (0, 0) increases. A suitable distance function is, for example, a Gaussian distribution. As a result of the convolution operation, the density function is a superposition of many distance functions (e.g., many Gaussian normal distributions), whose maximum values coincide with the locations of the logical elements of the respective class being sought. A large value of the density function at a specific point indicates that a high density of the logical elements being sought exists at that point, i.e., at least one of the logical elements being sought is located nearby.
[0026] (c) A gradient function is created on the plane of the partition to be stolen, said gradient function having its maximum value at the current partition boundary and decreasing strictly monotonically in the direction of the opposite boundary.
[0027] (d) Create the selection function by multiplying all created density functions with the gradient function.
[0028] The selection function is represented in particular as a matrix, wherein each matrix element of the matrix represents a logical element on a partition, and the value of each entry represents the value of the selection function at the corresponding point. Alternatively, the conversion can be performed at the end, or all functions can be represented as matrices from the beginning and the selection function can be represented as a matrix by element-by-element multiplication (Hadamard product).
[0029] Mathematical convolution is a function that expresses the degree to which two functions are superimposed, depending on their relative offset. (See, for example, the Wikipedia entry for "Convolution (mathematics).")
[0030] Advantageously, provision is made for the allocation between the individual partitions and the program routines in step (i) to be provided by dividing the plane of the programmable gate array into these partitions.
[0031] In a preferred embodiment of the present invention, a third category of logic assemblies and a third resource allocation function are provided. The overall resource allocation function is the product of the gradient function, the first, second and third resource allocation functions.
[0032] According to yet another preferred embodiment of the present invention, it is provided that at least one of the different categories of logic elements is a category from the following list of categories:
[0033] CLB category (CLB: Configurable Logic Block),
[0034] DSP category (DSP: Digital Signal Processor), and
[0035] RAM category (RAM: Random Access Memory).
[0036] These categories of logic elements are very typical for FPGAs.
[0037] Finally, with regard to the method, it is advantageously provided that, for planning the design of the further partitions, the method described above is followed, wherein one of the further partitions is considered the first partition and the assigned program routine is considered the first program routine. In this way, the design of all partitions is carried out step by step.
[0038] The present invention also relates to a method for planning a programmable gate array, the programmable gate array comprising different types of logic elements at predetermined locations, the programmable gate array being used for a plurality of program routines, the plurality of program routines comprising at least one first program routine and at least one further program routine, wherein each partition is designed for a program routine. It is provided that the design of each partition is carried out with the aid of the above-mentioned method for planning the design of partitions for a programmable gate array. The method may also be referred to as a floorplanning method and is in particular a computer-aided method. In addition to the design of the partitions, the method also comprises planning the connections of the logic elements.
[0039] In other words, the method describes the application of the above-described method for planning a partitioned design to planning an entire programmable gate array comprising different categories of logic elements at predetermined locations.
[0040] In this case, it is preferably provided that the method is a step-by-step method.
[0041] With regard to a computer-based system having a processor unit, it is provided that the system is designed to carry out the above-described method.
[0042] With regard to the computer program product, it is provided that the computer program product comprises a program part which is loaded into a processor of a computer-based system and is provided for carrying out the above-described method. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The present invention will be described in more detail below with reference to the accompanying drawings using preferred embodiments.
[0044] Figure 1 shows a sub-plane of a programmable gate array constructed as an FPGA,
[0045] Figure 2 shows a flow chart of a method for planning a design of a partition for a programmable gate array according to a preferred embodiment,
[0046] Figure 3 shows a schematic flow chart of the floorplanning method,
[0047] Figure 4 Shows the programmable gate array designed as FPGA, divided into logic element categories,
[0048] Figure 5 Schematic diagram showing 2D convolution of RAM resources with Gaussian filter,
[0049] Figure 6 FIG. 1 shows a portion of a programmable gate array designed as an FPGA.
[0050] Figure 7 An example gradient matrix is shown on the left; in the middle, a horizontal subdivision of the upper half of the FPGA is shown, where the resource matrix is calculated based on the CLBs, RAMs, and gradient matrices; and on the right, resources converted from further partitions are shown.
[0051] Figure 8 Graphical illustration of a series of experiments showing the placement efficiency of the new floorplanning method,
[0052] Figure 9 Show that the found planar graph is decomposed into rectangles, and
[0053] Figure 10 A programmable gate array configured as an FPGA with a wrapper partition and an FPGA model partition is shown, to which a floorplan for a modular construction process can be applied. DETAILED DESCRIPTION
[0054] Figure 1The left and right sides show sub-planes of a programmable gate array 10 configured as an FPGA (Field Programmable Gate Array). Logic elements 12 are arranged on the illustrated sub-planes according to a predetermined structure. In this example, the logic elements shown here are CLBs (Configurable Logic Blocks). CLBs are one of several types of logic elements 12 commonly used in programmable gate arrays 10.
[0055] If the programmable gate array 10 is now used for a plurality of program routines, it must be divided into different partitions 14, 16. The following relates to a first partition 14 and at least one further partition 16. The division is carried out dynamically and is therefore changeable. By means of the division, corresponding partition boundaries 18 are produced between the first partition and the further partitions 14, 16. The change in the division into partitions 14, 16 is achieved by changing the course of the partition boundaries 18 between these partitions 14, 16. Figure 1 In FIG. 1 , the current course of the partition boundary 18 before a corresponding change in its boundary is shown on the left, while the course of the partition boundary 18 after a corresponding change in the partition boundary 18 is shown on the right.
[0056] Figure 2 Shown is a flow chart of a method for planning a design of a partition for a programmable gate array 10. The programmable gate array 10 has a plurality of different categories of logic elements 12 at predetermined locations.
[0057] Now, the programmable gate array 10 should be designed for processing at least one first program routine and at least one further program routine. To this end, the following steps are performed:
[0058] Step S1: Assigning the first partition 14 to the first program routine and assigning the at least one further partition 16 to the at least one further program routine. For this purpose, the gate array 10 is usually divided into consecutive spatial partitions 14, 16, the number of which is equal to the number of program routines.
[0059] Step S2: Determine the requirements of the first program routine for the various categories of logic elements 12 .
[0060] Step S3 : Matching the demand with the resources of the corresponding category of the logic elements 12 existing in the first partition 14 .
[0061] Step S4: If query A determines that the first program routine's demand for logic elements 12 of this category exceeds the corresponding resources of the logic elements 12 of the first partition 14, then at least one logic element 12 of the corresponding category is converted from the further partition 16 or at least one of the further partitions to the first partition 14 by moving the partition boundary between these partitions 14, 16.
[0062] In the following, the approach will now be discussed in the example of a floor plan for an FPGA sub-area, wherein what has been said also applies to any other type of programmable gate array 10 .
[0063] The basic approach to achieving speed advantages by subdividing (user) models and hierarchical design processes is described in the aforementioned patent US 9870440 B2, so it will not be repeated here. The core task of the present invention is to find a suitable floor plan in the FPGA so that as many submodels (program routines) as possible can be implemented in parallel. Therefore, there are a certain number of submodels / modules or program routines and FPGAs or FPGA sub-planes. The result should be a floor plan that satisfies the following criteria as optimally as possible for each submodel:
[0064] ●The required logic cells / CLBs of the sub-model must be provided;
[0065] ●The required memory requirements / RAM of the sub-model must be provided.
[0066] ●The required DSP requirements of the sub-model must be provided;
[0067] ● Resource waste on logic elements 12 (CLBs, RAMs, DSP) should be as small as possible;
[0068] The shape of the FPGA area should be as uniform as possible (e.g., square or slightly rectangular, no extreme shapes) so that the line lengths within the sub-model, and therefore the signal transit times and maximum clock frequency, are not unnecessarily negatively affected; and
[0069] ●The number and length of lines between modules must be kept as small as possible.
[0070] A classic approach for creating floor plans or for placing gates in VLSI circuits (Very Large Scale Integration) is dual partitioning. To this end, the available plane is divided into two halves, and the modules are distributed across the two planes, for example, in all possible combinations, and the results are then evaluated using an evaluation function. The evaluation function can reflect different criteria for acceptable distribution of modules, such as the minimum number of communication lines crossing the partition boundaries or the distribution of resources across the partitions as evenly as possible. The plane is then recursively divided into two halves in each of the two halves until only one module remains in each partition.
[0071] exist Figure 3 The following shows a rough outline of the floorplanning process, i.e., the planning of a programmable gate array for a plurality of program routines. As a first step, the size of the FPGA and the sum of all program routines or the sum of the corresponding modules for these program routines are analyzed in order to find out whether the submodel / module corresponding to the program routine is also suitable for the FPGA in terms of resources (logic elements: CLBs, RAM, DSP). At this point, bottlenecks are also directly analyzed, which can then flow into the evaluation function of the partitioning / placement in the form of variable weights and into the conversion ("stealing") of resources from other partitions 16. For example, if the sum of all RAM and DSP requirements of the submodels / modules is close to the total size of the FPGA, then a uniform distribution of RAM and DSP in the case of subdivision / double partitioning should be evaluated as higher than a uniform distribution of possibly abundant CLBs.
[0072] If the required module resources are substantially less than the FPGA's module resources, the FPGA plane can be partitioned for the first time. In this case, a horizontal split is first performed, and then all n modules are distributed across the two partitions in all 2n-1-1 combinations, and the results are evaluated (see Figure 3 ). -1 is present in this formula because the case where all modules are in one partition is not a meaningful result. For all m∈{1…(n-1)}, the combination results from m modules in partition 1 and nm modules in partition 2. The horizontal partition is again canceled, and the evaluation is repeated for the vertical partition through the plane.
[0073] Especially with many modules, the computation time increases significantly due to the large number of possible module combinations per partitioning step. Various methods exist to reduce this (e.g., heuristic methods such as Nonlinear Integer Programming (NLP)), so they will not be described here in detail. In the present invention, we always start with a small number of modules, corresponding, for example, to the number of subsystems in the top-level client model, so that all combinations can be tried. This takes only a few seconds.
[0074] After the FPGA has been split in the middle both horizontally and vertically, it is checked which partitioning (horizontal or vertical) and which division (division) of the module has been evaluated best. For this purpose, an evaluation function / suitability function is provided, which takes into account the above-mentioned criteria (uniform CLB distribution, etc.). The best division is indicated at the end of the first program part (see "Remembering the most promising division"). If the correct division has been found, the two resulting divisions are recursively subdivided according to the above-mentioned scheme until only one module remains in each division 14, 16 and the floor plan plane for this module is thus found.
[0075] However, in many cases a simple split in the middle will not work, so the split boundary may be moved until a given module can be placed (see Figure 3 However, in many cases, even moving the boundaries will not allow placement. This is simply due to the fact that while a given number of modules might fit within a given FPGA plane purely computationally, it might not be practical to place them due to the linear horizontal / vertical partitioning shown above.
[0076] Already mentioned Figure 1 This example illustrates two modules placed in a 5x4 = 20 CLB FPGA sub-plane. One module requires 18 CLBs, and the other requires 1 CLB. Computationally, these two modules also fit within the FPGA sub-plane. However, linear segmentation does not allow for subdivision. Figure 1 The left side of the diagram shows the best possible partitioning (4 CLBs versus 16 CLBs). However, if some CLBs can be "stealed," or converted, then subdivision is possible (as shown on the right). This is a very simple example; it becomes more complex if, in addition to the CLBs, unevenly distributed RAM and DSP must be considered. Especially when FPGAs are heavily used, simple dual partitioning and dual partitioning with boundary shifting often do not provide a complete FPGA floorplan.
[0077] Therefore, in Figure 3 In the flowchart, starting from "reaching the maximum boundary movement", there is a Figure 2 14 , 16 ). In this branch, the purely computational number of logic elements 18 of different categories (here: CLBs, RAMs and DSPs) of the FPGA plane is greater than or equal to the sum of the module resources, but nonetheless cannot be placed (not even using boundary shifting). In this case, some resources must be converted ("stolen") from adjacent partitions 16. In order to keep the overlap of adjacent partitions 14, 16 as small as possible, it is important to convert logic elements 18 of the corresponding categories that are close to the boundary. If multiple categories of logic elements 12 are missing, such as CLBs and RAMs, these should also be as close together as possible so that the routing resources of the partition do not have to be used by another partition at multiple locations. For example, it must be avoided that a CLB is converted from one corner and a RAM from a completely different corner.
[0078] For data processing, the resources are modeled as separate layers in the form of a matrix. Figure 4 On the right side of FIG, three resource layers are shown by way of example.
[0079] Next, the critical resource matrix is 2D-convolved with a filter 20 (eg, a Gaussian filter) so that each resource logic element obtains a larger influence area in its matrix. Figure 5 This is shown schematically. The distribution of available logic elements 22 (e.g. available FPGA RAM) and the Gaussian filter 20 is converted into a corresponding distribution of 2D filtered logic elements 22 by a 2D convolution 24. In order to give a higher weight to the resources close to the border of the further partition 16, an additional matrix is generated which implements a weighting that decreases starting from the border 18 - see Figure 6 and the gradient (arrow 28). Then, all the required 2D filtered resource matrices C CLB , R RAM , D DSP and the gradient matrix G settle with each other (see Figure 7 The left part of the partition from which the resource is to be "stolen" (in Figure 7 The calculation of the resulting matrix S (shown as a further partition 16 in FIG) is carried out here by element-by-element multiplication of the matrix (Hadamard product) including scalar weighting according to the following formula: wherein the scalars α, β, γ are each 0 (no resources required) or 1 (resources required).
[0080]
[0081] After the 2D filtered resource matrix and the gradient matrix are settled with each other, only the required resources with the highest evaluation values need to be gradually allocated to other partitions. In this way, a planar graph that cannot be achieved with a simple two-partition can be created.
[0082] The efficiency of the method has been checked in a test series of 1500 placements (FPGA fill levels of 10%, 20% ... 90% to 100% for each of the three strategies). A fill level of 50% means that the sum of the module CLBs occupies 50% of the FPGA. The same applies to RAM and DSP. The results for the three placement strategies are given in Figure 8 . It can be clearly seen that the pure bipartitioning strategy (BP) already fails significantly at a fill level of 30%. The bipartitioning strategy with boundary shifting (BP+) can still consistently place all modules until the FPGA fill level reaches approximately 40%, but then fails. Only the new approach (NFP: New Floorplan) can reach 100%.
[0083] In the following, the conversion of matrix-based planar graphs into FPGA representation is discussed.
[0084] As a result, the proposed algorithm provides a floor plan of the FPGA in the form of a matrix. Here, each element of the matrix has the value of the module number assigned to its module. Thus, the partitions 14, 16 for the individual modules can be easily identified (also visually) (see right). Figure 7 ).
[0085] In FPGA tools, the partitioning of FPGA planes is usually implemented in different encodings. Rectangular partitions are usually placed together; in the case of Xilinx FPGAs, these partitions are called Pblocks. Since the number of Pblocks affects the performance when converting the floor plan, the goal here is to keep the number of Pblocks as small as possible. For this reason, the conversion of matrix partitions into Pblock partitions is based on the "maximum rectangle algorithm" (D. Vandevoorde), which finds the largest possible rectangle in each plane. The algorithm described in this message is iteratively applied to this transformation until the entire plane of partitions 14, 16 is represented by rectangles (Pblocks).
[0086] Figure 9 The plan view is shown broken down into rectangles. In this case, multiple rectangles can belong to one partition 14 , 16 .
[0087] Because the FPGA model development process undergoes multiple iterations, as described in the aforementioned US Patent No. 9,870,440 B2, the subsequent build of the FPGA application first analyzes whether the existing floor plan is still usable for the potentially changed module size or number of modules. It is desirable to reuse the existing floor plan as much as possible because, if the module partitioning remains unchanged, unchanged modules do not need to be rebuilt. This reduces build time ("setup time").
[0088] To prevent the creation of new floor plans, the following measures can be taken in advance:
[0089] If there is sufficient reserve in the FPGA, partitions of size p can be reserved in the initial floorplan as placeholders for partitions of future modules, where p∈{0...(FPGA resources-resources of all modules)}.
[0090] Before the initial floorplanning, the required resources of all modules are multiplied by a factor r so that the resulting partition sets the reserve for all modules, where r∈{1...FPGA resources / resources of all modules}.
[0091] Then, you can take the following actions:
[0092] - Install new modules
[0093] o Module partitions that are no longer in use (if adjacent to each other) can be merged first and then used for new modules when there are no more reserve partitions.
[0094] - Expand existing partitions
[0095] The initial floorplanning algorithm can reserve partitions for this purpose and only execute the steps of moving partition boundaries and "stealing" resources in modules whose resource requirements exceed the resources of their partition. In this process, resources from other partitions are occupied (beansprechen) with the following priority:
[0096] ■Neighborhood partitions no longer used
[0097] ■ The adjacent partition with the largest reserve or the least rebuild time (but only if multiple modules do not have to be rebuilt)
[0098] ■ Adjacent partitions of adjacent partitions can still be well connected by wiring.
[0099] ■ Here, as a first step, the actually used resources of the modules that have been placed in the partition from which the resources are to be taken can be masked in the resource matrix (CLB, RAM, DSP).
[0100] ● As a result, only the resources that are actually still available are occupied by the floorplanning method described. Modules that are reduced in size in this way do not have to be restructured.
[0101] If this does not produce any results, the masking can be omitted in the second step. This way, more potential resources are available at the border of the occupied partition. However, the modules of the occupied resources must be rebuilt in their reduced partition.
[0102] If moving partition boundaries and "stealing" resources does not result in success, a completely new floor plan can be created.
[0103] ■ The overall resource situation can thereby be reassessed and the introduced weighting function can lead to a better overall picture for the new situation.
[0104] Using Floorplans in Sub-Regions of an FPGA
[0105] The described floorplanning and updating of the floorplan can be transferred 1:1 from the planning of the entire FPGA plane to the sub-areas. This is important because usually not the entire plane of the FPGA is available for free customer modeling. In the case of the dSPACE FPGA Programming Blockset, the customer models are embedded, for example, in a VHDL wrapper that provides a fixed interface between the customer models on the one hand and the processor bus and dSPACE I / O channels on the other hand. In this case, the wrapper takes over the complete connection of the customer model to the outside world and must therefore be placed in an FPGA area where there are I / O drivers for the FPGA pins (see Figure 10 ). In general, multiple wrappers can also be considered in an FPGA. To reduce the waste caused by wrappers, when planning the layout of the Pblock, the resources occupied by the wrapper and the components of the used "component stack" can be marked as occupied.
[0106] In the following, the important elements of the present invention will be summarized in bullet points:
[0107] Dynamically adapt the weighting function to the current bottleneck. For example, RAM is critical because the sum of all module requirements is approaching the limit of FPGA RAM, while CLBs and DSPs are relaxed, so the weights are adapted accordingly to pay more attention to scarce resources when evaluating / selecting the placement.
[0108] ● The method whereby resources will be expanded to other regions
[0109] ○ Simple dual partitioning usually does not work because it keeps too many resources unused
[0110] ○ Identify missing resources (such as CLB and RAM)
[0111] o The resource matrix is convolved with a filter function (eg, a Gaussian filter 20).
[0112] ○Create gradient matrix
[0113] ○ The filtered resources are settled with the gradient matrix
[0114] o "Steal" some resources so that the partitions 14, 16 are changed as "least invasively" as possible, in such a way that the most highly valued resources are "stealed" first.
[0115] ●A method for updating a floor plan in which an area is expanded into another area without rebuilding the modules of the area that was shrunk.
[0116] • Combining floorplanning with additional wrappers in the FPGA and a “component stacking” approach by masking resources in such occupied areas in the resource matrix for floorplanning.
[0117] Reference Signs List
[0118] Programmable Gate Array 10
[0119] Logic element 12
[0120] Division, first 14
[0121] Division 2, 16
[0122] Partition boundaries, currently 18
[0123] Gaussian filter 20
[0124] Available FPGA RAM 22
[0125] 2D Convolution 24
[0126] 2D filtered RAM 26
[0127] Arrow (Gradient) 28
[0128] Query A
[0129] Dual-zone BP
[0130] Expanded dual-zone BP+
[0131] New Layout Plan (NFP)
[0132] Step 1 S1
[0133] Step 2 S2
[0134] Step 3 S3
[0135] Step 4 S4
Claims
1. A method for planning the design of partitions (14, 16), each partition being for a programmable gate array (10) and a plurality of program routines, the programmable gate array comprising logic elements (12) of different categories at predetermined locations, the plurality of program routines comprising at least one first program routine and at least one further program routine, the method comprising the following method steps: Provision is made for allocating a first partition (14) of a programmable gate array (10) to a first program routine and for allocating at least one further partition (16) of the programmable gate array to the at least one further program routine, wherein These partitions are separated from each other (S1) by the current partition boundaries (18), determining the requirements of the first program routine for the respective classes of logic elements (12) (S2), matching (S3) the demand with the resources present in the first partition (14) of the corresponding class of logical elements (12), and If the demand for logic elements (12) of the class by the first program routine exceeds the corresponding resources of the first partition (14), at least one logic element (12) of the corresponding class is selected from the further partition (16) or at least one of the further partitions and transferred to the first partition (14) by changing the course of the partition boundary (18) between these partitions (S4), The partition boundary (18) extends in a curved manner in at least one section before and / or after its course is changed, The method is characterized in that some logic elements (12) are successively transferred from a boundary region of the further partition (16) or of at least one of the further partitions, which region is directly adjacent to a partition boundary (18), into the first partition (14) until the need for additional logic elements (12) in the first partition (14) is met. In each individual conversion, a boundary region is preferably selected in which at least two of the required logic elements (12) are present with the highest possible spatial density.
2. The method according to claim 1, characterized in that When transferring the at least one corresponding logic element (12) from the further partition (16) into the first partition (14), preference is given to selecting a logic element (12) which is closer to a current partition boundary (18) between the partitions than to at least one other logic element (12) of the corresponding class.
3. The method according to claim 1 or 2, characterized in that The selection of the logic element (12) to be switched depends on a weighted distance function of the distance of the respective logic element from the current course of the partition boundary (18).
4. The method according to claim 3, characterized in that The weighted distance function is implemented by a gradient matrix.
5. The method according to claim 1 or 2, characterized in that In order to select the logical element (12) to be switched next, a two-dimensional selection function is created with the help of the following steps: - creating a distribution function of the distribution of logic elements as a superposition of delta functions having a delta peak wherever the corresponding logic element (14) is present; - creating a density function by convolving a distribution function with a distance function, wherein the distance function is characterized in that it has a global maximum at its reference point and decreases strictly monotonically in every direction as the distance from the reference point increases; - creating a gradient function on the plane of the further partition (16), which gradient function has its maximum at the current partition boundary (18) and decreases strictly monotonically towards the center of the further partition (16), and - Create a selection function by multiplying all created density functions and gradient functions.
6. The method according to claim 5, characterized in that The distance function is a Gaussian function.
7. The method according to claim 1 or 2, characterized in that At least one of the different categories of logic elements is a category from the following list of categories: Configurable logic block categories, Digital Signal Processor category, and Random Access Memory Class.
8. The method according to claim 1 or 2, characterized in that For planning the configuration of the further partitions (16), one of the further partitions (16) is considered as the first partition (14) and the assigned program routine is considered as the first program routine.
9. A method for programming a programmable gate array (10), said programmable gate array comprising different categories of logic elements at predetermined locations, said programmable gate array being used for a plurality of program routines, said plurality of program routines comprising at least one first program routine and at least one further program routine, wherein: The partitions (14, 16) are designed for program routines, characterized in that the partitions (14, 16) are designed by means of a method according to any one of claims 1 to 8.
10. The method according to claim 9, characterized in that The method described is a step-by-step method.
11. A computer-based system having a processor unit, wherein: The system is configured to perform the method according to any one of claims 1 to 10 .
12. A computer program product comprising a program portion arranged to be loaded into a processor of a computer-based system for performing the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Method for automatically generating a netlist of an FPGA program
US9870440B2
FPGA having a virtual array of logic tiles, and method of configuring and operating same
CN110506393A
Multi-processor chip with shared FPGA execution unit and a design structure thereof
US20110307661A1