A Packing Method for FPGA Adaptive Logic Modules
Through a boxing method for FPGA adaptive logic modules, the problem of lack of effective boxing methods in the prior art is solved, and a low-cost and high-performance FPGA chip is designed, which expands the application scope.
Patent Information
- Application Number
- CN202111373586.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-19
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-11-19
AI Technical Summary
There is a lack of practical and effective packing methods for FPGA adaptive logic modules in the prior art, and it is difficult to assist in the design of low-cost and high-performance FPGA chips.
A boxing method for FPGA adaptive logic module is proposed, including obtaining boxing input, performing boxing process and output boxing results. The boxing process involves pre-boxing the combined logic unit and register unit into the adaptive logic module unit and boxing it into the adaptive logic module cluster.
This method has a novel and reasonable design, simple operation, and can assist in the design of low-cost and high-performance FPGA chips, and expand them to other types of FPGA modules, with high promotion and application value.
Smart Images

Figure CN114282471B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of software design of field programmable gate array (FPGA), and in particular relates to a packaging method for FPGA adaptive logic modules. Background Art
[0002] Logic Array Block (LAB) and ALM are the basic building blocks of logic in the FPGA device structure. LAB is composed of ALMs, and these ALMs can be configured to realize logic functions, arithmetic functions and register functions. Each LAB consists of ten ALMs, various carry chains, shared arithmetic chains, LAB control signals, local interconnects and register chain connection lines. FPGA EDA software loads the relevant logic into the LAB and lays out the LAB on the FPGA chip, realizing the functions of the circuit by using local, shared arithmetic chain and register chain connections.
[0003] FPGA EDA binning processing module is an important configuration item of FPGA application circuit design software. The binning module that supports ALM structure has the following main functions: obtain binning rule data and basic logic unit information in user design (UDM), pre-bin the logic unit according to the corresponding pre-binning algorithm, generate ALM module, and then use the binning algorithm to generate UDM data structure with LAB information for subsequent layout and routing modules. The binning algorithm based on ALM structure is different from the traditional LE binning. The binning based on LE structure only needs to bin several LEs into CLB according to the netlist after comprehensive mapping. This is because LE is the smallest logic unit and the netlist after comprehensive mapping is a circuit described based on LE structure; while the ALM structure includes two combinational LUT logic blocks and two registered FF logic blocks. The combinational logic block and the registered logic block are the smallest logic units. Therefore, the binning process will include two stages: first, several LUTs and FFs are packed into the ALM structure, and then the ALM is packed into the CLB, both of which are subject to the constraints of ALM usage. In the prior art, there is still a lack of practical and effective packaging methods for FPGA adaptive logic modules. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide a packing method for FPGA adaptive logic modules in view of the deficiencies in the above-mentioned prior art. The method has a novel and reasonable design and is easy to operate. It can assist FPGA hardware architects in designing low-cost, high-performance FPGA chips and can also be expanded to other types of FPGA modules, with high value for promotion and application.
[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is: a packing method for FPGA adaptive logic modules, the method comprising the following steps:
[0006] Step 1: Obtaining a packing input; the packing input includes user-designed logic unit information, user circuit constraint information, and packing rule information;
[0007] Step 2, executing the packing process; the packing process includes a pre-packing process of pre-packing the combinational logic unit (lcell_comb) and the register unit (dffeas) into the adaptive logic module unit (alm_molecule) according to the pre-packing mode for the FPGA adaptive logic module structure, and a packing process of packing the pre-packed adaptive logic module unit (alm_molecule) into the adaptive logic module cluster (alm_cluster);
[0008] Step 3: Output the packing results: process the packed data, write the processed results back to the user design model (UDM), and output the packing result file.
[0009] In the above-mentioned packing method for FPGA adaptive logic module, the specific process of obtaining the packing input in step 1 is as follows:
[0010] Step 101: Read the logic unit information designed by the user, including the type of the logic unit, the starting point and end point of the signal, and whether it is an internal signal;
[0011] Step 102, read the constraint information of the user circuit, including the project name, project path, top-level entity file name, packing rule file, chip name, packing algorithm and strategy;
[0012] Step 103, reading packing rule information, including obtaining information of logical blocks and physical blocks and port information corresponding to each logical block and physical block, obtaining pre-packing rule information, obtaining formal packing rule information, and obtaining port mapping information of logical blocks and physical blocks.
[0013] In the above-mentioned packing method for FPGA adaptive logic modules, the types of logic units in step 101 include combinational logic units and register units.
[0014] In the above-mentioned packing method for FPGA adaptive logic modules, the information of the logic block and the physical block in step 103 includes four levels of block information, namely, logic unit (ATOM), logic module (MOLECULE), logic block (LCBlock), and physical block (PhyBlock);
[0015] The pre-packaging rule information in step 103 includes the port mapping relationship between the logic unit (ATOM) and the corresponding logic module (MOLECULE) in different pre-packaging modes;
[0016] The formal packing rule information in step 103 includes the port mapping relationship between the logic module (MOLECULE) and the corresponding logic block (LCBlock) in different packing modes.
[0017] In the above-mentioned packing method for FPGA adaptive logic module, the pre-packing mode in step 2 includes standard mode, extended LUT mode, arithmetic mode, shared arithmetic mode and LUT register mode;
[0018] The pre-packing process of pre-packing the combinational logic unit (lcell_comb) and the register unit (dffeas) into the adaptive logic module unit (alm_molecule) according to the pre-packing mode in step 2 includes:
[0019] Step A: Pre-packaging according to different situations:
[0020] Case A1: When the pre-packing mode is arithmetic mode, pre-packing is performed according to the carry signal;
[0021] Case A2: When the pre-packing mode is the shared arithmetic mode, pre-packing is performed according to the carry signal and the shared signal;
[0022] Case A3: When the pre-binning mode is the extended LUT mode, the combinational logic cell (lcell_comb) with 7 inputs is pre-binned as a separate ALM;
[0023] Case A4: When the pre-packing mode is the standard mode, the combinational logic unit (lcell_comb) and the register unit (dffeas) are directly pre-packed into the adaptive logic module unit (alm_molecule);
[0024] Case A5: When the pre-binning mode is the LUT register mode, the register cell (dffeas) associated with the combinatorial logic cell (lcell_comb) is pre-binned into the adaptive logic module (ALM) according to the pre-binning results of Case A1, Case A2, Case A3, and Case A4;
[0025] Step B, mapping each port of the combinatorial logic unit (lcell_comb) and the register unit (dffeas) with the adaptive logic module unit (alm_molecule) according to the pre-packed box mode;
[0026] Step C: modify the lookup table mask value (LUT_MASK value) of the combinatorial logic unit (lcell_comb) according to the pre-packing condition.
[0027] In the above-mentioned method for packing an FPGA adaptive logic module, the specific process of mapping each port of the combinational logic unit (lcell_comb) and the register unit (dffeas) with the adaptive logic module unit (alm_molecule) according to the pre-packing mode in step B is as follows:
[0028] Step B1, select a non-pre-packed combinatorial logic cell (lcell_comb) as a seed of a new adaptive logic module (ALM);
[0029] Step B2, sorting the unpre-packed combinatorial logic units (lcell_comb) according to the size of the input numbers shared with the current combinatorial logic unit (lcell_comb), and storing them in the shared combinatorial logic unit (lcell_comb) set;
[0030] Step B3, sequentially select combinatorial logic cells (lcell_comb) from the set to add to the current adaptive logic module (ALM), and check whether the pre-packing rule is met. If the pre-packing rule is not met, select the next combinatorial logic cell (lcell_comb) in the set until a combinatorial logic cell (lcell_comb) that meets the requirement is found; when there are no other combinatorial logic cells (lcell_comb) in the circuit that have signal connections with the current combinatorial logic cell (lcell_comb), find a combinatorial logic cell (lcell_comb) that meets the pre-packing rule from all the combinatorial logic cells (lcell_comb) that are not pre-packed and add it to the current adaptive logic module (ALM);
[0031] Step B4: Load the register unit (dffeas) associated with each combinatorial logic unit (lcell_comb) into the corresponding adaptive logic module (ALM).
[0032] In the above-mentioned method for packing an FPGA adaptive logic module, the packing process of packing the pre-packed adaptive logic module unit (alm_molecule) into the adaptive logic module cluster (alm_cluster) in step 2 includes the following three packing cases:
[0033] Case D1, Carry chain packing: For the adaptive logic module unit (alm_molecule) in arithmetic mode, pack according to the carry signal;
[0034] Case D2, shared chain packing: For the adaptive logic module unit (alm_molecule) in the shared arithmetic mode, packing is performed according to the carry signal and the shared signal;
[0035] Case D3: Except for cases D1 and D2, greedy packing based on cost function is performed: according to the idea of greedy algorithm, the unpacked adaptive logic module units (alm_molecule) are added to the adaptive logic module cluster (alm_cluster) in sequence according to the size of attraction until the adaptive logic module cluster (alm_cluster) is full.
[0036] The above-mentioned packing method for FPGA adaptive logic modules, as described in case D1, for the adaptive logic module unit (alm_molecule) in the arithmetic mode, the specific process of packing according to the carry signal is: the adaptive logic modules (ALM) located on the carry chain are tried to be added to the logic array block (LAB) in sequence according to the order of the carry signal, and at the same time, it is determined whether the packing constraints are met, and the adaptive logic modules (ALM) that meet the packing constraints are packed into one logic array block (LAB). When the packing constraints are not met, a new logic array block (LAB) is created, and the packing operation is repeated until all the ALMs on the entire carry chain are loaded into the logic array block (LAB).
[0037] The above-mentioned packing method for FPGA adaptive logic modules, as described in case D2, for the adaptive logic module unit (alm_molecule) in the shared arithmetic mode, the specific process of packing according to the carry signal and the shared signal is: the adaptive logic modules (ALM) located on the shared chain are tried to be added to the logic array block (LAB) in the order of the shared carry signal, and at the same time, it is determined whether the packing constraints are met, and the adaptive logic modules (ALM) that meet the packing constraints are packed into one logic array block (LAB). When the packing constraints are not met, a new logic array block (LAB) is created, and the packing operation is repeated until all the adaptive logic modules (ALM) on the entire shared chain are loaded into the logic array block (LAB).
[0038] In the above-mentioned binning method for the FPGA adaptive logic module, the process of processing the binned data in step 3 includes:
[0039] Step 301, setting the internal signal flag of each signal after packaging;
[0040] Step 302: Set the parameter values of each block after packing, including the subloc and the name of the cluster.
[0041] Step 303: performing port mapping of the packed adaptive logic module unit (alm_molecule) and the adaptive logic module cluster (alm_cluster);
[0042] Step 304: perform port mapping of the packed adaptive logic module cluster (alm_cluster) and the adaptive logic module physical block (alm_tile).
[0043] Compared with the prior art, the present invention has the following advantages: the present invention innovatively proposes a packing method for FPGA adaptive logic modules, which is novel and reasonable in design, simple to operate, can assist FPGA hardware architects in designing low-cost, high-performance FPGA chips, and can also be expanded to other types of FPGA modules, with high value for promotion and application.
[0044] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a structural diagram of the FPGA adaptive logic module;
[0046] Figure 2 This is a schematic diagram of the structure of lcell_comb;
[0047] Figure 3 It is a schematic diagram of the structure of dffeas;
[0048] Figure 4 This is a schematic diagram of the arrangement of ALM in LAB;
[0049] Figure 5 A flowchart of a method for packing a FPGA adaptive logic module according to the present invention;
[0050] Figure 6 Shown is a schematic diagram of the standard mode of ALM;
[0051] Figure 7 Shown is a schematic diagram of the extended LUT mode of ALM;
[0052] Figure 8 Shown is a schematic diagram of the arithmetic mode of ALM;
[0053] Fig. 9 Shown is a schematic diagram of the shared arithmetic mode of ALM;
[0054] Fig.10 The figure shows the LUT register mode diagram of ALM;
[0055] Fig.11It is a flow chart of mapping various ports of lcell_comb and dffeas with alm_molecule according to the pre-packaging mode of the present invention;
[0056] Fig.12 is a flow chart for modifying the LUT_mask value of the present invention;
[0057] Fig.13 It is a flow chart of the present invention for boxing alm_molecule in arithmetic mode according to the carry signal;
[0058] Fig.14 It is a flow chart of the present invention for boxing alm_molecule in the shared arithmetic mode according to the carry signal and the shared signal;
[0059] Fig.15 It is a schematic diagram of the process of greedy packing based on cost function of the present invention;
[0060] Fig.16 This is a schematic diagram of the description format of the packing result file of the present invention. DETAILED DESCRIPTION
[0061] Figure 1The structural diagram of the FPGA adaptive logic module (ALM) is given. The ALM contains a variety of LUT register-based resources, which can be divided from the combination of adaptive LUT (ALUT) and two registers. By using the 8 inputs of the two combined ALUTs, an ALM can implement various combinations of these two functions. This adaptability makes the ALM fully backward compatible with the 4-input LUT architecture. An ALM can also implement any function through 6 inputs and some 7-input functions. In addition to the adaptive LUT-based resources, each ALM also includes two programmable registers, two dedicated full adders, a carry chain, a shared arithmetic chain, and a register chain. With these dedicated resources, an ALM can effectively implement various arithmetic functions and shift registers. Each ALM can drive all types of interconnects: local, row, column, carry chain, shared arithmetic chain, register chain, and direct link. An ALM contains two programmable registers. Each register has data, clock, clock enable, synchronous and asynchronous clear, and synchronous load and clear inputs. Global signals, general purpose I / O pins, or internal logic can drive the registers’ clock and clear control signals. General purpose I / O pins or internal logic can drive the clock enable signals. For combinatorial logic functions, the registers are bypassed and the outputs of the LUTs drive directly to the outputs of the ALMs. Each ALM has two sets of outputs that drive local, row, and column routing resources. LUT, adder, or register outputs can drive these output drivers. For each set of output drivers, two ALM outputs can drive column, row, or direct link routing connections. One of the ALM outputs can also drive local interconnect resources. This allows a register to drive one output while a LUT or adder drives another output. This feature, called register packing, can improve device utilization because the device can use registers and combinatorial logic for unrelated functions. Another special packing mode allows register outputs to be fed back into the LUTs of the same ALM, allowing registers to be packed with their own fan-out LUTs. This provides improved placement and routing for another mechanism. ALMs can also drive registered LUTs as well as unregistered LUT or adder outputs. Each ALM has eight fracturable look-up table (LUT) inputs, two dedicated embedded adders, two dedicated registers, and additional enhancement logic.
[0062] FPGA adaptive logic module (ALM) is mainly composed of two parts: lcell_comb and dffeas. Figure 2The structural diagram of lcell_comb is given. lcell_comb is mainly composed of a LUT and a full adder. The LUT is composed of four 4-input LUTs, which are paired with two of each other, followed by six 2-to-1 muxes divided into two groups, forming two 4-to-1 muxes, which are divided into two outputs. The full adder is a special adder that can perform ordinary addition and can also cooperate with the LUT to add three numbers. The adder adopts a carry-skip structure to speed up the calculation. Figure 3 The structural diagram of dffeas is given. dffeas can support asynchronous reset, synchronous reset and synchronous set functions, with the priorities decreasing in turn; the input source of the d terminal can be the output of lcell_comb or the input of cascade, and the input of register_packing is the asdata terminal.
[0063] A logic array block (LAB) contains 10 ALMs, each of which can be composed of 2 lcell_combs and 2 dffeas. The arrangement of ALMs in LAB is as follows: Figure 4 As shown, the numbers are from 0 to 9, the lcell_comb inside ALM is numbered as an even number, and the dffeas is numbered as an odd number. The total number of lcell_comb and dffeas is 40 cells, numbered from 0 to 39.
[0064] The port definition of the basic logic unit in ALM is shown in Table 1:
[0065] Table 1 Port definition table of basic logic units in ALM
[0066]
[0067] ALM includes two units: lcell_comb and dffeas. lcell_comb is a combinational unit, whose input ports include data ports (dataa, datab, datac, datad, datae, dataf, datag) and carry input port (cin), and whose output ports include combinational output port (combout) and carry output port (cout). dffeas is a register unit, whose input ports include data port (d), clock port (clk), clear port (clrn), clock enable port (ena), register input port (asdata), asynchronous set port (aload), synchronous clear port (sclr), synchronous set port (sload), and whose output ports include register output port (q).
[0068] The port definitions of the basic logic modules in ALM are shown in Table 2:
[0069] Table 2 Port definition table of basic logic modules in ALM
[0070]
[0071]
[0072] The ports of the ALM module include input ports and output ports. The input ports include 8 data ports (dataa, datab, datac, datad, datae0, dataf0, datae1, dataf1), 2 asynchronous clear ports (nclr0, nclr1), 1 asynchronous set port (aload), 1 synchronous clear port (sclr), 1 synchronous set port (sload), 2 clock ports (clk0, clk1), 2 clock enable ports (ena0, ena1), 1 carry input port (cin), 1 register control port (re gscan), 1 register cascade input port (regchain_i), 1 shared arithmetic input port (shared_arith_i), the output ports include 1 registered cascade output port (regchain_o), 1 carry output port (cout), 1 shared arithmetic output port (shared_arith_out), and 6 output ports (alm_out1d, alm_out1u, alm_out2d, alm_out2u, alm_out0d, alm_out0u).
[0073] like Figure 5 As shown, the packing method for FPGA adaptive logic module of the present invention comprises the following steps:
[0074] Step 1, obtain packing input; the packing input includes user-designed logic unit information (User Design Model, UDM), user circuit constraint information (User Constraint Model, UCM) and packing rule information (pack_guide.xml);
[0075] In this embodiment, the specific process of obtaining the boxing input in step 1 is:
[0076] Step 101, read the logic unit (Atom_Set) information designed by the user, including the type of the logic unit, the starting point and end point of the signal, and whether it is an internal signal;
[0077] In this embodiment, the types of the logic cells in step 101 include a combinatorial logic cell (lcell_comb) and a register cell (dffeas).
[0078] Step 102, read the constraint (Constraint_Set) information of the user circuit, including the project name, project path, top-level entity file name, packing rule file, chip name, packing algorithm and strategy;
[0079] Step 103, reading packing rule information, including obtaining information of logical blocks and physical blocks and port information corresponding to each logical block and physical block, obtaining pre-packing rule information, obtaining formal packing rule information, and obtaining port mapping information of logical blocks and physical blocks.
[0080] In this embodiment, the information of the logical block and the physical block in step 103 includes four levels of block information, namely, logical unit (ATOM), logical module (MOLECULE), logical block (LCBlock), and physical block (PhyBlock);
[0081] In this embodiment, the pre-packaging rule information in step 103 includes the port mapping relationship between the logic unit (ATOM) and the corresponding logic module (MOLECULE) in different pre-packaging modes;
[0082] In this embodiment, the formal packing rule information in step 103 includes the port mapping relationship between the logic module (MOLECULE) and the corresponding logic block (LCBlock) in different packing modes.
[0083] Step 2, executing the packing process; the packing process includes a pre-packing process of pre-packing the combinational logic unit (lcell_comb) and the register unit (dffeas) into the adaptive logic module unit (alm_molecule) according to the pre-packing mode for the FPGA adaptive logic module (ALM) structure, and a packing process of packing the pre-packed adaptive logic module unit (alm_molecule) into the adaptive logic module cluster (alm_cluster);
[0084] That is, the pre-packing algorithm is executed to generate a pre-packed adaptive logic module unit (alm_molecule), and the packing algorithm is executed to generate a packed adaptive logic module cluster (alm_cluster).
[0085] In this embodiment, the pre-packing modes in step 2 include a standard mode (Normal), an extended LUT mode (Extended LUT), an arithmetic mode (Arithmetic), a shared arithmetic mode (Shared Arithmetic) and a LUT register mode (LUT-Register);
[0086] In specific implementation, each mode uses ALM resources in a different way; in each mode, the eleven inputs of ALM are directed to different destinations to implement the required logic functions (these eleven inputs include eight data inputs from the LAB local interconnect, carry-in and shared arithmetic chain connections from the previous ALM or LAB, and register chain connections); full LAB signals provide clock, asynchronous clear, synchronous clear, synchronous load, and clock enable control signals for registers. These full LAB signals are available in all ALM modes.
[0087] Figure 6 The figure shows the normal mode, which is suitable for general logic applications and combinational functions. In this mode, the eight data inputs from the LAB local interconnect are the inputs of the combinational logic. The normal mode supports two functions in one Stratix IV ALM, or one function with six inputs. The ALM supports some completely independent combinational functions, as well as various combinational functions with common inputs.
[0088] Figure 7 Shown is the Extended LUT mode (ExtendedLUT) to implement a specific set of 7-input functions. This specific set must be a 2-to-1 multiplexer driven by two arbitrary 5-input functions that share four inputs. In this mode, if the 7-input function is unregistered, the unused 8th input can be used for register packing.
[0089] Figure 8Shown is the arithmetic mode, which is ideal for implementing adders, counters, accumulators, full parity functions, and comparators. The ALM in arithmetic mode uses two sets of two four-input LUTs and two dedicated full adders. The dedicated adders support the LUTs for performing pre-adder logic; therefore, each adder can add the outputs of two four-input functions. The four LUTs share the dataa and datab inputs. The carry-in signal drives to adder0, and the carry-out signal from adder0 drives to the carry-in of adder1. The carry-out from adder1 drives to adder0 of the next ALM in the LAB. The ALM in arithmetic mode can drive registered adder outputs and / or unregistered adder outputs. When operating in arithmetic mode, the ALM supports the use of the carry output of the adder and the output of the combinational logic at the same time. In this operation, the adder output is ignored. Using adders and combinational logic together can save up to 50% of resources when this feature is used. In addition, arithmetic mode also supports clock enable, counter enable, synchronous up / down control, add / subtract control, synchronous clear, and synchronous load functions. The LAB local interconnect data input generates the clock enable, counter enable, synchronous up / down, and add / subtract control signals. These control signals are good choices for inputs shared between the four LUTs in the ALM. The synchronous clear and synchronous load options are LAB-wide signals that affect all registers in the LAB. These signals can also be disabled or enabled individually on a per-register basis. The carry chain provides a fast carry function between dedicated adders in arithmetic or shared arithmetic modes. The 2-bit carry select feature reduces the carry chain propagation delay in half in the ALM. The carry can start at the first ALM in the LAB, or the fifth ALM. The final carry-out signal is propagated to the ALM and driven to the local, row, or column interconnect.
[0090] Fig. 9Shown is the shared arithmetic mode (SharedArithmetic), which enables the ALM to implement 3-input addition operations in the ALM. In this mode, the ALM is configured with four 4-input LUTs. Each LUT will calculate the sum of three inputs or calculate the carry of three inputs. The output of the carry calculation is provided to the next adder (can be used for adder1 in the same ALM, or adder0 in the next ALM) by using a dedicated connection called a shared arithmetic chain. This shared arithmetic chain can significantly improve the performance of the adder tree by reducing the summation steps used to implement the adder tree. The shared arithmetic chain in enhanced arithmetic mode enables the ALM to implement three-input addition operations, significantly reducing the resources required to implement large adder trees or correlator functions. The shared arithmetic chain starts at the first ALM or the sixth ALM in the LAB.
[0091] Fig.10 LUT-Register mode supports the third register capability in the ALM. Two internal feedback loops enable combinational ALUT1 to implement the master latch required for the third register and combinational ALUT0 to implement the slave latch required for the third register. The LUT register shares its clock, clock enable, and asynchronous clear sources with the top dedicated register.
[0092] The pre-packing process of pre-packing the combinational logic unit (lcell_comb) and the register unit (dffeas) into the adaptive logic module unit (alm_molecule) according to the pre-packing mode in step 2 includes:
[0093] Step A: Pre-packaging according to different situations:
[0094] Case A1: When the pre-packing mode is the arithmetic mode, pre-packing is performed according to the carry signal (the signal on the ports cin and cout); this case is the pre-packing of the carry chain of the ALM structure;
[0095] Case A2: When the pre-binning mode is the shared arithmetic mode (shared_arith=ON of parameter lcell_comb), pre-binning is performed according to the carry signal (signal on ports cin and cout) and the shared signal (signal on ports sharein and shareout); this case is the pre-binning of the shared chain of the ALM structure, in which the carry signal and the shared signal exist at the same time;
[0096] Case A3: When the pre-binning mode is the extended LUT mode (Extended LUT), the combinational logic cell (lcell_comb) with 7 inputs (parameter extended_lut=ON) is pre-binned as a separate ALM;
[0097] Case A4: When the pre-packing mode is the standard mode (Normal) (i.e., except for the above cases A1, A2, and A3), the combinational logic unit (lcell_comb) and the register unit (dffeas) are directly pre-packed into the adaptive logic module unit (alm_molecule);
[0098] Case A5: When the pre-binning mode is the LUT register mode (LUT-Register), the register cell (dffeas) associated with the combinatorial logic cell (lcell_comb) is pre-binned into the adaptive logic module (ALM) according to the pre-binning results of Case A1, Case A2, Case A3, and Case A4;
[0099] Step B, mapping each port of the combinatorial logic unit (lcell_comb) and the register unit (dffeas) with the adaptive logic module unit (alm_molecule) according to the pre-packed box mode;
[0100] In this embodiment, Fig.11 As shown, the specific process of mapping the ports of the combinatorial logic unit (lcell_comb) and the register unit (dffeas) with the adaptive logic module unit (alm_molecule) according to the pre-packed box mode in step B is:
[0101] Step B1, select a non-pre-packed combinatorial logic cell (lcell_comb) as a seed of a new adaptive logic module (ALM);
[0102] Step B2, sorting the unpre-packed combinatorial logic units (lcell_comb) according to the size of the input numbers shared with the current combinatorial logic unit (lcell_comb), and storing them in the shared combinatorial logic unit (lcell_comb) set;
[0103] Step B3, sequentially select combinatorial logic cells (lcell_comb) from the set to add to the current adaptive logic module (ALM), and check whether the pre-packing rule is met. If the pre-packing rule is not met, select the next combinatorial logic cell (lcell_comb) in the set until a combinatorial logic cell (lcell_comb) that meets the requirement is found; when there are no other combinatorial logic cells (lcell_comb) in the circuit that have signal connections with the current combinatorial logic cell (lcell_comb), find a combinatorial logic cell (lcell_comb) that meets the pre-packing rule from all the combinatorial logic cells (lcell_comb) that are not pre-packed and add it to the current adaptive logic module (ALM);
[0104] Step B4: Load the register unit (dffeas) associated with each combinatorial logic unit (lcell_comb) into the corresponding adaptive logic module (ALM).
[0105] Step C: modify the lookup table mask value (LUT_MASK value) of the combinatorial logic unit (lcell_comb) according to the pre-packing situation (ie, when the port to which the signal of lcell_comb is connected changes).
[0106] During the pre-binning process, lcell_comb performs a data port swap operation to satisfy the pre-binning mode. The port swap will cause the LUT_mask value to change. LUT_mask has 64 bits in total. Replacing two ports in lcell will cause 32 bits in LUT_mask to change. The change rules are as follows:
[0107] Original LUT mask: m [63:0]
[0108] The input terminals are labeled according to the address: {f,e,d,c,b,a}=A [5:0]
[0109] lcell output is: LUTOUT = m[A [5] *32+A [4] *16+A [3] *8+A [2] *4+A [1] *2+A [0] ]
[0110] If lut_mask after replacing A[i] and A[j] is: m' [63:0]
[0111] lcell output is: LUTOUT = m'[A [5] *32+...+A [i]*2 j +A [j] *2 i ...+A [0] ]
[0112] m * [A [5] *32+...+A [i] *2 j +A [j] *2 i ...+A [0] ]=
[0113] m[A [5] *32+A [4] *16+A [3] *8+A [2] *4+A [1] *2+A [0] ]
[0114] According to the above change rules, the modification flow chart of LUT_mask value is as follows Fig.12 shown.
[0115] In this embodiment, the packing process of packing the pre-packed adaptive logic module unit (alm_molecule) into the adaptive logic module cluster (alm_cluster) in step 2 includes the following three packing situations:
[0116] Case D1, Carry chain (arithmetic mode) packing: For the adaptive logic module unit (alm_molecule) in the arithmetic mode (Arithmetic), packing is performed according to the carry signal (signal on ports cin and cout);
[0117] In this embodiment, Fig.13 As shown, the specific process of boxing the adaptive logic module unit (alm_molecule) in the arithmetic mode (Arithmetic) in case D1 according to the carry signal (signal on the port cin, cout) is as follows: according to the unique structure of the carry chain, the ALMs located on the carry chain are tried to be added to the LAB in sequence according to the order of the carry signal, and at the same time, it is determined whether the boxing constraints are met, and the ALMs that meet the boxing constraints are packed into one LAB. When the boxing constraints are not met, a new LAB is created, and the boxing operation is repeated until all the ALMs on the entire carry chain are packed into the LAB.
[0118] Case D2, shared chain (shared arithmetic mode) packing: for the adaptive logic module unit (alm_molecule) in shared arithmetic mode (parameter shared_arith=ON of lcell_comb), packing is performed according to the carry signal (signal on ports cin and cout) and the shared signal (signal on ports sharein and shareout);
[0119] In this embodiment, Fig.14 As shown, in case D2, for the adaptive logic module unit (alm_molecule) in the shared arithmetic mode (parameter shared_arith=ON of lcell_comb), the specific process of boxing according to the carry signal (signal on ports cin and cout) and the shared signal (signal on ports sharein and shareout) is as follows: according to the unique structure of the shared chain, the ALMs located on the shared chain are tried to be added to the LAB in sequence according to the order of the shared carry signal, and at the same time, it is determined whether the boxing constraints are met, and the ALMs that meet the boxing constraints are packed into one LAB. When the boxing constraints are not met, a new LAB is created, and the boxing operation is repeated until all the ALMs on the entire shared chain are packed into the LAB.
[0120] In specific implementation, the port mappings of ALM and lcell_comb in different sharing modes are shown in Table 3-1, Table 3-2, Table 3-3 and Table 3-4:
[0121] Table 3-1 Port mapping between ALM ports and lcell_comb in different sharing modes
[0122]
[0123] Table 3-2 Port mapping between ALM ports and lcell_comb in different sharing modes
[0124]
[0125] Table 3-3 Port mapping between ALM ports and lcell_comb in different sharing modes
[0126]
[0127] Table 3-4 Port mapping between ALM ports and lcell_comb in different sharing modes
[0128]
[0129] 1 ALM can contain 2 lcell_combs, and the data ports of the 2 lcell_combs can be shared or independent, but they need to meet certain pre-packing rules, that is, the total data input of the 2 lcell_combs cannot exceed 8, and the ports of the shared signal must be specific ports. Installing 2 lcell_combs into 1 ALM can be divided into the following cases: independent 4-input LUT; dual 4-input LUT, 1 shared input port; dual 4-input LUT, 2 shared inputs; dual 5-input LUT, 3 shared inputs; dual 5-input LUT, 2 shared inputs; independent 5LUT and 3LUT; 5LUT and 4LUT, 1 shared input; dual 6-input LUT, 4 shared inputs; dual 5-input LUT, 4 shared inputs; 6-input LUT; extended mode 7-input LUT.
[0130] Case D3: Except for cases D1 and D2, greedy packing based on cost function is performed: according to the idea of greedy algorithm, the unpacked adaptive logic module units (alm_molecule) are added to the adaptive logic module cluster (alm_cluster) in sequence according to the size of attraction until the adaptive logic module cluster (alm_cluster) is full.
[0131] When implementing it, Fig.15 As shown in the figure, firstly, an unpacked ALM is selected as the seed of the new LAB, and then the unpacked ALMs are added in turn according to the attractiveness until the LAB is filled. The cost function needs to consider both the number of logic units in the logic block and the number of connections between logic blocks on the critical path.
[0132] In specific implementation, when the corresponding packing constraints are met, it means that the alm_cluster is full.
[0133] Step 3: Output the packing results: process the packed data, write the processed results back to the user design model (UDM), and output the packing result file.
[0134] In the specific implementation, in order to facilitate the user to view the packing results, the packing results are finally written into the .cluster file. The user can check and analyze the association between the logic units in the circuit based on the results. In this embodiment, the process of processing the packed data in step 3 includes:
[0135] Step 301, setting the internal signal flag of each signal after packaging;
[0136] Step 302: Set the parameter values of each block after packing, including the subloc and the name of the cluster.
[0137] Step 303: performing port mapping of the packed adaptive logic module unit (alm_molecule) and the adaptive logic module cluster (alm_cluster);
[0138] Step 304: perform port mapping of the packed adaptive logic module cluster (alm_cluster) and the adaptive logic module physical block (alm_tile).
[0139] like Fig.16 The following is the description format of the packing result file. According to the packing level, the adaptive logic module cluster (alm_cluster) contains the adaptive logic module unit (alm_molecule), and the adaptive logic module unit (alm_molecule) contains the corresponding combinational logic unit (lcell_comb) and register unit (dffeas). The naming method of the adaptive logic module cluster (alm_cluster) is "type + number", and the naming method of the adaptive logic module unit (alm_molecule) is "type + location of the adaptive logic module unit (alm_molecule) + number of the adaptive logic module cluster (alm_cluster)". For the adaptive logic module cluster (alm_cluster) on the carry chain, shared arithmetic chain and register cascade chain, the corresponding linked list number and the number of the adaptive logic module cluster (alm_cluster) in the linked list are given.
[0140] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0141] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0142] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0143] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0144] The foregoing description of specific exemplary embodiments of the present invention is for the purpose of illustration and demonstration. These descriptions are not intended to limit the present invention to the precise form disclosed, and it is clear that many changes and variations can be made based on the above teachings. The purpose of selecting and describing the exemplary embodiments is to explain the specific principles of the present invention and its practical application, so that those skilled in the art can realize and utilize various different exemplary embodiments of the present invention and various different selections and changes. The scope of the present invention is intended to be limited by the claims and their equivalents.
Claims
1. A packaging method for FPGA adaptive logic modules, characterized in that: The method comprises the following steps: Step 1: Obtaining a packing input; the packing input includes user-designed logic unit information, user circuit constraint information, and packing rule information; Step 2, executing a packing process; the packing process includes a pre-packing process of pre-packing a combinational logic unit and a register unit into an adaptive logic module unit according to a pre-packing mode for the FPGA adaptive logic module structure, and a packing process of packing the pre-packed adaptive logic module unit into an adaptive logic module cluster; Step 3: Output the packing result: process the packed data, write the processed result back to the user design model, and output the packing result file; The pre-packing modes in step 2 include standard mode, extended LUT mode, arithmetic mode, shared arithmetic mode and LUT register mode; The pre-packing process of pre-packing the combinational logic unit and the register unit into the adaptive logic module unit according to the pre-packing mode in step 2 includes: Step A: Pre-packaging according to different situations: Case A1: When the pre-packing mode is arithmetic mode, pre-packing is performed according to the carry signal; Case A2: When the pre-packing mode is the shared arithmetic mode, pre-packing is performed according to the carry signal and the shared signal; Case A3: When the pre-binning mode is the extended LUT mode, the combinational logic unit with 7 inputs is pre-binned as a separate ALM; Case A4: When the pre-packing mode is the standard mode, the combinational logic unit and the register unit are directly pre-packed into the adaptive logic module unit; Case A5: When the pre-binning mode is the LUT register mode, the register unit associated with the combinational logic unit is pre-binned into the adaptive logic module according to the pre-binning results of Case A1, Case A2, Case A3, and Case A4; Step B, mapping each port of the combinational logic unit and the register unit to the adaptive logic module unit according to the pre-packed box mode; Step C, modifying the lookup table mask value of the combinational logic unit according to the pre-packing condition; In addition, the packing process of packing the pre-packed adaptive logic module units into the adaptive logic module cluster in step 2 includes the following three packing situations: Case D1, Carry chain packing: For the adaptive logic module unit in arithmetic mode, packing is performed according to the carry signal; Case D2, shared chain packing: For the adaptive logic module unit in the shared arithmetic mode, packing is performed according to the carry signal and the shared signal; Case D3: Except for cases D1 and D2, greedy packing based on cost function is performed: according to the idea of greedy algorithm, the unpacked adaptive logic module units are added to the adaptive logic module cluster in sequence according to the size of attraction until the adaptive logic module cluster is full.
2. A packing method for FPGA adaptive logic modules according to claim 1, characterized in that: The specific process of obtaining the boxing input described in step 1 is: Step 101: Read the logic unit information designed by the user, including the type of the logic unit, the starting point and end point of the signal, and whether it is an internal signal; Step 102, read the constraint information of the user circuit, including the project name, project path, top-level entity file name, packing rule file, chip name, packing algorithm and strategy; Step 103, reading packing rule information, including obtaining information of logical blocks and physical blocks and port information corresponding to each logical block and physical block, obtaining pre-packing rule information, obtaining formal packing rule information, and obtaining port mapping information of logical blocks and physical blocks.
3. A packing method for FPGA adaptive logic modules according to claim 2, characterized in that: The types of logic units in step 101 include combinational logic units and register units.
4. A packaging method for FPGA adaptive logic modules according to claim 2, characterized in that: The information of the logical block and the physical block in step 103 includes four levels of block information, namely, logical unit, logical module, logical block, and physical block; The pre-packaging rule information in step 103 includes the port mapping relationship between the logic units and the corresponding logic modules in different pre-packaging modes; The formal packing rule information in step 103 includes the port mapping relationship between the logic modules and the corresponding logic blocks in different packing modes.
5. A packaging method for FPGA adaptive logic modules according to claim 1, characterized in that: The specific process of mapping the ports of the combinational logic unit and the register unit to the adaptive logic module unit according to the pre-packed box mode in step B is as follows: Step B1, selecting a non-pre-packed combinational logic unit as a seed of a new adaptive logic module; Step B2, sorting the combinational logic units that are not pre-packed according to the number of inputs shared with the current combinational logic unit, and storing them in a shared combinational logic unit set; Step B3, sequentially selecting combinational logic units from the set to add to the current adaptive logic module, and checking whether the pre-packing rule is satisfied. If the pre-packing rule is not satisfied, selecting the next combinational logic unit in the set until a combinational logic unit that meets the requirement is found; if there is no other combinational logic unit in the circuit that has a signal connection with the current combinational logic unit, finding a combinational logic unit that meets the pre-packing rule from all combinational logic units that are not pre-packed and adding it to the current adaptive logic module; Step B4: Load the register units associated with each combinational logic unit into the corresponding adaptive logic module.
6. A packaging method for FPGA adaptive logic modules according to claim 1, characterized in that: The specific process of boxing according to the carry signal for the adaptive logic module unit in the arithmetic mode described in situation D1 is as follows: the adaptive logic modules located on the carry chain are tried to be added to the logic array block in sequence according to the order of the carry signal, and whether the boxing constraints are met is determined at the same time, and the adaptive logic modules that meet the boxing constraints are packed into one logic array block. When the boxing constraints are not met, a new logic array block is created, and the boxing operation is repeated until all the ALMs on the entire carry chain are packed into the logic array block.
7. A packaging method for FPGA adaptive logic modules according to claim 1, characterized in that: The specific process of boxing the adaptive logic module units in the shared arithmetic mode according to the carry signal and the shared signal described in situation D2 is as follows: the adaptive logic modules on the shared chain are tried to be added to the logic array block in sequence according to the order of the shared carry signal, and whether the boxing constraints are met are determined at the same time. The adaptive logic modules that meet the boxing constraints are packed into one logic array block. When the boxing constraints are not met, a new logic array block is created, and the boxing operation is repeated until all the adaptive logic modules on the entire shared chain are packed into the logic array block.
8. A packaging method for FPGA adaptive logic modules according to claim 1, characterized in that: The process of processing the boxed data in step 3 includes: Step 301, setting the internal signal flag of each signal after packaging; Step 302: Set the parameter values of each block after packing, including the subunit position and the name of the cluster; Step 303, performing port mapping of the packed adaptive logic module units and adaptive logic module clusters; Step 304: perform port mapping of the packed adaptive logic module cluster and the adaptive logic module physical block.
Citation Information
Patent Citations
Three-dimensional FPGA device layout optimization method based on thermal simulation
CN107729704A
FPGA logic synthesis method and device for realizing summation operation based on yosys
CN113568598A