An FPGA parallel processing method based on netlist division
By dividing the FPGA circuit netlist into multiple subnetlists for parallel processing, the problem of excessively long synthesis time for large-scale FPGAs is solved, achieving a faster synthesis speed.
Patent Information
- Application Number
- CN202511300110.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-09-12
AI Technical Summary
As FPGA chip size and user application circuit design become larger, traditional FPGA synthesis tools, which use single-threaded processing, suffer from excessively long synthesis times.
A parallel processing method based on netlist partitioning is adopted. By executing concurrently, the circuit netlist is divided into multiple sub-netlists for parallel processing, and the results are generated by merging them.
It significantly reduces the time required for application circuit synthesis and improves the running speed of FPGA synthesis.
Smart Images

Figure CN120822469B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of FPGA, and particularly relates to an FPGA parallel processing method based on netlist division. BACKGROUND
[0002] Logic synthesis is an important step in the design process of FPGA EDA software, which is to convert the user input behavior level or register transfer level (RTL) Verilog / VHDL circuit file into a netlist file composed of FPGA basic logic units such as lookup table (Lut) and flip-flop (FF).
[0003] FPGA logic synthesis includes two stages: synthesis and mapping. Synthesis is to convert the behavior level or RTL circuit file into a logic netlist composed of gate circuits; mapping is to map the logic netlist composed of gate circuits into a netlist file composed of FPGA basic logic units.
[0004] With the increase of the scale of FPGA chips, the scale of user application circuit design is also increasing, which brings great challenges to FPGA synthesis tools, because the time of complex application circuit synthesis will become very long. The business of FPGA synthesis tools is complex, the processing process is multiple, and because the traditional FPGA synthesis tool is processed in a single thread mode, the time of application circuit synthesis is relatively long. SUMMARY
[0005] The application provides an FPGA parallel processing method based on netlist division, which reduces the time required for application circuit synthesis by concurrent execution, and further improves the running speed of FPGA synthesis.
[0006] Other purposes and advantages of the application can be further understood from the technical features disclosed in the application.
[0007] To achieve one or part or all of the above purposes or other purposes, the application provides an FPGA parallel processing method based on netlist division.
[0008] An FPGA parallel processing method based on netlist division comprises:
[0009] Step S1: generating N threads according to the number N of computer CPU cores;
[0010] Step S2: dividing the total netlist of the current circuit according to the N threads, and calculating the sub-netlist processed by each thread;
[0011] Step S3: processing the sub-netlist in each thread, obtaining the logic unit and signal of each thread sub-netlist, and merging into the total netlist;
[0012] Step S4: processing the merged total netlist, outputting all logic units and signals into a synthesis result file.
[0013] The specific process of step S2 includes:
[0014] obtaining a top module of the circuit total netlist and a list of sub-modules called by the top module;
[0015] analyzing the top module and the list of sub-modules to obtain the number of logic units contained in each module;
[0016] sorting each module in descending order of the number of logic units, and adding the sorted modules to a queue one by one;
[0017] iteratively processing the queue until the queue is empty.
[0018] After the end of the iterative processing of the queue, deleting threads that do not contain any modules.
[0019] The iterative processing process includes:
[0020] taking the first module from the queue and placing it in the sub-netlist W1 of thread T1, setting the index i, and determining whether the queue is empty;
[0021] If the queue is not empty, continuously taking the first module from the queue and placing it in the sub-netlist Wi of thread Ti until the total number of logic units contained in the sub-netlist Wi is greater than or equal to the total number of logic units contained in all modules in the sub-netlist W1, until the processing of the Nth thread or the queue Q is empty;
[0022] If the current queue Q is not empty, iteratively processing the threads T1, T2, …, TN in the next round, and the end condition of iteration is that the queue Q is empty.
[0023] If the queue is empty, the iterative processing process ends, and threads that do not contain any modules are deleted.
[0024] Setting the index i to 2 to N.
[0025] The specific process of processing the sub-netlist in each thread in step S3 includes:
[0026] extracting the logic units of each module in the sub-netlist and performing rough optimization to remove useless logic units and merge equivalent logic units;
[0027] Processing the carry chain units in each module to generate a set of combinational logic block units of carry chain functions;
[0028] Processing the memory unit in each module, converting each memory unit into a set of memory basic block units;
[0029] Processing the digital processing logic unit in each module, converting each digital processing logic unit into a set of basic multiplication unit and multiplication-accumulation output unit.
[0030] The extracted logic units include: multiplexer units, register units, carry chain units, memory units and digital processing logic units.
[0031] The multiplexer units, basic gate circuit units and register units are roughly optimized, useless logic units are removed, and equivalent logic units are merged.
[0032] The specific process of processing the merged total netlist includes:
[0033] Processing each register unit in the netlist, extracting the control signal of each register unit;
[0034] The multiplexer units, basic gate circuit units and register units in the netlist are deeply optimized, useless logic units and line nets are removed, and equivalent logic units and line nets are merged;
[0035] Converting each multi-bit wide input logic unit in the netlist into a set of unit width input logic units;
[0036] Process mapping is performed on all combinational logic units of the netlist, generating a series of lookup table units;
[0037] From the netlist, find the combination of the output signal of the lookup table unit and the input signal of the register unit, and merge them into a basic logic unit.
[0038] The control signal of the register unit includes: clock enable signal, asynchronous clear signal and asynchronous set signal.
[0039] Compared with the prior art, the beneficial effects of the present application mainly include:
[0040] The present application provides an FPGA parallel processing method based on netlist division, which reduces the time required for application circuit synthesis through concurrent execution, further improving the running speed of FPGA synthesis.
[0041] In order to make the above and other objects, features and advantages of the present application more apparent and easy to understand, the following preferred embodiments are specifically described below, and the accompanying drawings are used for detailed description. BRIEF DESCRIPTION OF DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the specific embodiments of the present application, the drawings needed to be used in the description of the embodiments will be briefly introduced. Obviously, the drawings described below only constitute some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative effort on the basis of these drawings.
[0043] Figure 1 A flowchart of an FPGA parallel processing method based on netlist partitioning provided for an embodiment of the present application Figure 1 .
[0044] Figure 2 A flowchart of an FPGA parallel processing method based on netlist partitioning provided for an embodiment of the present application Figure 2 .
[0045] Figure 3 A flowchart of an FPGA parallel processing method based on netlist partitioning provided for an embodiment of the present application Figure 3 .
[0046] Figure 4 A flowchart of an FPGA parallel processing method based on netlist partitioning provided for an embodiment of the present application Figure 4 . DETAILED DESCRIPTION
[0047] The foregoing and other technical contents, features and effects of the present application will be clearly presented in the following detailed description of a preferred embodiment in cooperation with the drawings. The directions mentioned in the following embodiments, such as up, down, left, right, front or back, etc., are only the directions of the drawings. Therefore, the directions used are used to illustrate and not to limit the present application.
[0048] The embodiments of the present application will be described in detail below in conjunction with the drawings. However, those skilled in the art can understand that in the embodiments of the present application, many technical details are presented in order to make the reader better understand the present application. However, the technical solutions claimed by the present application can be realized even without these technical details and various changes and modifications based on the following embodiments.
[0049] Embodiment:
[0050] As shown in Figure 1 , an FPGA parallel processing method based on netlist partitioning comprises:
[0051] Obtaining the number N of current computer CPU cores and generating N threads T1, T2, …, TN;
[0052] According to the number N of threads, the netlist of the current circuit is partitioned, and the sub-netlist Wi corresponding to each thread i is calculated.
[0053] Start to execute each thread, in each thread i, process the sub-netlist table Wi until each thread is executed to the end;
[0054] After each thread is executed to the end, merge the logic cells and signals of each thread's sub-netlist table into the netlist table W;
[0055] Process the merged netlist table W to generate a synthesized netlist file, and output all the logic cells and signals of the netlist table W to the synthesis result file.
[0056] As shown in FIG. 1, divide the netlist of the current circuit, and calculate the sub-netlist table Wi processed by each thread i. Figure 2
[0057] Specifically, first analyze the circuit netlist, analyze the top-level module M0 and the sub-module list M1, M2, …, Mm called by the top-level module M0; then analyze the m+1 modules M0, M1, M2, …, Mm, and count the number of logic cells contained in each module; then sort the m+1 modules in descending order of the number of logic cells to obtain the sorted module list M0, M1, M2, …, Mm; add each sorted module to the queue Q in turn; then perform iterative processing, and take the first module from the queue Q and add it to the sub-netlist table W1 of the thread T1;
[0058] Then, process the threads Ti, i=2…N as follows:
[0059] Continuously take the first module from the queue Q and add it to the sub-netlist table Wi of the thread Ti until the total number of logic cells contained in the sub-netlist table Wi >= the total number of logic cells contained in all the modules in the sub-netlist table W1; this process continues until the processing of the Nth thread or the queue Q is empty; if the current queue Q is not empty, perform the next round of processing of the threads T1, T2, …, TN, and the iteration ends when the queue Q is empty.
[0060] In addition, the present application deletes the threads that do not contain any modules after the iteration ends.
[0061] Specifically, analyze the circuit netlist, analyze the top-level module M0 and the sub-module list M1, M2, …, Mm called by the top-level module M0;
[0062] Analyze the m+1 modules M1, M2, …, Mm, and count the number of logic cells contained in each module;
[0063] Sort the m+1 modules in descending order of the number of logic cells to obtain the sorted module list M1, M2, …, Mm;
[0064] add each sorted module to the queue Q in turn;
[0065] put the first module from the queue Q into the subnet table Wi of thread Ti;
[0066] for thread Ti, i = 2…N, perform the following operations:
[0067] if the queue Q is empty, delete the thread that does not contain any module, otherwise put the first module from the queue Q into the subnet table Wi of thread Ti;
[0068] and determine whether the total number of logic units contained in the subnet table Wi is less than the total number of logic units contained in all modules in the subnet table W1, if so, return to continue to determine whether the queue Q is empty, until the total number of logic units contained in the subnet table Wi is greater than or equal to the total number of logic units contained in all modules in the subnet table W1;
[0069] This process continues until the processing of the Nth thread or the queue Q is empty; if the current queue Q is not empty, the next round of processing of threads T1, T2, …, TN is performed, and the end condition of iteration is that the queue Q is empty;
[0070] If the queue Q is empty, the iteration process ends, and the thread that does not contain any module is deleted.
[0071] As shown in Figure 3 , start to execute each thread described above, and in each thread i, process the subnet table Wi.
[0072] Specifically, first extract the multiplexer unit (Mux), register unit (DFF), carry chain unit (Carry), memory unit (Bram) and digital processing logic unit (Dsp) of each module in the subnet table Wi; then perform rough optimization on the multiplexer unit, basic gate unit and register unit described above, remove useless logic units, and merge equivalent logic units; process each carry chain unit described above to generate a set of combinational logic block units of carry chain function; next, process each memory unit described above to convert each memory unit into a set of storage basic block units (Ram_block); finally, process each digital processing logic unit described above to convert each digital processing logic unit into a set of basic multiplication units (Mac_mult) and multiplication-accumulation output units (Mac_out).
[0073] The process of processing the merged netlist W is as follows Figure 4As shown, firstly, each register unit in the netlist W is processed to extract the control signals of each register unit, such as clock enable signal, asynchronous clear signal, asynchronous set signal and the like; then, the multiplexer units, basic gate circuit units and register units in the netlist W are optimized in depth to remove useless logic units and wire nets, and to merge equivalent logic units and wire nets; next, each multi-bit width input logic unit in the netlist W is converted into a group of unit width input logic units; then, all the combinational logic units of the netlist W are processed mapping, to generate a series of lookup table (LUT) units; finally, from the netlist W, the combinations whose output signals and register unit input signals are the same are found and merged into a basic logic unit (LE).
[0074] In summary, the application provides an FPGA parallel processing method based on netlist division, which reduces the time required for application circuit synthesis through concurrent execution, and further improves the running speed of FPGA synthesis.
[0075] Some commonly used English names or letters used for the convenience of clear description in the application are only used for exemplary reference and are not limited to the protection scope of the application.
[0076] It should also be noted that, in this paper, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between the entities or operations.
Claims
1. A parallel processing method for FPGAs based on netlist partitioning, characterized in that, include: Step S1: Generate N threads based on the number of CPU cores N in the computer; Step S2: Divide the total netlist of the current circuit according to the N threads, and calculate the subnetlist processed by each thread; Step S3: Process the subnetlist in each thread, obtain the logical units and signals of each thread's subnetlist, and merge them into the total netlist; Step S4: Process the merged netlist and output all logic units and signals to the synthesis result file; The specific process of step S2 includes: Obtain the top-level module of the overall circuit netlist and the list of sub-modules called by the top-level module; The top-level module and the list of sub-modules are analyzed to obtain the number of logical units contained in each module; Sort each module in descending order of the number of logical units, and add the sorted modules to the queue in sequence. The queue is iteratively processed until it is empty; The iterative processing procedure includes: Take the first module from the queue and put it into the subnet list W1 of thread T1, set the index i, and determine whether the queue is empty; If the queue is not empty, the first module is continuously taken out of the queue and put into the subnet list Wi of thread Ti until the total number of logical units contained in subnet list Wi is greater than or equal to the total number of logical units contained in all modules in subnet list W1, until the processing of the Nth thread or the queue Q is empty. If the current queue Q is not empty, the next round of processing of threads T1, T2...TN will be carried out. The iteration ends when queue Q is empty. If queue Q is empty, the iterative process ends, and threads that do not contain any modules are deleted.
2. The FPGA parallel processing method based on netlist partitioning according to claim 1, characterized in that, After the queue iteration process is completed, the thread that does not contain any modules is deleted.
3. The FPGA parallel processing method based on netlist partitioning according to claim 1, characterized in that, Set the index i to 2 to N.
4. The FPGA parallel processing method based on netlist partitioning according to claim 1, characterized in that, The specific process of processing the subnet table in each thread in step S3 includes: Extract the logical units of each module in the subnet table and perform rough optimization, removing useless logical units and merging equivalent logical units; Process the carry chain unit in each module to generate a set of combinational logic block units with carry chain function; The memory units in each module are processed, and each memory unit is converted into a set of basic storage block units; The digital processing logic units in each module are processed, and each digital processing logic unit is converted into a set of basic multiplication units and multiply-accumulate output units.
5. The FPGA parallel processing method based on netlist partitioning according to claim 4, characterized in that, The extracted logic units include: multiplexer unit, register unit, carry chain unit, memory unit, and digital processing logic unit.
6. The FPGA parallel processing method based on netlist partitioning according to claim 5, characterized in that, The multiplexer unit, basic gate circuit unit, and register unit are roughly optimized by removing useless logic units and merging equivalent logic units.
7. The FPGA parallel processing method based on netlist partitioning according to claim 1, characterized in that, The specific process of processing the merged netlist includes: Each register cell in the netlist is processed to extract the control signals for each register cell; Deep optimization is performed on the multiplexer units, basic gate units, and register units in the netlist, removing useless logic units and nets, and merging equivalent logic units and nets. Convert each multi-bit wide input logic unit in the netlist into a single-bit wide input logic unit; Perform process mapping on all combinational logic units in the netlist to generate a series of lookup table units; Find the combination of the same output signal of the lookup table unit and the same input signal of the register unit from the netlist, and combine them into a basic logic unit.
8. The FPGA parallel processing method based on netlist partitioning according to claim 7, characterized in that, The control signals for the register unit include: clock enable signal, asynchronous clear signal, and asynchronous preset signal.
Citation Information
Patent Citations
Multi-thread processing method and device, computer equipment and storage medium
CN117193986A