Circuit-partitioning-free method and system for FPGA parallel technology mapping
Patent Information
- Application Number
- PCT/CN2025/141828
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-25
- Filing Date
- 2025-12-11
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025141828_01102026_PF_FP_ABST
Abstract
Description
FPGA Parallel Technology Mapping Method and System Without Circuit Partitioning Technical Field
[0001] This application relates to the field of integrated circuit design, and in particular to the field of FPGA technology mapping and optimization. Background Technology
[0002] In the FPGA circuit design flow, technology mapping is a crucial step. Its goal is to map technology-independent logic circuits into circuits implementable on the FPGA using a given lookup table technology library. As the scale of FPGA applications continues to expand, improving technology mapping efficiency while ensuring mapping quality has become a significant challenge.
[0003] Existing technical mapping methods are mainly divided into two categories: single-threaded methods and multi-threaded parallel methods. Single-threaded technical mapping methods perform topological sorting of circuit nodes, then sequentially enumerate the cuts on each node and select the optimal cut to complete the mapping. The advantage of this method is that it can consider circuit optimization from a holistic perspective, avoiding the introduction of additional circuit structures. However, its main disadvantage is that it has low computational efficiency and is time-consuming for large-scale circuits.
[0004] To address the inefficiency of single-threaded methods, the industry has proposed parallel technical mapping methods based on multi-threading. These methods first divide the entire circuit into several independent sub-circuits, then assign a thread to each sub-circuit for technical mapping, with each thread running in parallel to improve computational efficiency. However, this circuit-partition-based parallel method has the following problems:
[0005] In the process of circuit partitioning, in order to obtain sub-circuits of similar size to improve parallel efficiency and ensure the correct function of each sub-circuit, it is often necessary to introduce additional circuit structures at the partitioning boundary, which increases the area overhead of the final circuit.
[0006] When a circuit is divided into independent sub-circuits, it becomes difficult to perform technical mapping optimization from a holistic perspective. In particular, when the partitioning process cuts off the timing-critical path in the circuit, it may lead to a significant deterioration in the timing performance of the final circuit.
[0007] The circuit partitioning and subsequent sub-circuit splicing process also require a certain amount of computation time, and this overhead may offset the efficiency improvement brought by parallel computing.
[0008] Therefore, how to achieve efficient parallel technology mapping without circuit partitioning, while maintaining the ability to optimize the overall circuit, has become a key problem that urgently needs to be solved in the field of FPGA technology mapping. This requires a new method to balance the relationship between parallel computing efficiency and circuit optimization quality. Summary of the Invention
[0009] The purpose of this application is to provide an FPGA parallel technology mapping method and system that does not require circuit partitioning, so as to solve the problems mentioned in the background art.
[0010] This application discloses a parallel technology mapping method for FPGAs without circuit partitioning, comprising the following steps:
[0011] Obtain the number of output ports N and the number of available threads T of the circuit to be mapped. Based on the structural characteristics of the transmission-type fan-in cone circuit corresponding to each output port, divide the output ports into T groups according to similarity. Each group contains at least one output port to maintain the integrity of the entire circuit to be mapped in the technical mapping process.
[0012] Based on the transmission-type fan-in cone circuit of each output port, determine its node topology sequence as the execution order of subsequent technology mapping;
[0013] Each output port is assigned an independent processing thread. The processing thread performs enumeration cuts node by node based on the corresponding node topology sequence and obtains the optimal cut that satisfies the optimization objective.
[0014] Based on the inverse topological order of the entire circuit to be mapped, the optimal cut on the node is selected to cover the entire circuit's AIG netlist, thus obtaining the final technical mapping circuit.
[0015] In a preferred embodiment, the step of grouping the output ports by similarity includes: calculating a distance metric between any two output ports i and j based on the node set of the transmission-type fan-in cone circuit corresponding to each output port.
[0016] Where Si and Sj represent the node sets of the transmission fan-in cone circuits corresponding to output ports i and j, respectively, and the smaller the distance metric value, the higher the similarity. The k-means clustering algorithm is used to cluster the output ports into T groups based on the distance metric.
[0017] In a preferred embodiment, when the number of output ports N is less than the number of available threads T, each output port is grouped into a separate group.
[0018] In a preferred embodiment, a depth-first search method is used to generate the topological sequence, and backtracking ensures that all predecessor nodes precede their successors.
[0019] In a preferred embodiment, each output port is assigned an independent processing thread. The processing thread performs enumerated cuts node by node based on the corresponding node topology sequence to obtain the optimal cut that satisfies the optimization objective: initialize all nodes to an "unvisited" state; before a thread accesses a node, check its locking status and ensure that all input nodes have been visited; lock unlocked nodes, and if other threads have already locked them, wait or select other processable nodes; after mapping is completed, unlock the node and mark it as "visited", and notify other threads waiting for that node.
[0020] In a preferred embodiment, the selection of the optimal cut is based on at least one of the following optimization objectives:
[0021] • Maximum circuit delay;
[0022] • Logical series;
[0023] • Number of times the lookup table is used.
[0024] In a preferred embodiment, during the parallel technique mapping process, nodes on the timing critical path are identified and marked, and the fragmentation of the critical path is reduced during thread allocation to optimize the timing performance of the circuit.
[0025] In a preferred embodiment, multiple rounds of iterative optimization are implemented in each processing thread, with each round dynamically adjusting the enumeration cut strategy to adapt to the current circuit optimization objective.
[0026] In a preferred embodiment, the circuit mapping result is obtained based on the inverse topology of the entire circuit to optimize the timing performance, area utilization, and power consumption of the circuit.
[0027] This application also discloses an FPGA parallel technology mapping system that does not require circuit partitioning, including:
[0028] The output port grouping module is used to obtain the number of output ports N and the number of available threads T of the circuit to be mapped. Based on the structural characteristics of the transmission type fan-in cone circuit corresponding to each output port, the output ports are divided into T groups according to similarity. Each group contains at least one output port to maintain the integrity of the entire circuit to be mapped in the technical mapping process.
[0029] The topology sequence determination module is used to determine the node topology sequence of each group of output ports of the transmission-type fan-in cone circuit, so as to serve as the execution order of subsequent technology mapping.
[0030] The parallel technology mapping module is used to allocate an independent processing thread to each group of output ports. The processing thread performs enumeration cuts node by node based on the corresponding node topology sequence to obtain the optimal cut that satisfies the optimization objective.
[0031] The mapping result generation module is used to select the optimal cut on the node based on the inverse topological order of the entire circuit to be mapped, and obtain the final technical mapping circuit.
[0032] This application provides an FPGA parallel mapping method that does not require circuit partitioning, breaking through the dependence on pre-partitioning of the circuit structure in existing technologies and effectively solving the limitations of traditional parallel mapping methods in terms of computational efficiency, timing optimization, and resource utilization. The technical solution of this application, while maintaining the integrity of the overall circuit structure, achieves greater efficiency, flexibility, and global optimization capabilities in the FPGA circuit mapping process through the coordinated efforts of intelligent grouping strategies based on transmission-type fan-in cone circuits at circuit output ports, topology sorting to control the parallel computation order, thread synchronization management using node locking mechanisms, and dynamic adjustment of multi-round optimization strategies. Specifically, this application brings the following technical effects:
[0033] First, it eliminates the need for circuit partitioning, avoiding the introduction of additional circuit structure overhead and improving mapping efficiency. Specifically, traditional parallel mapping methods rely on first partitioning the circuit and then mapping each sub-circuit separately. This method often requires copying parts of the circuit structure during partitioning, leading to additional circuit hardware resource consumption. Furthermore, the partitioned sub-circuits are difficult to optimize globally, affecting the final mapping effect. This application partitions parallel computing units directly based on the transmission-type fan-in cone structure of the FPGA circuit's output port. It achieves efficient parallel computing without physically partitioning the circuit structure, avoiding the additional circuit hardware overhead caused by partitioning, while ensuring global circuit optimization capabilities.
[0034] Furthermore, this method employs an intelligent grouping strategy based on k-means clustering to improve parallel computing efficiency. Specifically, traditional methods divide parallel computing units based solely on simple region partitioning or static grouping, which may lead to unbalanced computational load and affect overall computational efficiency. This application utilizes the k-means clustering algorithm to perform intelligent grouping based on the transmission-type fan-in cone circuit structure of the output port. This results in high computational correlation among nodes within the same group and low computational dependency between different groups, thereby maximizing the computational independence between threads, reducing synchronization wait, and increasing parallel computing throughput.
[0035] Furthermore, topological sorting is employed to optimize the computation order and ensure the correctness of data dependencies. Specifically, during parallel mapping, if the computation order is unreasonable, a situation may arise where the predecessor node has not been completed while the successor node has already begun computation, leading to data errors. This application employs a depth-first search (DFS)-based topological sorting algorithm within each parallel computing unit to sort the internal nodes of the transmission-type fan-in cone circuit at the circuit output port. This ensures that all predecessor nodes of each node are computed before it, guaranteeing the correctness of the computational data and avoiding data conflicts in parallel computing.
[0036] Furthermore, a node locking mechanism is introduced to improve the stability and consistency of parallel computing. This application proposes a thread synchronization mechanism based on node locking, which means that when a thread is processing a node, other threads cannot modify the node simultaneously, and the thread is only allowed to perform technical mapping calculations after all the predecessor nodes of the node have completed their calculations. This ensures data consistency, avoids thread conflicts, and improves the stability of parallel computing.
[0037] Furthermore, a multi-round iterative optimization strategy comprehensively improves circuit performance. Specifically, traditional technical mapping methods often employ a one-time optimal cut selection strategy, lacking global optimization capabilities and easily leading to local optima while resulting in global suboptimal outcomes. This application implements a multi-round optimization strategy in each thread's technical mapping process. After the first round of mapping, the enumeration cut strategy is dynamically adjusted based on the current circuit structure, and the optimal cut is reselected. This progressively optimizes the mapping results in subsequent rounds, achieving a comprehensive improvement in circuit area, logic levels, lookup table efficiency, and timing performance.
[0038] Furthermore, intelligent optimization of timing critical paths improves circuit operating speed. Specifically, traditional parallel mapping methods may fragment critical paths during circuit partitioning, leading to a decline in global timing performance. This application automatically identifies timing critical paths during the mapping process, ensuring that nodes on critical paths are not partitioned into different computing units by multiple threads, thereby reducing critical path fragmentation, increasing the maximum operating frequency of the final FPGA circuit, and optimizing the circuit's timing performance.
[0039] Furthermore, the reverse topological order optimization of the mapping results ensures optimal performance of the final mapped circuit. Specifically, traditional methods may employ local optimization during mapping result integration, neglecting global mapping structure optimization, which limits the performance of the final generated circuit. In the mapping result integration stage, this application selects the optimal cut of each node sequentially based on the reverse topological order, ensuring that the final mapped circuit is as close as possible to the optimal state in multiple dimensions such as area, power consumption, and timing.
[0040] In summary, this application constructs an efficient, stable, and circuit-partition-free FPGA parallel mapping method through a series of innovative technologies, including intelligent grouping of output ports, topology sorting to optimize computational order, node locking thread synchronization mechanism, multi-round optimization strategy, and timing critical path optimization. Compared with existing methods, the proposed solution significantly improves the computational efficiency of FPGA mapping, reduces the additional overhead caused by circuit partitioning during the mapping process, and ensures that the final mapping result is as close as possible to the optimal solution in terms of circuit area, lookup table utilization efficiency, logic levels, and timing performance. Furthermore, the technical solution of this application is applicable to multiple application scenarios such as high-performance computing, artificial intelligence acceleration chips, communication processors, and image processors, and has broad engineering application value and industrialization potential.
[0041] The specification of this application contains numerous technical features distributed across various technical solutions. Listing all possible combinations of these technical features (i.e., technical solutions) would make the specification excessively lengthy. To avoid this problem, the various technical features disclosed in the above-described invention, the various technical features disclosed in the following embodiments and examples, and the various technical features disclosed in the accompanying drawings can be freely combined to form various new technical solutions (all of which are considered to have been described in this specification), unless such a combination of technical features is technically infeasible. For example, one example discloses feature A+B+C, and another example discloses feature A+B+D+E. Features C and D are equivalent technical means that serve the same function, and technically only one needs to be used; they cannot be used simultaneously. Feature E can technically be combined with feature C. Therefore, the solution A+B+C+D should not be considered as described because it is technically infeasible, while the solution A+B+C+E should be considered as described. Attached Figure Description
[0042] Figure 1 is a flowchart illustrating an FPGA parallel mapping method without circuit partitioning according to a first embodiment of this application.
[0043] Figure 2 is a schematic diagram of the structure of an FPGA parallel technology mapping system without circuit partitioning according to the second embodiment of this application.
[0044] Figure 3 is a schematic diagram of a parallelization technique mapping method for FPGA without prior circuit partitioning according to an embodiment of this application.
[0045] Figure 4 is a schematic diagram of the technology mapping algorithm for each thread that is executed in parallel as described in step 7 of Figure 3.
[0046] Figure 5 is an NAND diagram of an example circuit according to an embodiment of this application. Detailed Implementation
[0047] In the following description, many technical details are presented to help the reader better understand this application. However, those skilled in the art will understand that the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments.
[0048] Explanation of some concepts:
[0049] FPGA (Field Programmable Gate Array): A type of reconfigurable integrated circuit that allows users to define its logic function through programming after manufacturing. It is widely used in digital circuit design, signal processing, and embedded systems.
[0050] Technology mapping: In logic synthesis, the process of converting a technology-independent logic representation into a specific hardware structure (such as an FPGA lookup table LUT) to ensure that the circuit can be implemented efficiently.
[0051] Look-Up Table (LUT): The basic logic unit of an FPGA, used to store and execute specific logic functions. It can perform arbitrary Boolean logic operations and is a core component of FPGA circuits.
[0052] And-Inverter Graph (AIG): A directed acyclic graph (DAG) used to represent logic circuits, where nodes represent AND gates and the negation properties of edges represent NOT gates. It is widely used for logic optimization and technology mapping.
[0053] Transitive Fan-in Cone (TFI Cone): For a given target node, the sub-circuit formed by all reachable predecessor nodes of that node reflects the logical input relationship of that node.
[0054] Topological Sorting: For all nodes in a directed acyclic graph (DAG), construct a linear sequence such that all predecessor nodes of each node are placed before that node, ensuring that the computation order conforms to the logical signal flow.
[0055] Reversed Topological Order: The reverse of the topological sorting, ensuring that each node precedes all its successors. It is often used in bottom-up circuit optimization and technology mapping steps.
[0056] Cut: In the LUT mapping process, a sub-circuit is divided for the target node, such that all paths from the circuit input port to the target node pass through at least one node in the sub-circuit. This cut set is used for LUT mapping optimization.
[0057] Cut Enumeration: In the process of technical mapping, the process of generating all possible cuts that meet the requirements by traversing the cut set on the input node of a node is used to determine the optimal cut scheme.
[0058] Best Cut: Among all feasible cuts, the best cut is selected based on the optimization objective (such as minimum latency, minimum number of logic levels, or minimum LUT usage) and is used as the final mapping scheme.
[0059] k-means clustering algorithm: an unsupervised learning algorithm that divides a dataset into k clusters by calculating the distance between data points. This application uses this method to group the output ports of FPGA circuits to optimize the thread allocation for parallel computing.
[0060] Node Locking Mechanism: In parallel computing, a synchronization mechanism is used to prevent multiple threads from accessing and modifying the same node at the same time. That is, when one thread is performing mapping calculations on a node, other threads cannot modify the node at the same time, so as to ensure data consistency and calculation correctness.
[0061] Multi-Round Optimization Strategy: During the technology mapping process, multiple rounds of iterative optimization are performed based on the current mapping results. The mapping strategy is adjusted in each round to gradually optimize the circuit area, logic levels, and timing performance.
[0062] The timing critical path is the path with the longest signal propagation delay in a circuit, which directly affects the circuit's maximum operating frequency. It needs to be optimized first during the technology mapping process to improve the timing performance of the circuit.
[0063] Parallel computing: Dividing a computational task into multiple subtasks and having them executed simultaneously by multiple threads or processing units to improve computational efficiency. This application achieves parallelization of technology mapping through reasonable task division and thread management.
[0064] Weighted Optimization Objective: During the technology mapping process, different weights are assigned to different optimization objectives (such as timing optimization, area optimization, and lookup table utilization optimization) according to the circuit design requirements, and the optimal solution is selected by comprehensively considering the weight factors.
[0065] The following is a brief summary of some of the innovative aspects of this application:
[0066] In summary, existing methods for FPGA mapping typically employ a single-threaded mapping strategy, resulting in computational efficiency that fails to meet the practical requirements of large-scale circuits. Parallel computing-based optimization methods often rely on circuit partitioning preprocessing, which not only introduces additional circuit structure overhead but may also lead to timing performance degradation due to the fragmentation of critical paths, thus limiting the overall circuit optimization space. This application proposes an FPGA parallel mapping method that eliminates the need for circuit partitioning, overcoming the dependence on circuit structure preprocessing in existing technologies. By maintaining the integrity of the overall circuit structure to be mapped, parallel computing units are constructed based on the transmission-type fan-in cone circuits at the output ports. The k-means clustering algorithm is then used to group the output ports to fully exploit the independence of each part within the circuit, achieving the optimal thread partitioning strategy.
[0067] Building upon this foundation, this application further linearizes the logical relationships of each computational unit through a topological sorting method. This ensures that the mapping process strictly follows the signal propagation direction of the circuit, thereby eliminating the distortion of computational results caused by the fragmentation of the critical path in traditional parallel mapping methods. Simultaneously, to ensure synchronization and coordination among multiple threads during parallel computing, this application employs a node locking mechanism. Each thread locks the node it needs to compute during the mapping process, leveraging the independence between different parts of the specific circuit structure to maximize parallel processing and effectively improve the overall throughput of parallel computing. Furthermore, this application implements a multi-round iterative optimization strategy, adjusting the enumeration cut strategy and optimal cut selection criteria at each node. This ensures that the final mapping result balances multiple optimization objectives, including circuit area, logic levels, lookup table utilization, and timing performance, thereby achieving global optimization of the mapping scheme.
[0068] Compared to existing methods, the technical solution proposed in this application not only significantly improves the technical mapping efficiency of FPGA circuits and avoids the additional structural overhead caused by circuit partitioning, but also maintains a global perspective during the technical mapping process, ensuring that the final generated circuit is as close as possible to the optimal state in multiple dimensions such as timing optimization, logic optimization, and area optimization. In particular, the synergistic cooperation of the transmission-type fan-in cone grouping strategy, topology sorting control of parallel computation order, node locking-based thread synchronization mechanism, and multi-round iterative optimization strategy proposed in this application enables the mapping process to obtain an optimized FPGA logic implementation scheme with minimal computational cost, thereby overcoming the circuit performance loss problems caused by circuit partitioning, thread competition, and critical path interruption in existing parallel mapping methods. Therefore, the solution proposed in this application not only provides a more efficient FPGA technical mapping method in theory, but also significantly improves the mapping efficiency and final implementation performance of large-scale FPGA circuits in practical applications, possessing extremely high engineering application value and innovation.
[0069] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0070] The first embodiment of this application relates to an FPGA parallel technology mapping method without circuit partitioning, the process of which is shown in Figure 1. The method includes the following steps:
[0071] Step 100: Output Port Grouping
[0072] The number of output ports N and the number of available threads T of the circuit to be mapped are obtained. Based on the structural characteristics of the transmission-type fan-in cone circuit corresponding to each output port, the output ports are divided into T groups according to similarity. Each group contains at least one output port to maintain the integrity of the entire circuit to be mapped during the technical mapping process. Optionally, for example, the T groups can be divided according to the overlap of the node sets of their transmission-type fan-in cone circuits. The node set overlap is defined as the ratio of the intersection size to the union size of the node sets of their transmission-type fan-in cone circuits for two output ports i and j, i.e., |S_i∩S_j| / |S_i∪S_j|. The larger this value, the more similar the circuit structures of the two output ports are.
[0073] Step 200: Determine the topological order
[0074] Based on the transmission-type fan-in cone circuit of each output port, its node topology sequence is determined as the execution order of subsequent technology mapping.
[0075] Step 300: Parallel technology mapping
[0076] Each output port is assigned an independent processing thread. The processing thread performs enumeration cuts node by node based on the topological sequence of the nodes in the corresponding output port transmission type fan-in cone circuit to obtain the optimal cut that satisfies the optimization objective at that node.
[0077] In other words, an independent processing thread is assigned to each group of output ports. Based on the corresponding node topology sequence, the processing thread performs enumerated cuts on each node it is responsible for, selecting and obtaining the optimal cut that satisfies the optimization objective for each node. The threads work together, jointly affecting the AIG graph of the entire circuit, with each node being processed only once by one thread.
[0078] Furthermore, in the parallel technical mapping process, each thread does not independently generate a complete mapping scheme. Instead, based on the global topological order, it only performs enumerated cuts on the nodes it is responsible for and selects the optimal cut. This means that the technical mapping of the entire AIG graph is completed collaboratively by all threads, rather than each thread independently calculating multiple complete mapping schemes and then merging them. Therefore, each node is processed only once by one thread throughout the entire process, ensuring the uniqueness and consistency of the mapping results.
[0079] To ensure the correctness and efficiency of parallel execution, multiple threads are prevented from accessing the same node simultaneously during thread allocation. Specifically, before accessing a node, a thread checks its dependencies to ensure that all its predecessor nodes have been mapped. Furthermore, a node locking mechanism is employed, preventing other threads from accessing the same node while one thread is processing it, thus avoiding data races and uncertainty in mapping results. After completing the mapping of a node, each thread releases it, allowing subsequent threads that depend on that node to continue their mapping tasks.
[0080] The advantage of this mapping method is that it avoids the traditional multi-threaded mapping method of dividing the circuit into multiple sub-circuits, thereby reducing additional circuit structure overhead and better optimizing the overall circuit mapping quality. Furthermore, due to the use of global topological order, each thread executes in an orderly manner according to data dependencies, ensuring data consistency by preventing the problem of the predecessor node not being fully calculated while the successor node has already started calculation.
[0081] Ultimately, after all threads collaboratively complete the mapping across the entire AIG graph, the result is a set of optimal cuts based on each node, rather than a merging of multiple complete mapping schemes. This ensures that the technical mapping result achieves optimal performance across multiple optimization objectives, such as logic level, lookup table utilization, timing performance, and circuit area.
[0082] Step 400: Generating Mapping Results
[0083] Based on the inverse topological order of the entire circuit to be mapped, the optimal cut at each node is selected to cover the entire circuit's AIG netlist, thus obtaining the final technical mapping circuit.
[0084] In other words, based on the inverse topological order of the entire circuit to be mapped, the optimal cuts determined by the thread at each node are selected sequentially to cover the entire circuit's AIG netlist, resulting in the final technically mapped circuit. Each optimal cut corresponds to a lookup table (LUT) implementation.
[0085] Furthermore, in the mapping result generation stage, the ultimate goal of technical mapping is to construct a complete FPGA logic circuit using the optimal cut of each node. Since each node is enumerated and its optimal cut is determined by only one thread during parallel computing, the final mapping result is not a merging of mapping schemes calculated by multiple threads individually. Instead, it involves progressively covering the entire circuit with the optimal cuts of all nodes in reverse topological order. This process ensures the consistency and optimization effect of the mapping scheme, resulting in high timing performance, logic optimization, and resource utilization in the final technical mapping result.
[0086] The use of reverse topological order for integrating mapping results is primarily to ensure consistency between the computation order and logical dependencies. In FPGA logic synthesis, the computation result of a node often depends on its predecessor nodes. Therefore, when integrating mapping results, processing must proceed in order from circuit input to output. This bottom-up approach ensures that when selecting the optimal cut for a node, the mapping schemes of all its predecessor nodes are already determined, thus avoiding data conflicts or suboptimal selections due to undetermined predecessor mapping schemes.
[0087] During the integration process, the optimal cut of each node directly corresponds to a LUT (Look-Up Table), which is responsible for implementing the logic function represented by that node. Since the LUT is the basic unit of FPGA logic implementation, one of the optimization goals of the mapping process is to minimize the number of LUTs, improve LUT utilization, and optimize circuit timing. In the integration phase of the mapping results, global optimization strategies can also be combined, such as prioritizing cuts with lower delays on critical paths to optimize the maximum operating frequency of the circuit, or selecting cuts with smaller areas on non-critical paths to reduce FPGA resource consumption.
[0088] Ultimately, the mapping result generation process is a global optimization process based on inverse topological order, which not only ensures the logical integrity and correctness of the entire circuit, but also optimizes FPGA resource utilization and circuit performance to the greatest extent.
[0089] Specifically, in order to realize the parallel technology mapping of FPGA circuits, the method proposed in this embodiment includes four main steps: output port grouping, topology order determination, parallel technology mapping, and mapping result generation.
[0090] First, in the output port grouping step, the method obtains the total number N of output ports of the circuit to be mapped and the number T of available threads in the system. The core idea of this step is to perform intelligent grouping based on circuit structural characteristics, rather than the physical partitioning of traditional methods. Specifically, the method analyzes the structural characteristics of the transmission-type fan-in cone circuit corresponding to each output port, grouping output ports with similar structural characteristics into the same group, ultimately forming T groups. This grouping method not only ensures that each group contains at least one output port, but also maintains the integrity of the entire circuit in the technical mapping process by avoiding physical partitioning of the circuit.
[0091] Next, in the topology order determination step, the method determines the node topology sequence of the corresponding transmission-type fan-in cone circuit for each group of output ports. This topology sequence will serve as the basis for the execution order of the subsequent technology mapping process. By establishing a clear topology sequence, it can be ensured that the technology mapping process proceeds in an orderly manner according to the logical dependencies of the circuits, avoiding circular dependencies or deadlocks. In particular, this topology order also ensures that when processing a certain node, all its input nodes have already been processed.
[0092] In the parallel mapping step, the method assigns an independent processing thread to each output port group. These threads process the nodes in the circuit one by one according to the previously determined topology sequence. During processing, each thread independently performs an enumeration cut operation and selects the optimal cut on the node based on preset optimization objectives (such as circuit delay, logic levels, or lookup table usage).
[0093] Finally, in the mapping result generation step, the method selects the optimal cuts already determined at the nodes based on the inverse topological order of the entire circuit to be mapped, thereby covering the entire circuit's AIG netlist and obtaining the technically mapped circuit result.
[0094] This step-by-step design ensures the efficiency of parallel computing while avoiding the performance loss caused by circuit partitioning in traditional methods. Through an intelligent grouping strategy based on the characteristics of transmission-type fan-in cone circuits and strict execution order control, this method can fully utilize the advantages of multi-threaded parallel computing while maintaining circuit integrity, achieving efficient and high-quality technology mapping.
[0095] The following provides a further explanation of each step in the above process.
[0096] Optionally, in step 100, the step of grouping the output ports by similarity includes: calculating the distance metric between any two output ports i and j based on the node set of the transmission-type fan-in cone circuit corresponding to each output port.
[0097] Where Si and Sj represent the node sets of the transmission fan-in cone circuits corresponding to output ports i and j, respectively, and the smaller the distance metric value, the higher the similarity. The k-means clustering algorithm is used to cluster the output ports into T groups based on the distance metric.
[0098] Alternatively, when the number of output ports N is less than the number of available threads T, each output port can be grouped into a separate group.
[0099] More specifically, in step 100, this embodiment proposes an intelligent grouping method for output ports based on similarity. The core of this method lies in determining the degree of correlation between output ports by analyzing circuit structure characteristics, thereby achieving the optimal grouping strategy.
[0100] First, the method introduces the concept of a Transmission Fan-In Cone (TFI cone). For any output port, its TFI cone includes all logic nodes that can affect that output port. For example, in a simple circuit, if output port O1 is affected by AND gates G1 and G2, then G1, G2, and their input nodes all belong to the TFI cone of O1.
[0101] To quantify the similarity between two output ports, the method innovatively proposes a distance metric formula: in:
[0102] Si and Sj represent the node sets of the transmission-type fan-in cone circuits corresponding to output ports i and j, respectively.
[0103] |Si∪Sj| represents the size of the union of two sets, i.e., the number of nodes in it.
[0104] |Si∩Sj| represents the size of the intersection of two sets, i.e., the number of nodes in it.
[0105] This distance metric has the following characteristics:
[0106] The distance metric is 1 (minimum) when the transmission fan-in cones of the two output ports are completely aligned.
[0107] The distance metric is maximized when the transmission fan-in cones of the two output ports are completely non-overlapping.
[0108] The smaller the distance metric, the higher the logical correlation between the two output ports.
[0109] For example, suppose there are three output ports O1, O2, and O3:
[0110] The transmission-type fan-in cone of O1 contains nodes {A,B,C,D}.
[0111] The transmission-type fan-in cone of O2 contains nodes {B,C,D,E}.
[0112] If the transmission-type fan-in cone of O3 contains nodes {F,G,H}, then:
[0113] D(O1,O2)=|{A,B,C,D,E}| / |{B,C,D}|=5 / 3
[0114] D(O1,O3)=|{A,B,C,D,F,G,H}| / |{}|=∞
[0115] Based on this distance metric, the method employs the k-means clustering algorithm to divide the output ports into T groups. The k-means algorithm, through iterative optimization, maximizes the similarity of output ports within each group while maximizing the dissimilarity between groups. This grouping method offers the following advantages:
[0116] Structural dependency optimization: Grouping output ports with high logical correlation into the same group and output ports with low logical correlation into different groups can reduce data dependencies between different threads and improve parallel efficiency.
[0117] Resource sharing optimization: Similar transmission-type fan-in cone circuits often share a lot of logic resources. Grouping them into the same group can reduce redundant calculations.
[0118] Load balancing: Through the adaptive adjustment of the k-means algorithm, a relative balance of workload among groups can be achieved while ensuring the rationality of grouping.
[0119] Specifically, when the number of output ports N is less than the number of available threads T, the method employs a simplification strategy: each output port is grouped into a separate group, resulting in a total of N groups. In this case, due to sufficient computing resources, there is no need to consider the similarity between ports, allowing for maximum parallelism to be achieved directly.
[0120] This intelligent grouping strategy ensures both the efficiency of parallel computing and the integrity of the circuit, avoiding the performance loss caused by physical partitioning in traditional methods. Furthermore, this structural feature-based grouping approach provides a solid foundation for subsequent parallel mapping, facilitating the acquisition of superior mapping results.
[0121] Optionally, in step 200, a depth-first search method is used to generate a topology sequence, and backtracking is used to ensure that all predecessor nodes precede their successors.
[0122] More specifically, in step 200, this embodiment employs a depth-first search (DFS) method to generate the topology sequence of circuit nodes, which is crucial for ensuring the correctness and efficiency of subsequent technical mapping processes. The DFS method starts with each set of output ports and progressively visits its input nodes along the logical connections of the circuit until an input port or a previously visited node is reached. By backtracking, the nodes are recorded in the order they were visited, and reversing this order yields the desired topology sequence.
[0123] For example, suppose a circuit has input ports A and B, output ports C and E, and an intermediate node D. A depth-first search starts at output ports E and C, visits D, and then visits A and B. The resulting access sequence is E->C->D->B->A. Reversing this sequence yields the topological sequence: A->B->D->C->E.
[0124] This method is computationally efficient, with a time complexity of O(V+E), where V is the number of nodes and E is the number of connections. Its space complexity is O(V), used to store access markers and the recursion stack. This allows the method to efficiently handle large-scale circuits.
[0125] This method supports parallel computing, allowing the topology sequence of each output port to be generated independently, and the generation process of topology sequences for different groups can be performed in parallel. This, combined with the intelligent grouping in step 100, greatly improves the efficiency of the overall mapping process.
[0126] Step 200 generates an independent topology sequence for each group of output ports, which, combined with the grouping results from step 100, provides the correct node processing order for the parallel mapping in step 300. Ultimately, the topology sequence generated in step 200 also ensures the integration of the mapping results in step 400.
[0127] By using this depth-first search-based topology sequence generation method, this embodiment not only ensures the correctness of the technology mapping process, but also provides a reliable execution order guarantee for subsequent parallel processing, thereby improving the efficiency and reliability of the entire technology mapping process.
[0128] Optionally, in step 300, the parallel technology mapping step includes: initializing all nodes to an "unvisited" state; before a thread accesses a node, checking its locking status and ensuring that all input nodes have been accessed; locking unlocked nodes, and if other threads have already locked them, waiting for or selecting other processable nodes; after mapping is completed, unlocking the node and marking it as "visited", and notifying other threads waiting for that node.
[0129] Further, optionally, the selection of the optimal cut is based on at least one of the following optimization objectives:
[0130] Maximum circuit delay;
[0131] Logical series;
[0132] Number of times the lookup table is used.
[0133] Further optionally, during the parallel technology mapping process, nodes on the timing critical path are identified and marked, and the fragmentation of the critical path is reduced during thread allocation to optimize the timing performance of the circuit.
[0134] Alternatively, multiple rounds of iterative optimization can be implemented in each processing thread, with each round dynamically adjusting the enumeration cut strategy to adapt to the current circuit optimization objective.
[0135] More specifically, in step 300, this embodiment designs a complete parallel technology mapping mechanism, which achieves efficient and reliable parallel processing through fine-grained thread coordination and multi-level optimization strategies. The core of this mechanism includes three main aspects: thread coordination, optimization target selection, and iterative optimization.
[0136] In terms of thread coordination, this mechanism employs a state management scheme. All circuit nodes are initialized to an "unaccessed" state before parallel processing begins. Each thread performs a series of checks when processing a node, including verifying the node's locking status to ensure it is not occupied by other threads, checking whether all input nodes of that node have been accessed, and locking the node after gaining processing rights. This multi-check mechanism effectively prevents data races and deadlocks. When a thread encounters a locked node, the system supports flexible processing strategies such as entering a waiting queue or switching to other processable nodes. This dynamic scheduling significantly improves parallel efficiency and avoids wasting thread resources. After completing the mapping, the thread executes a standard post-processing procedure, including unlocking the node, updating its state to "accessed," and notifying all threads waiting for that node.
[0137] In the process of selecting the optimal cut, the system comprehensively considers three key indicators: the maximum circuit delay, the number of logic levels, and the number of lookup tables used. By evaluating the impact of the cut scheme on the timing of the critical path, analyzing the resulting changes in logic depth, and statistically analyzing resource utilization, the system can generate a mapping scheme that achieves a balance in terms of timing, logic complexity, and resource utilization.
[0138] In terms of optimization strategy, a multi-round iterative optimization method is adopted. The first round uses a basic enumeration cut strategy to establish an initial mapping scheme. Subsequent rounds use different optimization objectives to adjust the cut enumeration strategy and optimize the mapping results. This progressive optimization method can effectively improve the overall performance of the circuit. The correctness of parallel processing is ensured through low-level thread coordination and state management, the quality assurance is provided by mid-level optimization objective synthesis, and the performance improvement is achieved through top-level iterative optimization strategies. The entire mechanism forms a complete multi-level collaborative working system. This mechanism has strong flexibility and scalability, and can adapt to FPGA circuit designs of different scales and performance requirements.
[0139] To ensure the global optimality of the final circuit during the process of obtaining the final circuit, this application employs an inverse topological order approach to perform optimal cut coverage of the entire circuit. Specifically:
[0140] First, by using the topological order of each group of output ports obtained in step 200, the topological order of the entire circuit can be quickly constructed, and the reverse topological order can be obtained by reversing the order.
[0141] Following this inverse topological order, the optimal cut of each node is selected to obtain the final circuit.
[0142] Furthermore, in step 400, this embodiment proposes a mapping result generation method based on inverse topological order, which aims to ensure that the final generated FPGA circuit achieves global optimization in multiple performance indicators.
[0143] Reverse topology ordering refers to sorting all nodes in the circuit from output port to input port, so that all successor nodes of each node precede it. This sorting method is opposite to the actual direction of signal propagation, but in the technical mapping process, it ensures that when processing a certain node, all dependent successor nodes have been mapped, thus providing a basis for optimization.
[0144] In the parallel mapping process of this application, the optimal cut of each node is calculated independently by different threads. To construct a complete circuit mapping, these scattered cut information need to be integrated in a reasonable order. Using inverse topological order ensures that when selecting the optimal cut of a node, the mapping schemes of all its successor nodes have been determined, thereby avoiding suboptimal selections due to improper mapping order.
[0145] For example, suppose node A is the common predecessor of nodes B and C, and B and C are processed by different threads. If the optimal cut of A is chosen before the mapping result of B and C is determined, the mapping scheme of A may not fully consider the optimization needs of its successor nodes, resulting in a suboptimal overall solution.
[0146] By using the reverse topological order, the optimal cut between B and C is determined first, and then the optimal cut between A is selected based on the mapping results of B and C. In this way, the mapping scheme of A can comprehensively consider the logical requirements of its successor nodes, thereby making the entire mapping scheme more globally optimized in terms of LUT (lookup table) utilization, logical levels, and timing performance.
[0147] Based on inverse topological ordering, this application employs a bottom-up approach to map the optimal cut of each node onto the FPGA's lookup table (LUT) to generate a complete mapped circuit. During the mapping process, these optimal cuts comprehensively consider multiple optimization objectives such as area, delay, and power consumption, achieving a global balance in circuit performance. Since the LUT is the basic unit of FPGA logic implementation, using inverse topological ordering can reduce the number of LUT stages, improve LUT resource utilization, and optimize circuit timing.
[0148] It is worth emphasizing that on the critical path, reverse topology ordering can prioritize the optimal cut with shallower logic depth and lower latency, thereby optimizing the circuit's maximum operating frequency. On non-critical paths, reverse topology ordering can appropriately select cuts with smaller areas to reduce FPGA resource consumption and further optimize the overall circuit performance.
[0149] The reverse topological order mapping method is not isolated, but closely coordinated and mutually supportive with the aforementioned intelligent grouping, topological sorting, and multi-threaded parallel mapping: Intelligent grouping (step 100): Improves computational independence between threads, reduces competition between threads, and enhances parallel computing efficiency; Topological sorting (step 200): Ensures that the correct data dependency order is followed during the technical mapping process, providing a computational basis for the reverse topological order; Parallel technical mapping (step 300): Pre-calculates the optimal cut for each node, so that step 400 only requires integration and optimization, without additional computation, thus improving efficiency. Based on this, the reverse topological mapping acts as a global coordinator, integrating the optimal cuts of each node to generate the final technical mapping scheme. This divide-and-conquer, layer-by-layer optimization strategy ensures computational efficiency while maximizing LUT resource utilization, circuit logic optimization, and timing performance, ultimately generating a highly efficient, globally optimized FPGA logic mapping scheme.
[0150] To better understand the technical solution of this application, a specific example is provided below. The details listed in this example are mainly for ease of understanding and are not intended to limit the scope of protection of this application.
[0151] This example proposes a novel parallel technology mapping method for FPGAs, which aims to accelerate the technology mapping stage of FPGA circuits while avoiding premature circuit partitioning. This avoids the performance degradation caused by partitioning and merging, and saves the resulting runtime overhead.
[0152] The flowchart for this example is shown in Figure 3. Figure 4 shows the specific flowchart of the technology mapping performed by each thread in parallel in step 7 of Figure 3. The flowchart in Figure 3 will be described in detail below.
[0153] Start: Input the FPGA circuit to be mapped, which is usually represented in the form of a NAND diagram.
[0154] Step 1: Group the output ports of the input circuit to be mapped: Determine if the number of output ports N is greater than the number of available threads T. If yes, proceed to Step 2; otherwise, proceed to Step 3.
[0155] Step 2: When the number of output ports N of the circuit is greater than the number of threads T, use the k-means clustering algorithm to divide the N output ports into T groups, with each group containing at least one output port. Then proceed to step 4.
[0156] Step 3: When the number of output ports N of the circuit is less than or equal to the number of threads T, divide the N output ports into N groups, with each group containing one output port. Then proceed to Step 4.
[0157] Step 4: For each set of output ports, find the topology sequence of the corresponding transmission-type fan-in cone circuit. Then proceed to Step 5.
[0158] Step 5: Assign a thread to each output port group to perform technical mapping on its corresponding transmission-type fan-in cone circuit.
[0159] Step 6: Set the state of all nodes in the circuit to "unvisited".
[0160] Step 7: Execute all threads in parallel to perform the technology mapping for the current round. Each round of mapping can be performed for a specific optimization objective, such as optimizing the maximum delay, number of logic levels, or circuit area. Then proceed to Step 8.
[0161] Step 8: Determine if all rounds of technology mapping have been completed. If completed, proceed to Step 9; otherwise, return to Step 6 and proceed to the next round of technology mapping.
[0162] Step 9: Based on the inverse topological order of the NAND graph nodes of the circuit, select the optimal cut for each node, cover the entire circuit, and finally obtain the circuit with completed technical mapping.
[0163] End: Output the final mapped circuit.
[0164] In step 2, when N is greater than T, a clustering algorithm is used to divide the N circuit output ports into T groups. This method improves the efficiency of parallel technology mapping by considering the correlation of the transmission-type fan-in cone circuits corresponding to each circuit output port and using the k-means clustering algorithm for optimized grouping. This grouping method will be explained in detail below.
[0165] The k-means clustering method is used to group the circuit output ports: First, a distance metric is defined between the circuit output ports. The set of nodes of the transmission-type fan-in cone for the i-th output port is defined as S_i. For the i-th and j-th output ports, the number of common nodes and the total number of nodes in their transmission-type fan-in cone circuits are S_i and S_j, respectively. i ∩S j |and S i ∪S j |, where |S i ∩S j |and S i ∪S j Let S_i∪S_j be the intersection of the two sets, and S_i∪S_j be the union of the two sets. The distance metric between two output ports can be defined as:
[0166] A larger value indicates a greater distance between two output ports, a lower percentage of shared nodes, and lower node overlap. Conversely, a smaller value indicates a closer distance between two output ports, a higher percentage of shared nodes, and higher node overlap. Then, the k-means clustering algorithm is used to cluster all output ports: First, T output ports are randomly selected as cluster centers. The distance between each output port and its cluster center is calculated using the formula above, and each output port is assigned to the nearest cluster center. After all output ports are assigned, the cluster centers are recalculated based on the distances between output ports within each cluster. This process is repeated until the number of output ports in each cluster no longer changes, meaning no output ports are reassigned to different clusters. At this point, the process stops, and the final output port groupings are obtained.
[0167] This clustering method allows for grouping based on the characteristics and interrelationships of the transmission fan-in cone circuits corresponding to the output ports. This results in higher overlap among the transmission fan-in cone circuits within each group, while maintaining lower overlap between different groups. This helps improve thread utilization efficiency during subsequent parallel technology mapping, thereby accelerating the entire technology mapping process.
[0168] The following describes the process of technology mapping for each thread running in parallel in step 7 of Figure 3, following the algorithm shown in Figure 4. Step 7 in Figure 3 is the process of multiple threads executing the algorithm in parallel.
[0169] Start: Input the topology sequence of the transmission fan-in cone circuit corresponding to the output ports of a set of circuits.
[0170] Step 701: Initialize, set the value of k to 1.
[0171] Step 702: Determine whether all nodes in the topological order have been processed. If yes, end the entire process; otherwise, proceed to step 703.
[0172] Step 703: Select the kth node as the current node according to the topological order, and proceed to step 704.
[0173] Step 704: Determine if the current node is in the "visited" state. If yes, proceed to step 705; otherwise, proceed to step 706.
[0174] Step 705: Increment k by 1, then return to step 702.
[0175] Step 706: Determine if the current node is locked by another thread. If yes, proceed to step 705; otherwise, proceed to step 7.
[0176] Step 7: The current thread locks the current node to prevent other threads from processing the node simultaneously. Proceed to Step 8.
[0177] Step 8: Determine whether both input nodes of the current node are in the "visited" state. If yes, proceed to step 710; if no, proceed to step 9 and wait until both input nodes become "visited", then proceed to step 710.
[0178] Step 710: Perform cut enumeration operation on the current node: pair the cut sets of the two input nodes of this node, add the cuts that meet the conditions to the cut set of the current node, sort these cuts according to the optimization strategy of the current mapping round, and select the optimal cut. Proceed to step 711.
[0179] Step 711: Unlock the current node and mark its status as "visited", then proceed to step 705.
[0180] End: After determining in step 702 that all nodes have been processed, the technical mapping process of this thread ends.
[0181] Next, the operation process of the proposed method will be explained with reference to the example circuit in Figure 5.
[0182] Assume the number of threads used is T = 2, and the number of circuit output ports is N = 3. First, the output ports are grouped. Using the k-means clustering algorithm, the output ports are grouped according to the node overlap of the corresponding transmission-type fan-in cone circuits, resulting in the first group containing output port po1, and the second group containing output ports po2 and po3. Then, thread T1 is assigned to the first group, and thread T2 is assigned to the second group.
[0183] Next, the node topology order of the transmission-type fan-in cone circuit is obtained for each group of output ports. The topology order for each group is obtained through a depth-first search:
[0184] Group 1 (po1): pi1,pi2,n1,pi3,n2,pi4,pi5,n3,pi6,n4,n5,n6,po1
[0185] The second group (po2,po3): pi4,pi5,n3,pi6,n4,n5,pi7,pi8,n7,n8,n9,n11,po2,n10,po3
[0186] This example assumes only one round of full-circuit mapping. If multiple rounds are performed, the algorithmic behavior for each round is similar to that described below.
[0187] Because parallel computing is used, the following shows the running status of different threads at different times. When each node is "enumerating cuts", it will form a cut through its own nodes.
[0188] Initially, the state of all nodes is set to "unvisited".
[0189] At time t1: T1 begins locking and processing pi1; T2 begins locking and processing pi4.
[0190] At time t2: T1 completes the processing of pi1 and unlocks it, sets its status to "accessed", and then begins locking and processing pi2; T2 completes the processing of pi4 and unlocks it, sets its status to "accessed", and then begins locking and processing pi5.
[0191] At time t3: T1 completes the processing of pi2 and unlocks it, sets its status to "accessed", and then begins locking and processing n1; T2 completes the processing of pi5 and unlocks it, sets its status to "accessed", and then begins locking and processing n3.
[0192] At time t4: T1 completes the processing of n1 and unlocks it, sets its status to "accessed", and then begins locking and processing pi3; T2 completes the processing of n3 and unlocks it, sets its status to "accessed", and then begins locking and processing pi6.
[0193] At time t5: T1 completes the processing of pi3 and unlocks it, sets its status to "accessed", and then begins locking and processing n2; T2 completes the processing of pi6 and unlocks it, sets its status to "accessed", and then begins locking and processing n4.
[0194] At time t6: T1 completes the processing of n2 and unlocks it, setting its state to "visited". Next, T1 jumps to pi4, finds pi4 has been visited, then jumps to pi5, finds pi5 has been visited, then jumps to n3, finds n3 has been visited, then jumps to pi6, finds it has been visited, then jumps to n4, finds it is locked by T2, jumps to n5, finds n5 has not been visited. At this point, T1 locks n5 and checks the two input nodes, finding n3 has been visited and n4 has not, therefore waiting for n4 to become "visited". T2 continues to lock and process n4.
[0195] At time t7: T1 discovers that n4 has been unlocked and marks it as "accessed", continues to lock and process n5; T2 unlocks n4 and marks it as "accessed", jumps to n5, finds it is locked by T1, continues to jump to pi7, and begins to lock and process.
[0196] At time t8: T1 continues to lock and process n5; T2 completes the processing of pi7 and unlocks it, sets its status to "accessed", and then begins to lock and process pi8.
[0197] At time t9: T1 continues to lock and process n5; T2 completes the processing of pi8 and unlocks it, sets its status to "visited", then jumps to n7, finds that n7 has not been visited and both input nodes have been visited, and begins to lock and process n7.
[0198] At time t10: T1 continues to lock and process n5; T2 completes the processing of n7 and unlocks it, sets its status to "visited", then jumps to n8, finds that n8 has not been visited and both input nodes have been visited, and begins to lock and process n8.
[0199] At time t11: T1 continues to lock and process n5; T2 completes the processing of n8 and unlocks it, sets its state to "visited", then jumps to n9, finds that n9 has not been visited and n5 has not been visited among the two input nodes, while n8 has been visited, so lock n9 and wait for n5 to become "visited".
[0200] At time t12: T1 completes the processing of n5 and unlocks it, sets its status to "visited", then jumps to n6, finds that it is not visited and both input nodes have been visited, locks n6 and starts processing; at T2, n5 has become visited, and both input nodes of n9 have been visited, so n9 is started to be processed.
[0201] At time t13: T1 continues to lock and process n6; T2 completes the processing of n9 and unlocks it, sets its status to "visited", then jumps to n11, finds that it has not been visited and that both input nodes have been visited, and begins to lock n11.
[0202] At time t14: T1 completes the processing of n6 and unlocks it, sets its status to "visited", then jumps to po1, finds that the topology order that T1 has completed is already completed, and ends; T2 continues to process n11.
[0203] At time t15: T2 completes the processing of n11 and unlocks it, sets its status to "accessed", then jumps to po2, finds that po2 is an output port, continues to jump to n10, finds that it is not accessed and that the two input nodes have been accessed, locks n10 and starts processing.
[0204] At time t16:T2 completes the processing of n10 and unlocks it, sets its status to "visited", then jumps to po3, finds that the topological order has been completed, and ends.
[0205] Finally, after all threads have completed, the reverse topological order is executed to select the optimal cut and cover, resulting in the technically mapped circuit.
[0206] In this process, the time period during which threads T1 and T2 execute in parallel (e.g., from time t7 to t12) demonstrates the efficiency of parallel computing. For example, while T1 processes n5, T2 processes nodes pi7, pi8, n7, and n8 in parallel, thus significantly improving the efficiency of the mapping.
[0207] The key to the above process lies in the development of a systematic algorithm that allows for parallel technology mapping across the entire FPGA circuit without prior circuit partitioning. Compared to traditional algorithms, this avoids circuit partitioning and merging, preventing increased circuit area and decreased timing performance, while also saving the time overhead required for partitioning and merging.
[0208] Furthermore, by combining the k-means clustering algorithm to group the circuit output ports, the parallel utilization of threads is optimized, further improving the efficiency of technology mapping.
[0209] Technical effects:
[0210] The above embodiments provide an FPGA parallel technology mapping method that does not require circuit partitioning, breaking through the dependence on pre-partitioning of the circuit structure in existing technologies and effectively alleviating the limitations of traditional parallel mapping methods in terms of computational efficiency, timing optimization, and resource utilization. The technical solution of the above embodiments, while maintaining the integrity of the overall circuit structure, achieves greater efficiency, flexibility, and global optimization capabilities in the FPGA circuit technology mapping process through the coordinated efforts of intelligent grouping strategies based on the circuit output port transmission-type fan-in cone circuit, topology sorting control of the parallel computing order, thread synchronization management with node locking mechanisms, and dynamic adjustment of multi-round optimization strategies. Specifically, the above embodiments bring the following technical effects:
[0211] First, it eliminates the need for circuit partitioning, avoiding the introduction of additional circuit structure overhead and improving mapping efficiency. Specifically, traditional parallel mapping methods rely on first partitioning the circuit and then mapping each sub-circuit separately. This method often requires copying parts of the circuit structure during partitioning, leading to additional hardware resource consumption. Furthermore, the partitioned sub-circuits are difficult to optimize globally, affecting the final mapping effect. The above embodiment, by directly partitioning parallel computing units based on the output port of the FPGA circuit corresponding to a transmission-type fan-in cone structure, achieves efficient parallel computing without physically partitioning the circuit structure. This avoids the additional hardware overhead caused by partitioning while ensuring global circuit optimization capabilities.
[0212] Furthermore, this method employs an intelligent grouping strategy based on k-means clustering to improve parallel computing efficiency. Specifically, traditional methods divide parallel computing units based solely on simple region partitioning or static grouping, which may lead to unbalanced computational load and affect overall computational efficiency. The above embodiment utilizes the k-means clustering algorithm to perform intelligent grouping based on the transmission-type fan-in cone circuit structure of the output port. This results in high computational correlation among nodes within the same group and low computational dependency between different groups, thereby maximizing the computational independence between threads, reducing synchronization wait, and increasing parallel computing throughput.
[0213] Furthermore, topological sorting is employed to optimize the computation order and ensure the correctness of data dependencies. Specifically, during parallel mapping, if the computation order is unreasonable, a situation may arise where a predecessor node has not been completed while a successor node has already begun computation, leading to data errors. In the above embodiment, within each parallel computing unit, a topological sorting algorithm based on depth-first search (DFS) is used to sort the transmission-type fan-in cone circuits corresponding to the circuit output ports, ensuring that all predecessor nodes of each node are computed before it, thus ensuring the correctness of the computational data and avoiding data conflicts in parallel computing.
[0214] Furthermore, a node locking mechanism is introduced to improve the stability and consistency of parallel computing. The above embodiment proposes a thread synchronization mechanism based on node locking, that is, when a thread is processing a certain node, other threads cannot modify the node at the same time, and the thread is only allowed to perform technical mapping calculations after all the predecessor nodes of the node have completed their calculations, thereby ensuring data consistency, avoiding thread conflict problems, and improving the stability of parallel computing.
[0215] Furthermore, a multi-round iterative optimization strategy comprehensively improves circuit performance. Specifically, traditional technical mapping methods often employ a one-time optimal cut selection strategy, lacking global optimization capabilities and easily leading to local optima while resulting in global suboptimal outcomes. The above embodiment implements a multi-round optimization strategy during the technical mapping process of each thread. That is, after the first round of mapping, the enumeration cut strategy is adjusted, and the optimized cut is updated and reselected. This gradually optimizes the mapping results in subsequent rounds, achieving a comprehensive improvement in circuit area, logic levels, lookup table efficiency, and timing performance.
[0216] Furthermore, maintaining the timing critical path improves the final circuit's operating speed. Specifically, traditional parallel mapping methods may fragment the critical path during circuit partitioning, leading to a decrease in global timing performance. The above embodiment avoids fragmenting the critical path by not physically partitioning the circuit, thereby increasing the maximum operating frequency of the final FPGA circuit and optimizing the circuit's timing performance.
[0217] Furthermore, the reverse topological order optimization of the mapping results ensures optimal performance of the final mapped circuit. Specifically, traditional methods may employ local optimization during mapping result integration, neglecting global mapping structure optimization, which limits the performance of the final generated circuit. In the mapping result acquisition stage, the above embodiment selects the optimal cut at each node, determined by multiple rounds of mapping, based on the reverse topological order, thereby ensuring that the final mapped circuit is as close as possible to the optimal state in multiple dimensions such as area, power consumption, and timing.
[0218] In summary, the above embodiments construct a highly efficient, stable, and circuit-partition-free FPGA parallel technology mapping method through a series of innovative technical means, including intelligent grouping of output ports, topology sorting to optimize computational order, node locking thread synchronization mechanism, multi-round optimization strategy, and timing critical path optimization. Compared with existing methods, the above embodiments can significantly improve the computational efficiency of FPGA technology mapping, reduce the additional overhead caused by circuit partitioning during the mapping process, and ensure that the final mapping result is as close as possible to the optimal solution in terms of circuit area, lookup table utilization efficiency, logic levels, and timing performance. Furthermore, the technical solutions of the above embodiments are applicable to multiple application scenarios such as high-performance computing, artificial intelligence acceleration chips, communication processors, and image processors, and have broad engineering application value and industrialization potential.
[0219] The second embodiment of this application relates to an FPGA parallel technology mapping system without circuit partitioning, the structure of which is shown in Figure 2. This FPGA parallel technology mapping system without circuit partitioning includes:
[0220] The output port grouping module is used to obtain the number of output ports N and the number of available threads T of the circuit to be mapped. Based on the structural characteristics of the transmission type fan-in cone circuit corresponding to each output port, the output ports are divided into T groups according to similarity. Each group contains at least one output port to maintain the integrity of the entire circuit to be mapped in the technical mapping process.
[0221] The topology sequence determination module is used to determine the node topology sequence of each group of output ports of the transmission-type fan-in cone circuit, so as to serve as the execution order of subsequent technology mapping.
[0222] The parallel technology mapping module is used to allocate an independent processing thread to each group of output ports. The processing thread performs enumeration cuts node by node based on the corresponding node topology sequence to obtain the optimal cut that satisfies the optimization objective.
[0223] The mapping result generation module is used to select the optimal cut on each node based on the inverse topology order of the entire circuit to be mapped, covering the entire circuit's AIG netlist, to obtain the final technical mapping circuit.
[0224] The first embodiment is a method embodiment corresponding to this embodiment. The technical details in the first embodiment can be applied to this embodiment, and the technical details in this embodiment can also be applied to the first embodiment.
[0225] It should be noted that those skilled in the art should understand that the implementation functions of each module shown in the above-described implementation of the FPGA parallel technology mapping system without circuit partitioning can be understood with reference to the relevant description of the FPGA parallel technology mapping method without circuit partitioning. The functions of each module shown in the above-described implementation of the FPGA parallel technology mapping system without circuit partitioning can be implemented by a program (executable instructions) running on a processor, or by specific logic circuits. If the above-described FPGA parallel technology mapping system without circuit partitioning is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media that can store program code, such as USB flash drives, mobile hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.
[0226] Accordingly, this application also provides a computer storage medium storing computer-executable instructions, which, when executed by a processor, implement the various method implementations of this application.
[0227] Furthermore, this application also provides an FPGA parallel technology mapping system without circuit partitioning, including a memory for storing computer-executable instructions and a processor; the processor is used to implement the steps in the above-described method embodiments when executing the computer-executable instructions in the memory. The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The aforementioned memory can be read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or solid-state drive, etc. The steps of the methods disclosed in the various embodiments of this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor.
[0228] It should be noted that in this patent application, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. In this patent application, if it refers to performing an action according to an element, it means performing the action at least according to that element, including two cases: performing the action only according to that element, and performing the action according to that element and other elements. Expressions such as "multiple," "repeatedly," and "various" include two, two times, two kinds, and more than two, more than two times, and more than two kinds.
[0229] All documents mentioned in this application are considered to be incorporated in their entirety into the disclosure of this application so that they can serve as a basis for modifications if necessary. Furthermore, it should be understood that after reading the foregoing disclosure of this application, those skilled in the art can make various alterations or modifications to this application, and these equivalent forms also fall within the scope of protection claimed in this application.
Claims
1. A parallel mapping method for FPGAs without circuit partitioning, characterized in that, Includes the following steps: Obtain the number of output ports N and the number of available threads T of the circuit to be mapped. Based on the structural characteristics of the transmission-type fan-in cone circuit corresponding to each output port, divide the output ports into T groups according to similarity. Each group contains at least one output port to maintain the integrity of the entire circuit to be mapped in the technical mapping process. Based on the transmission-type fan-in cone circuit of each output port, determine its node topology sequence as the execution order of subsequent technology mapping; Each output port is assigned an independent processing thread. The processing thread performs enumeration cuts node by node based on the corresponding node topology sequence and obtains the optimal cut that satisfies the optimization objective. Based on the inverse topological order of the entire circuit to be mapped, the optimal cut on the node is selected to cover the entire circuit's AIG netlist, thus obtaining the final technical mapping circuit.
2. The method as described in claim 1, characterized in that, The step of grouping the output ports by similarity includes: calculating the distance metric between any two output ports i and j based on the node set of the transmission-type fan-in cone circuit corresponding to each output port. Where Si and Sj represent the node sets of the transmission fan-in cone circuits corresponding to output ports i and j, respectively, and the smaller the distance metric value, the higher the similarity. The k-means clustering algorithm is used to cluster the output ports into T groups based on the distance metric.
3. The method as described in claim 1, characterized in that, When the number of output ports N is less than the number of available threads T, each output port is grouped into a separate group.
4. The method as described in claim 1, characterized in that, A depth-first search method is used to generate the topology sequence, and backtracking is used to ensure that all predecessor nodes precede their successors.
5. The method as described in claim 1, characterized in that, Each output port group is assigned an independent processing thread. The processing thread performs enumerated cuts node by node based on the corresponding node topology sequence and obtains the optimal cut that satisfies the optimization objective: initialize all nodes to an "unvisited" state; before a thread accesses a node, check its locking status and ensure that all input nodes have been visited; lock unlocked nodes, and if other threads have already locked them, wait or select other processable nodes; after the mapping is completed, unlock the node and mark it as "visited", and notify other threads waiting for that node.
6. The method as described in claim 1, characterized in that, The selection of the optimal cut is based on at least one of the following optimization objectives: • Maximum circuit delay; • Logical series; • Number of times the lookup table is used.
7. The method as described in claim 1, characterized in that, In the parallel technology mapping process, nodes on the timing critical path are identified and marked, and the fragmentation of the critical path is reduced during thread allocation to optimize the timing performance of the circuit.
8. The method as described in claim 1, characterized in that, Multiple rounds of iterative optimization are implemented in each processing thread, and the enumeration cut strategy is dynamically adjusted in each round of optimization to adapt to the current circuit optimization goal.
9. The method as described in claim 1, characterized in that, The circuit mapping result is obtained based on the inverse topology order of the entire circuit to optimize the timing performance, area utilization and power consumption of the circuit.
10. An FPGA parallel technology mapping system that does not require circuit partitioning, characterized in that, include: The output port grouping module is used to obtain the number of output ports N and the number of available threads T of the circuit to be mapped. Based on the structural characteristics of the transmission type fan-in cone circuit corresponding to each output port, the output ports are divided into T groups according to similarity. Each group contains at least one output port to maintain the integrity of the entire circuit to be mapped in the technical mapping process. The topology sequence determination module is used to determine the node topology sequence of each group of output ports of the transmission-type fan-in cone circuit, so as to serve as the execution order of subsequent technology mapping. The parallel technology mapping module is used to allocate an independent processing thread to each group of output ports. The processing thread performs enumeration cuts node by node based on the corresponding node topology sequence to obtain the optimal cut that satisfies the optimization objective. The mapping result generation module is used to select the optimal cut on the node based on the inverse topological order of the entire circuit to be mapped, and obtain the final technical mapping circuit.