A layout method and application for scalable multi-die network-on-chip FPGA architecture
Through integer linear programming and hierarchical recursive layout algorithm, a multi-grain on-chip network FPGA architecture with hierarchical topology is constructed, which solves the problem of insufficient scalability of multi-grain FPGAs in the existing technology, and realizes larger-scale circuit design and more efficient layout algorithms.
Patent Information
- Application Number
- CN202211257475.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-14
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2042-10-14
AI Technical Summary
The scalability of existing multi-grain FPGAs and their supporting EDA tools cannot meet the scale growth of circuit designs. For example, the most advanced EDA tools on the Xilinx U250 can only complete a 13×16 scale layout at 316MHz.
It provides a layout method for scalable multi-grain on-chip network FPGA architecture, adopts integer linear planning problems and hierarchical recursive layout algorithms, and builds a hierarchical topology through NoC connections and central routers, and utilizes on-chip network resources to improve the scalability of FPGA architecture.
It improves the scalability of FPGA scale, reduces the complexity of layout algorithms, and provides more parallelization opportunities, achieving larger-scale circuit design.
Smart Images

Figure CN115935887B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a layout method for a scalable multi-die network-on-chip (FPGA) architecture and its application. Background Art
[0002] Emerging applications, such as convolutional network accelerators [1] and deep learning accelerators [2], require larger multi-die FPGAs. However, the scalability of existing architectures and related EDA tools is insufficient to keep pace with the growth in the number of FPGA die. In recent years, many works have explored innovations in interconnect architectures. For example, [3] and [4] demonstrate methods for using on-chip networks to improve system performance.
[0003] These architectural innovations have put forward new requirements for EDA tools. To address these challenges, [5] proposed a high-performance custom interconnect architecture for FPGAs with HBM and a novel high-level synthesis-based optimization technology to improve the performance of AXI on-chip network components. However, these methods only consider traditional substrate-based mesh topology FPGAs and cannot map the design to more complex die topologies. [6] After observing the traditional interconnect architecture on modern substrate-based multi-die FPGA architectures, the sub-modules in the design are distributed across multiple dies to improve the overall performance of the system. However, this method only focuses on traditional interconnect resources and ignores dedicated interconnect resources represented by on-chip networks. These existing systems can only handle traditional substrate-based architectures and their adaptation objects do not include scalable multi-die FPGA architectures.
[0004] References:
[0005] 【1】W.Jiang, H.Yu, X.Liu, and Y.Ha, “Energy efficiency optimizationoffpga-based CNN accelerators with full data reuse and VFS,” in 26 th IEEEInternational Conference on Electronics, Circuits and Systems, ICECS 2019, Genoa, Italy, November 27-29, 2019. IEEE, 2019, pp.446–449.
[0006] 【2】W.Jiang,H.Yu,X.Liu,H.Sun,R.Li,and Y.Ha,“Tait:One-shotfullintegerlightweight dnn quantization via tunable activationimbalancetransfer,”in 2021 58th ACM / IEEE Design Automation Conference(DAC).IEEE,2021,pp.1027–1032.
[0007] 【3】K.Khalil,O.Eldash,B.Dey,A.Kumar,and M.Bayoumi,“Anefficientembryonic hardware architecture based on network-on-chip,”in2021 IEEEInternational Midwest Symposium on Circuits and Systems(MWSCAS),2021,pp.449-452.
[0008] 【4】G.Passas,M.Katevenis,and D.Pnevmatikatos,“Crossbar nocsarescalable beyond 100nodes,”Trans.Comp.-Aided Des.Integ.Cir.Sys.,vol.31,no.4,p.573-585,Apr.2012.
[0009] 【5】Y.-k.Choi,Y.Chi,W.Qiao,N.Samardzic,and J.Cong,“Hbmconnect:High-performance hls interconnect for fpgahbm,”in The2021ACM / SIGDA InternationalSymposium on Field-ProgrammableGate Arrays,ser.FPGA’21.New York,NY,USA:Association forComputing Machinery,2021,p.116-126.
[0010] 【6】L.Guo, Y.Chi, J.Wang, J.Lau, W.Qiao, E.Ustun, Z.Zhang, and J.Cong, "Autobridge: Coupling coarse-grained floorplanning and pipelining for high-frequency hls design on multi-die fpgas," in The2021ACM / SIGDA InternationalSymposium on Field-ProgrammableGate Arrays, ser. FPGA'21. New Yoork, NY,, USA: Association for Computing Machinery, 2021, p.81-92. Summary of the Invention
[0011] The technical problem to be solved by the present invention is that the scalability of existing multi-die FPGAs and their supporting EDA tools cannot meet the scale growth of circuit design. For example, the most advanced EDA tools on the most advanced commercial FPGA Xilinx U250 can only complete a 13×16 scale layout for convolutional neural networks at 316MHz.
[0012] In order to solve the above technical problems, a technical solution of the present invention is to provide a layout method for a scalable multi-die on-chip FPGA architecture, characterized in that when the structural parameter is (), the FPGA architecture is a single die; when the structural parameter is (m), m is a positive integer, the FPGA architecture is m single die connected to a NoC router via NoC, and the NoC router is called the central router of the FPGA architecture; when the structural parameter is (m1, m2), m1 and m2 are positive integers, the FPGA architecture is m2 central routers of (m1) structure connected to a NoC router via NoC, and the NoC router is called the central router of the FPGA architecture; when the structural parameter is (m1, m2), m1 and m2 are positive integers, the FPGA architecture is m2 central routers of (m1) structure connected to a NoC router via NoC, and the NoC router is called the central router of the FPGA architecture; n ), m1, ..., m n is a positive integer, and the FPGA architecture is m n (m1, ..., m n-1 ) structure is connected to a NoC router via NoC, and the NoC router is called the central router of the FPGA architecture, and (m1, ..., m n-1 ) structure is called the second-level substructure;
[0013] Laying out the FPGA fabric involves an integer linear programming problem and a hierarchical recursive placement algorithm based on the integer linear programming problem, where:
[0014] The integer linear programming problem consists of the following steps:
[0015] Step 1: record the FPGA architecture model as GFPGA. For architecture topology, is the NoC link bandwidth of each layer, is the resource capacity of each die. The data flow design is recorded as graph G design , G design =(V, E, a(V), S(E), D(E), w(E)), where V is the data flow module, E is the data flow queue, a(V) is the area of the data flow module, S(E) is the starting point of the data flow queue, D(E) is the end point of the data flow queue, and w(E) is the width of the data flow queue;
[0016] Step 2: Remember For target layout, Representing the grain, the objective function dominated by the vertex is shown as follows:
[0017]
[0018] Where w(e) represents the data stream queue width, d m () represents the distance metric, S(e) represents the queue source module, Indicates the queue source module corresponding to the die, D(e) indicates the queue drain module, Indicates the die corresponding to the queue drain module.
[0019] Step 3: Use the one-hot code to encode the linearized vertex space, and the corresponding The linearized linear transformation Φ of , so the linearized objective function is shown as follows, which is the objective function of the integer linear programming problem:
[0020]
[0021] Where w T e represents the linear form of the data stream queue width, Se represents the linear form of the queue source module, ΦSe represents the linear form of the queue source module corresponding to the die, De represents the linear form of the queue drain module, and ΦDe represents the linear form of the queue drain module corresponding to the die;
[0022] Step 4: Each data flow module is placed on exactly one die, formalized as the constraint shown below:
[0023]
[0024] In the formula, x represents the target die, v represents the data flow module to be laid out, Φ xv represents a placement decision variable that is 1 if the dataflow module x is assigned to the die v and 0 otherwise;
[0025] Step 5: The total resources of the data flow modules on the same die must not exceed the total resources of the die; this can be formalized as the following constraint:
[0026]
[0027] Where a(v) represents the resource usage of data flow module v, and a(x) represents the resource capacity of die x;
[0028] Step 6: The user can provide a manual layout, formalized as constraints as shown below:
[0029]
[0030] Where, Indicates the user's manual allocation of the data flow module to the corresponding grain, V M It represents the data flow module that represents the manual allocation of grains by the design user;
[0031] The hierarchical recursive layout algorithm includes the following steps:
[0032] Step a: Place the data flow module in the FPGA topology The layout results on the substructure are summarized as As shown in the following two formulas:
[0033]
[0034]
[0035] Where, Represents the top-level substructure, represents the nth level substructure with structure parameter m and position x, Represents the tuple m without the tail item, Represents the tuple x without the first item;
[0036] Step b: Define the recursive layout operator φ:
[0037]
[0038] Where, represents the recursive layout of module v starting from the nth level substructure, Represents the secondary layout of module v on the nth level substructure, represents the grain at position y;
[0039] have Thus the original layout problem The solution is decomposed into substructure layout The solution of
[0040] Step c: The objective function on the substructure is transformed into an edge-dominated representation, as shown below:
[0041]
[0042] Where, represents the layout to be solved on the n-th substructure with structural parameter m and position x, represents the data flow queue assigned to the n-th substructure with structural parameter m and position x, d represents the distance metric of the on-chip network link, and Ξe represents the on-chip network link corresponding to the data flow queue on the n-th substructure.
[0043] Step d: When performing layout, establish constraints based on the following conditions:
[0044] The computation flow module is assigned to exactly one sub-level sub-structure in the sub-structure;
[0045] The computation flow queue is allocated to one of the links between the current substructure central router and the next-level substructure central router in the substructure.
[0046] The allocation of computing flow modules is consistent with the allocation of computing flow queues;
[0047] For resource estimation of the i-th level substructure, the congestion factor ρ is introduced i As a pair v∈V a(v)Φ xv ≤a(x), The correction is as follows:
[0048]
[0049] Where A represents the resource type.
[0050] The bit width of the computational flow module allocated to the link must not exceed the link bandwidth;
[0051] The layout on the substructure is consistent with the user's manual layout.
[0052] For the 0th level substructure, the structural parameters are (), and the FPGA architecture is expressed as follows:
[0053]
[0054] Where, It represents the nth level substructure with structure parameter m and position X. Indicates the 0th level substructure with structure parameter () and position X, Represents the grain at position X, m j represents the jth item of the total structural parameter, x i Represents the i-th item of tuple X.
[0055] For the nth level substructure, when the structure parameters are (m1, ..., m n ), the FPGA architecture is expressed as follows:
[0056]
[0057] Where, The structural parameters are The n-1th level substructure at position (x, x), where x represents the relative position of the n-1th level substructure in the current nth level substructure.
[0058] Another technical solution of the present invention is to provide an application of the aforementioned layout method for a scalable multi-die network-on-chip FPGA architecture, which is characterized in that it is used in the design of multi-die FPGAs to improve the scalability of the FPGA architecture and facilitate the scalable implementation of supporting EDA tools.
[0059] This paper discloses a scalable multi-die FPGA architecture based on a network-on-chip (NOC) and a corresponding hierarchical recursive placement algorithm. The goal is to directly map register-transfer-level dataflow designs generated by existing high-level synthesis onto the proposed interconnect architecture. The disclosed method can unlock the potential of hierarchical topologies and more effectively utilize dedicated interconnect resources, such as cross-die wire meshes, NOCs, and high-speed transceivers. Compared to existing solutions, this paper offers the following innovations:
[0060] 1) A multi-die network-on-chip FPGA with a hierarchical topology that improves FPGA scalability relative to the number of die and is more user-friendly for efficient implementation of placement algorithms.
[0061] 2) The integer linear programming problem formulation for the proposed interconnect architecture. This formulation contrasts with the traditional Cartesian grid layout problem, where the distance metric is simply the 11-norm of vertex coordinate differences. In contrast, the novel distance metric for the proposed hierarchical network-on-chip interconnect architecture is defined on the edges of the payload data flow and involves a complex combination of integer linear programming primitives, exemplified by cascaded conditional branches. Consistency constraints on the resulting vertex and edge layouts in the dataflow graph are also introduced.
[0062] 3) A novel recursive method is used to solve the aforementioned integer linear programming problem. Leveraging the hierarchical nature of the proposed architecture, the method disclosed in this paper divides the original problem into independent subproblems on the sub-architecture and solves them separately. This not only reduces the overall complexity of the problem but also introduces numerous opportunities for parallelization. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 The structure of the architecture proposed in the present invention is shown when the structural parameters are (8, 8).
[0064] Figure 2 The flowchart of the hierarchical layout algorithm of the present invention is shown;
[0065] Figure 3 The specific implementation of the algorithm provided by the present invention is demonstrated. DETAILED DESCRIPTION
[0066] Below in conjunction with specific embodiment, further set forth the present invention.Should be understood that these embodiments are only used to illustrate the present invention and are not used in limiting the scope of the present invention.In addition, should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall equally within the scope limited by the appended claims of the application.
[0067] For multi-die FPGA based on on-chip network, we recursively define the m-tree topology on it, where is the total structural parameter. For the basic case of l = 1, the topology is m l =m1 grains are connected to a central router, which is called the first-level router; for the case of l>1, the topology is m l indivual The level l-1 routers of the tree are connected to a central router, which is called the level l central router.
[0068] The architecture of the scalable multi-die network-on-chip FPGA proposed in this invention is as follows:
[0069] When the structure parameter is (), the FPGA fabric substructure is a single die, as shown in the following formula:
[0070]
[0071] in, Indicates the nth level substructure with structure parameter m and position X, Indicates the 0th level substructure with structure parameter () and position X, Represents the grain at position X, m j represents the jth item of the total structural parameters, x i Represents the i-th item of tuple X.
[0072] When the structure parameter is (m) (where m is a positive integer), the FPGA fabric refers to m single dies connected to a NoC router (called the central router of the fabric) via the NoC.
[0073] When the structure parameters are (m1, m2) (where m1 and m2 are positive integers), the FPGA architecture referred to is m2 central routers of (m1) structures connected to a NoC router via NoC, and the NoC router is called the central router of the FPGA architecture.
[0074] When the structural parameters are (m1, ..., m n ) (where m1, ..., m n is a positive integer), the FPGA structure is m n (m1, ..., m n-1 The central router of the FPGA fabric (called the second-level substructure) is connected to a NoC router via the NoC. The NoC router is called the central router of the FPGA fabric, as shown in the following formula:
[0075]
[0076] Where, The structural parameters are The n-1th level substructure at position (x, x), x represents the relative position of the n-1th level substructure in the current nth level substructure
[0077] The FPGA architecture when the structural parameters are (8, 8) is as follows Figure 1 shown.
[0078] The distance metric on the proposed architecture is as follows:
[0079]
[0080] Where, represents the grain at position x1, represents the grain at position x2, x1 represents the first grain of the pair of grains whose distance is to be determined, x2 represents the second grain of the pair of grains whose distance is to be determined, (x1) i represents the i-th item of the tuple x1, (x2) i represents the i-th item of tuple x2, Indicates that when (x1) i ≠(x2) i An indicator variable that is 1 when s is true and 0 otherwise.
[0081] The resource calculation on the proposed architecture substructure is shown in the following formula:
[0082]
[0083] Where, Indicates the resource capacity of the nth level substructure with structure parameter m and position x, T n-1 Indicates the n-1th level substructure in the current nth level substructure, represents the die at position X, and a(X) represents the resource capacity of the die X.
[0084] When NoC supports time division multiplexing, the time division multiplexing factor is k TDM , the bandwidth of the NoC link at level i is B i , NoC nominal operating frequency f NoC , the design nominal operating frequency is f op , then the equivalent bandwidth of the NoC link is The designed equivalent operating frequency is
[0085] The integer linear programming problem of the proposed layout problem is expressed as follows:
[0086] Step 1: The NoC FPGA hierarchical architecture model proposed above is recorded as graph GFPGA. For architecture topology, is the NoC link bandwidth of each layer, is the resource capacity of each die. The data flow design is recorded as graph G design , G design =(V, E, a(V), S(E), D(E), w(E)), V is the data flow module, E is the data flow queue, a(V) is the area of the data flow module, S(E) is the starting point of the data flow queue, D(E) is the end point of the data flow queue, and w(E) is the bit width of the data flow queue.
[0087] Step 2: Remember For target layout, Representing the set of all grains, the objective function dominated by the vertex is shown as follows:
[0088]
[0089] Where w(e) represents the data stream queue width, d m () represents the distance metric, S(e) represents the source module of the data flow queue, represents the die corresponding to the source module of the data flow queue, D(e) represents the sink module of the data flow queue, Indicates the die corresponding to the drain module of the data flow queue.
[0090] Step 3: Use the one-hot code to encode the linearized vertex space, and the corresponding The linearized linear transformation Φ of , so the linearized objective function is shown as follows, which is the objective function of the integer linear programming problem:
[0091]
[0092] Where w T e represents the linear form of the bit width of the data flow queue, Se represents the linear form of the source module of the data flow queue, ΦSe represents the linear form of the source module of the data flow queue corresponding to the grain, De represents the linear form of the drain module of the data flow queue, and ΦDe represents the linear form of the drain module of the data flow queue corresponding to the grain.
[0093] Step 4: Each data flow module should be placed on exactly one die, which can be formalized as the following constraint:
[0094]
[0095] In the formula, x represents the target die, v represents the data flow module to be laid out, Φ xv A placement decision variable that is 1 if dataflow module x is assigned to die v and 0 otherwise.
[0096] Step 5: The total resources of the data flow modules on the same die must not exceed the total resources of the die, which can be formalized as the following constraint:
[0097]
[0098] Where a(v) represents the resource usage of data flow module v, and a(x) represents the resource capacity of die x.
[0099] Step 6: The user can provide a manual layout, which is formalized as the following constraints:
[0100]
[0101] Where, Indicates the user's manual allocation of the data flow module to the corresponding grain, V M Represents the data flow module where the design user manually assigns die.
[0102] The proposed hierarchical recursive layout algorithm is formulated as follows:
[0103] Step 1: Place the data flow module in the FPGA topology The layout results on the substructure are summarized as
[0104] As shown in the following two formulas:
[0105]
[0106]
[0107] Where, Represents the top-level substructure, represents the nth level substructure with structure parameter m and position x, Represents the tuple m without the tail item, Represents the tuple x without the first item;
[0108] Step 2
[0109] Define the recursive layout operator φ:
[0110]
[0111] Where, represents the recursive layout of module v starting from the nth level substructure, Represents the secondary layout of module v on the nth level substructure, Represents the grain at position y.
[0112] have Thus the original layout problem The solution is decomposed into substructure layout The solution.
[0113] Step 3: The objective function on the substructure is transformed into an edge-dominated representation, as shown below:
[0114]
[0115] Where, represents the layout to be solved on the n-th substructure with structural parameter m and position x, represents the data flow queue assigned to the n-th substructure with structural parameter m and position x, d represents the distance metric of the on-chip network link, and Ξe represents the on-chip network link corresponding to the data flow queue on the n-th substructure.
[0116] Step 4: The computational flow module should be assigned to exactly one substructure in the substructure, which can be formalized as the following constraint:
[0117]
[0118] Where, An indicator variable indicating whether the data flow module v is allocated to the x-th secondary substructure in the n-th substructure with structure parameter m and position x.
[0119] Step 5: The calculated flow queue should be allocated to exactly one of the links between the current substructure central router and the next-level substructure central router on the substructure, which can be formalized as the constraint shown in the following formula:
[0120]
[0121] Where, ηe The placement decision variable indicating whether the data stream queue e is assigned to the link η, η represents the on-chip network link between the secondary sub-nodes in the current n-th sub-structure, E T Represents the total number of on-chip network links between secondary sub-nodes within the current n-th level sub-structure.
[0122] Step 6: The allocation of computational flow modules should be consistent with the allocation of computational flow queues, which can be formalized as the following constraint:
[0123]
[0124]
[0125] Where, Represents the source module mapping of the data queue that is laid out on the nth level substructure with structure parameter m and position x, Represents the drain module mapping of the data queue that is laid out on the n-th level substructure with structure parameter m and position x, S T Ξ represents the source substructure of the on-chip network link between the secondary subnodes in the current n-th substructure, D T Ξ represents the leaky substructure of the on-chip network link between the secondary sub-nodes in the current n-th substructure.
[0126] Step 7: Estimate the resources of the i-th level substructure and introduce the congestion factor ρ i As a pair v∈V a(v)Φ xv ≤a(x), The correction is as follows:
[0127]
[0128] Where A represents the resource type.
[0129] Step 8: The bit width of the computational flow module allocated to the link must not exceed the link bandwidth, which can be formalized as the following constraint:
[0130]
[0131] Where, δ η Indicates whether the source and drain of link n are the same substructure, w(e) indicates the data stream queue width, ηeThe placement decision variable indicating whether data flow queue e is placed on link n.
[0132] Step 9: The layout on the substructure should be consistent with the user's manual layout, which can be formalized as the following constraint:
[0133]
[0134] Where, Indicates that the tuple m removes the tail item, Indicates the relative position of the user's manual layout of the grain in the n-level substructure, Indicates the secondary substructure corresponding to the n-level substructure of the user's manual layout grain, V M Represents a data flow module that involves manual layout by the user.
[0135] The specific implementation of the algorithm proposed in this invention is as follows Figure 3 As shown, the following steps are included:
[0136] For the proposed layout problem, from k TDM = 0, and try the layout in a loop, such as Figure 3 As shown in lines 2 to 3.
[0137] First k TDM Self-increment, such as Figure 3 As shown in line 4, if k TDM If the upper limit given by the user is exceeded, the return will be unsolvable, such as Figure 3 As shown in lines 5 to 7.
[0138] Try recursively layer by layer and substructure by substructure, such as Figure 3 As shown in lines 8 to 9. Try to layout the content as a substructure, such as Figure 3 As shown in line 10. If there is no solution for any level or any substructure, give up the current round and try to enter the next cycle, as shown in Figure 3 Lines 11 to 13 show this. For successful attempts, statistics are and In the value of the current substructure, such as Figure 3 As shown in line 14, Indicates the upper level substructure, Indicates the current substructure, Represents the layout result on the previous level substructure.
[0139] If a feasible solution is found before kTDM exceeds the upper limit, Calculate the overall layout results, such as Figure 3 Row 18 shows the layout result under the time division multiplexing factor kTDM. like Figure 3 Shown in line 19.
[0140] The scalable architecture proposed in this invention can be applied to the design of new multi-die FPGAs, improving the scalability of the FPGA architecture and facilitating the scalable implementation of supporting EDA tools. The hierarchical layout algorithm proposed in this invention can be scalably applied to the EDA tools required by new multi-die FPGAs, significantly increasing the achievable design scale while reducing algorithm runtime without compromising design performance.
Claims
1. A layout method for a scalable multi-die network-on-chip FPGA architecture, characterized in that: When the structural parameter is (), the FPGA architecture is a single crystal; when the structural parameter is (m), m is a positive integer, the FPGA architecture is m single crystals connected to a NoC router via NoC, and the NoC router is called the central router of the FPGA architecture; when the structural parameter is (m1, m2), m1 and m2 are positive integers, the FPGA architecture is m2 (m1) structure central routers connected to a NoC router via NoC, and the NoC router is called the central router of the FPGA architecture; when the structural parameter is (m1, ..., m n ), m1,……,m n is a positive integer, and the FPGA architecture is m n The central routers are connected to a NoC router via NoC, and the NoC router is called the central router of the FPGA architecture, and (m1, ..., m n-1 ) structure is called the second-level substructure; Laying out the FPGA fabric involves an integer linear programming problem and a hierarchical recursive placement algorithm based on the integer linear programming problem, where: The integer linear programming problem consists of the following steps: Step 1: Record the FPGA architecture model as graph G FPGA , For architecture topology, is the NoC link bandwidth of each layer, is the resource capacity of each die; the data flow design is recorded as graph G design , V is the set of data flow modules, E is the set of data flow queues, a(v) is the resource occupancy of data flow module v, S(e) is the starting point of the data flow queue, D(e) is the end point of the data flow queue, and w(e) is the bit width of the data flow queue; Step 2: Remember For target layout, Representing the set of all grains, the objective function dominated by the vertex is shown as follows: Where w(e) represents the data stream queue width, d m () represents the distance metric, S(e) represents the source module of the data flow queue, represents the die corresponding to the source module of the data flow queue, D(e) represents the sink module of the data flow queue, Indicates the die corresponding to the drain module of the data flow queue; Step 3: Use the one-hot code to encode the linearized vertex space, and the corresponding The linearized linear transformation Φ of , so the linearized objective function is shown as follows, which is the objective function of the integer linear programming problem: Where w T e represents the linear form of the data stream queue width, Se represents the linear form of the queue source module, ΦSe represents the linear form of the queue source module corresponding to the die, De represents the linear form of the queue drain module, and ΦDe represents the linear form of the queue drain module corresponding to the die; Step 4: Each data flow module is placed on exactly one die, formalized as the constraint shown below: In the formula, x represents the target die, v represents the data flow module to be laid out, Φ xv represents a placement decision variable that is 1 if the dataflow module x is assigned to the die v and 0 otherwise; Step 5: The total resources of the data flow modules on the same die must not exceed the total resources of the die; this can be formalized as the following constraint: Where a(v) represents the resource usage of data flow module v, and a(x) represents the resource capacity of die x; Step 6: The user can provide a manual layout, formalized as constraints as shown below: Where, Indicates the user's manual allocation of the data flow module to the corresponding grain, V M Represents the data flow module where the design user manually allocates the die; The hierarchical recursive layout algorithm includes the following steps: Step a: Place the data flow module in the FPGA topology The layout results on the substructure are summarized as As shown in the following two formulas: Where, Represents the top-level substructure, represents the nth level substructure with structure parameter m and position x, Represents the tuple m without the tail item, Represents the tuple x without the first item; Step b: Define the recursive layout operator φ: Where, represents the recursive layout of module v starting from the nth level substructure, Represents the secondary layout of module v on the nth level substructure, represents the grain at position y; Thus the original layout problem The solution is decomposed into substructure layout The solution of Step c: The objective function on the substructure is transformed into an edge-dominated representation, as shown below: Where, represents the layout to be solved on the n-th substructure with structural parameter m and position x, represents the data flow queue assigned to the n-th substructure with structure parameter m and position x, d represents the distance metric of the on-chip network link, and Ξe represents the on-chip network link corresponding to the data flow queue on the n-th substructure; Step d: When performing layout, establish constraints based on the following conditions: The computation flow module is assigned to exactly one sub-level sub-structure in the sub-structure; The computation flow queue is allocated to one of the links between the current substructure central router and the next-level substructure central router in the substructure. The allocation of computing flow modules is consistent with the allocation of computing flow queues; For resource estimation of the i-th level substructure, the congestion factor ρ is introduced i As a pair The correction is as follows: Where A represents the resource type; The bit width of the computational flow module allocated to the link must not exceed the link bandwidth; The layout on the substructure is consistent with the user's manual layout.
2. A layout method for a scalable multi-die network-on-chip FPGA architecture according to claim 1, characterized in that: When the structure parameter is (), the FPGA architecture is expressed as follows: Where, It represents the nth level substructure with structure parameter m and position X. Indicates the 0th level substructure with structure parameter () and position X, Represents the grain at position X, m j represents the jth item of the total structural parameters, x i Represents the i-th item of tuple X.
3. The layout method for a scalable multi-die network-on-chip FPGA architecture according to claim 1, wherein: When the structural parameters are (m1, ..., m n ), the FPGA architecture is expressed as follows: Where, The structural parameters are The n-1th level substructure at position (x, X), x represents the relative position of the n-1th level substructure in the current nth level substructure.
4. An application of the layout method for a scalable multi-die network-on-chip FPGA architecture as claimed in claim 1, characterized in that: It is used in the design of multi-die FPGAs to improve the scalability of the FPGA architecture and facilitate the scalable implementation of supporting EDA tools.
Citation Information
Patent Citations
Built-in self-test structure and method for on-chip network resource node storage device
CN103310850A
FPGA-based mapping-oriented network-on-chip verification method and system
CN109450705A