Core particle interconnection network generation method and generation system
By optimizing the core-grain layout and network generation methods, communication bottlenecks and design cost problems in core-grain technology are solved, and efficient chip system communication and resource allocation are achieved.
Patent Information
- Application Number
- CN202510372967.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-08-01
AI Technical Summary
The existing core-grain technology has bottlenecks in the network bandwidth, delay and power consumption performance between chips in chip systems, and lacks a cross-layer collaborative optimization mechanism, resulting in low interconnection resource allocation efficiency, making it difficult to adapt to heterogeneous core-grain integration scenarios, increasing design costs and cycles.
The core particle layout, inter-chip network and on-chip network are generated using a specific execution sequence. The inter-chip and on-chip communication paths are optimized through the minimum cutting algorithm and the minimum cost maximum flow algorithm, rationalize resource allocation, and avoid redundant communication links and congestion.
It improves the communication efficiency and performance of the chip system, reduces communication overhead, avoids redundancy in the number of routers between and within the core and core particles, and reduces design costs and development cycles.
Smart Images

Figure CN120409404A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of chip manufacturing, specifically to the Chiplet technology in the field of chip manufacturing, and more specifically, to a method and system for generating a Chiplet interconnection network. Background Art
[0002] Chiplet technology, as a new chip design and manufacturing method, has received extensive attention. This technology splits a complex chip system into multiple smaller Chiplets with specific functions, designs and manufactures them separately, and then integrates these Chiplets together through advanced packaging technology to form a complete chip system.
[0003] Although Chiplet technology can meet the functional requirements of complex chip systems, the additional design complexity introduced has become a technical bottleneck. In the process of constructing a complete chip system, it is necessary to design the network between Chiplets (hereinafter simply referred to as the inter-chiplet network) and the network within Chiplets (hereinafter simply referred to as the on-chiplet network) simultaneously. However, the current wiring resource limitations of the Chiplet integration process have led to the bandwidth, latency, and power consumption performance of the inter-chiplet network becoming the communication bottleneck of the chip system, seriously restricting the overall communication efficiency and performance of the chip system.
[0004] In addition, the existing technology lacks a cross-layer collaborative optimization mechanism for the split processing of the inter-chiplet network and the on-chiplet network, resulting in a significant reduction in the allocation efficiency of interconnection resources. At the same time, although the rule topologies of traditional on-chip networks (such as Mesh, Torus) are extended to adapt to the interconnection requirements of chip systems, such rule topologies are difficult to match heterogeneous Chiplet integration scenarios (such as the different communication requirements of Chiplets with different sizes and functions), leading to redundant communication paths and a decrease in energy efficiency ratio.
[0005] On the other hand, although existing Chiplet technologies are committed to integrating multi-stage design processes to construct a complete chip system, however, the existing processes do not fully consider physical-level constraints during the design stage, resulting in the designed chip system facing physical implementation challenges during actual manufacturing, significantly increasing the design cost and development cycle.
[0006] In summary, the existing technology has obvious defects in aspects such as resource allocation for interconnection network design, topology adaptability in heterogeneous scenarios, and the coordination between physical constraints and design processes.
[0007] It should be noted that: This background art is only used to introduce relevant information of the present invention to facilitate understanding of the technical solution of the present invention, but it does not necessarily mean that the relevant information is prior art. Without evidence indicating that the relevant information was publicly available before the filing date of the present invention, the relevant information should not be regarded as prior art. Summary of the Invention
[0008] Therefore, an object of the present invention is to overcome the defects of the above-mentioned prior art and provide a method for generating a die interconnect network and a system for generating a die interconnect network.
[0009] The object of the present invention is achieved by the following technical solutions.
[0010] According to a first aspect of the present invention, there is provided a method for generating a die interconnect network for generating a die interconnect network required to form a complete chip based on a plurality of dies. The method includes: Step S1, determining a die layout by using a preset first strategy based on a plurality of dies required to construct a chip and layout constraint conditions; wherein, one or more functional modules are integrated in a die, and the communication requirements between the functional modules are known; the die layout represents the positional arrangement relationship between the dies; Step S2, determining the number of die interfaces and the number of inter-die routers on each die based on the communication requirements between the functional modules in each die and the functional modules in other dies, and determining the communication paths between each die and other dies by using a preset second strategy based on the number of die interfaces, the number of inter-die routers, and the die layout of each die to obtain an inter-chip network; wherein, the inter-die routers are responsible for the communication between the dies; the communication paths between each die and other dies include the communication links between all the die interfaces on this die and the corresponding die interfaces on other dies; Step S3, determining the number of intra-die routers based on the communication requirements between the functional modules in each die and the inter-chip network, and determining the internal communication paths of each die by using a preset second strategy based on the inter-chip network and the number of intra-die routers to obtain an intra-chip network; wherein, the internal communication paths of each die include the communication links between the die interfaces within the die, the communication links between the functional modules, and the communication links between the functional modules and the die interfaces within the die; wherein, the die layout, the inter-chip network, and the intra-chip network constitute a die interconnect network.
[0011] In some embodiments of the present invention, the preset second strategy is a minimum cut algorithm and a minimum cost maximum flow algorithm.
[0012] In some embodiments of the present invention, the die interfaces adopt serial interfaces or parallel interfaces.
[0013] In some embodiments of the present invention, step S2 includes: step S21, determining the total input / output bandwidth of each die based on the communication requirements between the functional modules within each die and the functional modules within other dies, dividing the total input / output bandwidth of each die by the preset single-die interface bandwidth to determine the number of die interfaces on each die, and evenly distributing the die interfaces on the boundary of the die according to the number of die interfaces on each die; wherein, the horizontal and vertical coordinates are used to represent the positions of each die interface on its own die; step S22, determining one or more target die interfaces corresponding to each die interface on each die and on other dies based on the number of die interfaces on each die and the communication requirements between the functional modules within each die and the functional modules within other dies; wherein, a die interface and a corresponding target die interface on other dies form an interface pair; step S23, starting from zero, setting the number of inter-die routers until the number of inter-die routers can support the communication between all functional modules within each die and the functional modules within other dies with the total number of die interfaces on all dies as a constraint; step S24, dividing all die interfaces into multiple interface groups by using the minimum cut algorithm based on the positions of each die interface on all dies; wherein, each interface group includes multiple die interfaces, and all die interfaces in each interface group are connected to the same inter-die router; step S25, obtaining the communication rate of all die interfaces in each interface group, respectively calculating the weighted sum of the product of the communication rate of all die interfaces in each interface group and its abscissa and the weighted sum of the product of the communication rate of all die interfaces in each interface group and its ordinate, and respectively dividing the weighted sum of the product of the abscissa and the weighted sum of the product of the ordinate by the sum of the communication rates of all die interfaces in the interface group to obtain the horizontal and vertical coordinates of the inter-die router connected by each interface group, and determining the position of the inter-die router in the die layout based on the horizontal and vertical coordinates; step S26, determining the positions of the dies where the two die interfaces in each interface pair are located based on the die layout, and selecting a matching interface type based on the positions of the dies where the two die interfaces in each interface pair are located to obtain multiple different interface insertion schemes; wherein, each interface insertion scheme represents the interface types adopted by all die interfaces, and the two die interfaces in each interface pair adopt the same interface type; step S27, selecting the interface insertion scheme with the lowest power consumption and latency from multiple different interface insertion schemes, constructing an inter-die router cost graph based on the selected interface insertion scheme, and calculating the communication link between each interface pair by using the minimum cost maximum flow algorithm based on the preset inter-die path constraint conditions; wherein, the inter-die router cost graph includes multiple nodes and multiple edges connecting the two nodes, the nodes represent die interfaces or inter-die routers, the edges represent the existence of a communication link between the two nodes, and each edge uses the communication cost between the two connected nodes as the weight.
[0014] In some embodiments of the present invention, step S3 includes: step S31, determining, based on the inter-die network, whether there is a communication requirement between any two die interfaces within each die, between a die interface and a functional module; step S32, based on the communication requirements between die interfaces within each die, between a die interface and a functional module, and between functional modules within the die, setting the number of routers within the die from scratch until the number of routers within the die can meet the communication requirements between die interfaces within the die, between a die interface and a functional module, and between functional modules within the die; step S33, dividing all die interfaces and all functional modules within each die into multiple communication groups by using the minimum cut algorithm based on the positions of functional modules and die interfaces within each die; wherein, each communication group includes multiple die interfaces and / or multiple functional modules, and all die interfaces and all functional modules in each communication group are connected to the same router within the die; step S34, obtaining the communication rates of all die interfaces and all functional modules connected to the corresponding router within each communication group, respectively calculating the weighted sum of the product of the communication rates of all die interfaces in each communication group and their abscissas and the product of the communication rates of all functional modules and their abscissas, and the weighted sum of the product of the communication rates of all die interfaces in each communication group and their ordinates and the product of the communication rates of all functional modules and their ordinates, and respectively dividing the weighted sum of the abscissas and the weighted sum of the ordinates by the sum of the communication rates of all die interfaces and all functional modules in the communication group to obtain the abscissa and ordinate of the router within the die connected by the communication group, and determining the position of the router within the die on the die based on the abscissa and ordinate; wherein, the abscissa and ordinate of the functional module represent the position of the functional module on the die where it is located; step S35, constructing a router overhead graph within the die, and calculating the communication links between die interfaces within each die, between a die interface and a functional module, and between functional modules within the die by using the minimum cost maximum flow algorithm based on the preset path constraints within the die; wherein, the router overhead graph within the die includes multiple nodes and multiple edges connecting two nodes, the nodes represent die interfaces, routers within the die or functional modules, the edges represent that there is a communication link between two nodes, and each edge uses the communication overhead between the two connected nodes as the weight.
[0015] In some embodiments of the present invention, the preset first strategy is a simulated annealing algorithm, a genetic algorithm or a particle swarm algorithm.
[0016] In some embodiments of the present invention, the method may iteratively execute steps S1, S2 and S3 multiple times to obtain the die layout, the inter-die network and the on-die network; wherein, in step S2, the inter-die network is generated based on the die layout obtained by executing step S1 for the last time; in step S3, the on-die network is generated based on the inter-die network obtained by executing step S2 for the last time.
[0017] According to a second aspect of the present invention, there is provided a die interconnect network generation system based on the method described in the first aspect of the present invention. The system includes: a layout generation module, configured to determine a die layout based on a plurality of dies required for constructing a chip and layout constraint conditions by using a preset first strategy; wherein, one or more functional modules are integrated in each die, and the communication requirements between the functional modules are known; the die layout represents the positional arrangement relationship between the dies; an inter-die network generation module, configured to determine the number of die interfaces on each die and the number of inter-die routers based on the communication requirements between the functional modules in each die and the functional modules in other dies, and determine the communication paths between each die and other dies by using a preset second strategy based on the number of die interfaces of each die, the number of inter-die routers, and the die layout to obtain an inter-die network; wherein, the inter-die routers are responsible for the communication between the dies; the communication path between each die and other dies includes the communication links between all the die interfaces on this die and the corresponding die interfaces on other dies; an on-chip network generation module, configured to determine the number of routers within each die based on the communication requirements between the functional modules within each die and the inter-die network, and determine the internal communication path of each die by using a preset second strategy based on the inter-die network and the number of routers within each die to obtain an on-chip network; wherein, the internal communication path of each die includes the communication links between the die interfaces within the die, the communication links between the functional modules, and the communication links between the functional modules and the die interfaces within the die; wherein, the die layout, the inter-die network, and the on-chip network constitute a die interconnect network.
[0018] Compared with the prior art, the advantages of the present invention are as follows: (1) The die layout, the inter-die network, and the on-chip network are generated in a specific execution order, avoiding the congestion of the on-chip network caused by not considering the communication between the dies; (2) When generating the inter-die network and the on-chip network, the corresponding inter-die network and on-chip network are generated guided by the communication requirements, so as to avoid the redundancy of the number of inter-die routers and the routers within the die, and further avoid the redundant communication links on the inter-die network and the on-chip network, reducing the communication overhead of the inter-die network and the on-chip network. Description of the Drawings
[0019] The following further describes embodiments of the present invention with reference to the drawings, where:
[0020] Figure 1 It is a schematic flowchart of a die interconnect network generation method according to an embodiment of the present invention;
[0021] Figure 2 It is a schematic comparison diagram of an inter-die network generated according to an embodiment of the present invention and an inter-die network generated by the prior art;
[0022] Figure 3A comparison schematic diagram of the network-on-chip generated according to an embodiment of the present invention and the network-on-chip generated by the prior art. Detailed implementation manners
[0023] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below through specific embodiments with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0024] As mentioned in the background art section, there are obvious defects in the prior art in aspects such as resource allocation in the design of the interconnected network, topological adaptability in heterogeneous scenarios, and the coordination of physical constraints and design processes.
[0025] To solve the above problems, the inventors propose a method for generating a die interconnect network. In this method, a dedicated die interconnect network is generated according to the communication requirements between dies, the communication requirements within dies, and layout constraint conditions in a specific execution order, so as to rationalize the resource allocation in the chip system, and it is not restricted by the traditional regular topology, and also avoids the waste of cost and bandwidth caused by using the regular topology for a dedicated chip system. At the same time, the introduced layout constraint conditions can pay attention to the physical implementation challenges faced by the chip system during the actual manufacturing process. The generation method includes steps S1 - S3. Among them, in step S1, the die layout is determined based on a plurality of dies and layout constraint conditions; in step S2, an inter-die network is generated based on the die layout generated in step S1 and the communication requirements between dies; in step S3, a network-on-chip is generated based on the inter-die network generated in step S2 and the communication requirements within dies.
[0026] Generally speaking, as Figure 1As shown in the figure, the present invention provides a method for generating a die interconnect network for generating a die interconnect network required to form a complete chip based on a plurality of dies. The method includes: Step S1: Determine the die layout based on a plurality of dies required to build the chip and layout constraint conditions by using a preset first strategy; wherein, one or more functional modules are integrated in the die, and the communication requirements between the functional modules are known; the die layout represents the positional arrangement relationship between the dies; Step S2: Determine the number of die interfaces on each die and the number of inter-die routers based on the communication requirements between the functional modules in each die and the functional modules in other dies, and determine the communication paths between each die and other dies by using a preset second strategy based on the number of die interfaces of each die, the number of inter-die routers, and the die layout to obtain an inter-chip network; wherein, the inter-die routers are responsible for the communication between the dies; the communication path between each die and other dies includes the communication links between all the die interfaces on this die and the corresponding die interfaces on other dies; Step S3: Determine the number of intra-die routers based on the communication requirements between the functional modules in each die and the inter-chip network, and determine the internal communication path of each die by using a preset second strategy based on the inter-chip network and the number of intra-die routers to obtain an on-chip network; wherein, the internal communication path of each die includes the communication links between the die interfaces within the die, the communication links between the functional modules, and the communication links between the functional modules and the die interfaces within the die; wherein, the die layout, the inter-chip network, and the on-chip network constitute the die interconnect network.
[0027] To better understand the present invention, the following will specifically describe each step in detail with reference to specific embodiments.
[0028] I. Step S1
[0029] In the step S1, determine the die layout based on a plurality of dies required to build the chip and layout constraint conditions by using a preset first strategy; wherein, one or more functional modules are integrated in the die, and the communication requirements between the functional modules are known; the die layout represents the positional arrangement relationship between the dies. According to an embodiment of the present invention, the preset first strategy is a simulated annealing algorithm, a genetic algorithm, or a particle swarm algorithm.
[0030] It should be noted that in the step S1, the layout constraint conditions are set according to actual requirements, and the present invention does not impose special restrictions. For example, the die heat dissipation can be used as a layout constraint condition, and the die layout can be obtained by using the simulated annealing algorithm, genetic algorithm or particle swarm algorithm, and it is judged whether the die heat dissipation in the obtained die layout meets the requirements. If not, step S1 is repeated until a die layout that meets the heat dissipation constraint is obtained, and the finally obtained die layout is used as a reference to participate in step S2. Another example is that the connection length between specific dies can be used as a layout constraint condition, and the die layout can be obtained by using the simulated annealing algorithm, genetic algorithm or particle swarm algorithm, and it is judged whether the connection length between specific dies in the obtained die layout meets the requirements. If not, step S1 is repeated until a die layout that meets the connection length constraint is obtained, and the finally obtained die layout is used as a reference to participate in step S2.
[0031] It can be seen from step S1 that the die layout is not a random permutation and combination, but a die layout that meets the requirements needs to be obtained in combination with the layout constraint conditions. And it is possible that a die layout that meets the layout constraint conditions cannot be obtained by executing step S1 once. Therefore, the present invention proposes that step S1 can be iteratively executed multiple times until a die layout that meets the requirements is obtained, and step S2 is executed based on the finally obtained die layout to obtain the inter-die network.
[0032] II. Step S2
[0033] In the step S2, based on the communication requirements between the functional modules in each die and the functional modules in other dies, the number of die interfaces on each die and the number of inter-die routers are determined, and based on the number of die interfaces of each die, the number of inter-die routers and the die layout, a preset second strategy is used to determine the communication path between each die and other dies to obtain the inter-die network; wherein, the inter-die router is responsible for the communication between dies; the communication path between each die and other dies includes the communication links between all die interfaces on this die and the corresponding die interfaces on other dies. Although the inter-die network can be obtained by executing step S2 once, the inter-die network obtained by executing step S2 once is not necessarily the inter-die network with the smallest communication overhead. Therefore, the present invention proposes that step S2 can be iteratively executed multiple times to obtain the inter-die network with the smallest communication overhead, and the inter-die network obtained in the last iteration is used as a reference to participate in step S3 to generate the on-chip network.
[0034] According to an embodiment of the present invention, the preset second strategy is the minimum cut algorithm and the minimum cost maximum flow algorithm.
[0035] According to an embodiment of the present invention, the die interface adopts a serial interface or a parallel interface.
[0036] According to an embodiment of the present invention, the step S2 includes sub-steps S21-S27, and the sub-steps in step S2 will be described one by one below.
[0037] In step S21, based on the communication requirements between the functional modules within each die and the functional modules within other dies, determine the total input / output bandwidth of each die, and divide the total input / output bandwidth of each die by the preset single-die interface bandwidth to determine the number of die interfaces on each die, and evenly distribute the die interfaces on the boundary of the die according to the number of die interfaces on each die; wherein, the horizontal and vertical coordinates represent the positions of each die interface on its own die. It should be noted that the bandwidth of the die interface is set according to the existing process level, and the serial interface and the parallel interface can be set to the same bandwidth.
[0038] In step S22, based on the number of die interfaces of each die and the communication requirements between the functional modules within each die and the functional modules within other dies, determine one or more target die interfaces corresponding to each die interface on each die and other dies; wherein, a die interface and a corresponding target die interface on other dies form an interface pair.
[0039] It should be noted that the load of the die interfaces in each interface pair cannot exceed the established rate limit, and the wiring distance between each die interface and all corresponding target die interfaces should be as short as possible. It should also be noted that the communication rate between a functional module within a die and a functional module within another die may exceed the communication bandwidth of its own single-die interface. Therefore, the communication requirements between these two functional modules may need to be realized by means of two or more die interfaces within their own die. For example, there is a communication requirement between functional module A in die 1 and functional module B in die 2, but the communication bandwidth of the die interfaces in die 1 and die 2 is less than the communication rate requirement between functional module A and functional module B. At this time, it is necessary to use two or more die interfaces on die 1 and die 2 to realize the communication between functional module A and functional module B, that is, at least two pairs of interface pairs are required to realize the communication between functional module A and functional module B.
[0040] In step S23, with the total number of die interfaces on all dies as a constraint, start setting the number of inter-die routers from zero until the number of inter-die routers can support the communication between all functional modules within each die and the functional modules within other dies.
[0041] It should be noted that the reason for setting the number of inter-die routers from scratch is that it can accurately analyze the number of inter-die routers required for the inter-die network and avoid wasting resources. It should also be noted that the wiring resources between dies are relatively scarce. If all die interfaces between dies are connected point-to-point by interconnect lines, there may be a shortage of space or congestion. Therefore, the communication between die interfaces on different dies is forwarded through inter-die routers.
[0042] In step S24, based on the positions of each die interface on all dies, the minimum cut algorithm is used to divide all die interfaces into multiple interface groups; where each interface group includes multiple die interfaces, and all die interfaces in each interface group are connected to the same inter-die router. It can be seen that when grouping, the grouping is carried out according to the positions of each die interface on all dies. Such a grouping method can preferentially allocate the communication of multiple die interfaces with close positions to the same inter-die router for forwarding, thereby saving wiring resources.
[0043] In step S25, obtain the communication rate of all die interfaces in each interface group, calculate the weighted sum of the product of the communication rate of all die interfaces in each interface group and their abscissas and the weighted sum of the product of the communication rate of all die interfaces in each interface group and their ordinates respectively, and divide the weighted sum of the product of the abscissas and the weighted sum of the product of the ordinates by the sum of the communication rates of all die interfaces in the interface group to obtain the abscissa and ordinate of the inter-die router connected to each interface group, and determine the position of the inter-die router in the die layout based on the abscissa and ordinate. Among them, the solution process of the abscissa and ordinate of the inter-die router corresponding to each interface group can be expressed as follows.
[0044]
[0045]
[0046] Among them, represents the abscissa of the inter-die router, represents the ordinate of the inter-die router, represents the communication rate of the first die interface in the interface group, and respectively represent the abscissa and ordinate of the first die interface in the interface group; represents the communication rate of the second die interface in the interface group, and respectively represent the abscissa and ordinate of the second die interface in the interface group; represents the communication rate of the (n + 1)-th die interface in the interface group, and respectively represent the abscissa and ordinate of the (n + 1)-th die interface in the interface group.
[0047] In step S26, based on the die layout, determine the positions of the dies where the two die interfaces in each interface pair are located, and select a matching interface type based on the positions of the dies where the two die interfaces in each interface pair are located to obtain multiple different interface insertion schemes; wherein, each interface insertion scheme represents the interface types adopted by all die interfaces, and the two die interfaces in each interface pair adopt the same interface type.
[0048] It should be noted that the serial interface and the parallel interface cannot be randomly assigned. It is necessary to analyze whether a certain interface can be used according to the positions of the dies where the two die interfaces in the die layout are located. For example, if there is a communication with a rate not exceeding the bandwidth between non-adjacent dies (the die layout result shows that the two dies have no overlapping edges), there is no possibility of using a pair of parallel interfaces for connection. There are only the following two schemes: using a pair of serial interfaces or using a series of parallel interfaces through other dies for transit. Another example is that if the communication distance of the serial interface is x and the communication distance of the parallel interface is y, when the distance d between the dies where the two die interfaces are located does not meet the communication distance of the parallel interface but meets the communication distance of the serial interface (x > d > y), then serial interface communication or communication through a series of parallel interfaces through other dies for transit is adopted.
[0049] Step S27: Select the interface insertion scheme with the lowest power consumption and latency from multiple different interface insertion schemes, construct an inter-die router cost graph based on the selected interface insertion scheme, and calculate the communication link between each interface pair using the minimum cost maximum flow algorithm based on the preset inter-die path constraint conditions; wherein, the inter-die router cost graph includes multiple nodes and multiple edges connecting the two nodes. The nodes represent die interfaces or inter-die routers, the edges represent the existence of a communication link between the two nodes, and each edge uses the communication cost between the two connected nodes as the weight. Among them, the communication cost of the communication link between each interface pair should be as small as possible, and the communication cost between the two nodes is calculated by the connection distance between the two nodes.
[0050] It should be noted that when constructing the overhead diagram of the inter-die router, for any two inter-die routers, if the distance between the two inter-die routers meets the established connection limit, it means that a communication link can be established between the two inter-die routers; for two die interfaces on the same die, a communication link can be established between the two die interfaces. It should also be noted that the preset inter-die path constraint conditions are determined according to actual needs, and the present invention does not make special restrictions. For example, the preset inter-die path constraint conditions include: restricting that the sum of the communication rates in any direction of any communication link should be less than or equal to the maximum available bandwidth; if the communication between two die interfaces needs to be relayed by other dies, at this time, the outflow of the die where the output die interface is located is abstracted as 1, and the outflow of the die where the input die interface is located is abstracted as -1, then the outflow of the relaying die passed through should be 0; restricting that the number of die interfaces connected to the inter-die router is less than or equal to its maximum connection number.
[0051] As can be seen from step S2, the communication between die interfaces on different dies needs to be forwarded by the inter-die router, and when the distance between the dies is relatively far, in addition to being forwarded by the inter-die router, it may also need to be relayed by the die interfaces on one or more dies. Based on this, the present invention proposes to generate an inter-chip network after completing the die layout. The reason for such processing is that if the on-chip network is generated first and then the inter-chip network is generated, since the load brought by inter-die communication is not considered when generating the on-chip network, it may cause congestion in the entire chip system.
[0052] To better understand the difference between the present invention and the prior art, the following is an illustration with Figure 2 the content shown. In Figure 2 , "R" represents the router (inter-die), the black rectangles on the die boundaries represent the die interfaces, the red lines represent the connections between the routers, and the green lines represent the connections between the die interfaces and the routers. Among them, Figure 2 Figure (a) shows the inter-chip network generated according to the prior art; Figure 2 Figure (b) shows the inter-chip network generated according to step S2 of the present invention.
[0053] As can be seen from Figure 2 Figure (a), in the inter-chip network generated according to the prior art, each die interface on each die is connected to an inter-die router (for the sake of simplified representation, Figure 2 the die interfaces on each die in Figure (a) are not drawn), and the inter-die routers are interconnected by a mesh. In this way, the number of inter-die routers in the generated inter-chip network may be redundant, and if the communication allocation from the die interface to the inter-die router is not reasonably designed, it may exceed the load of the inter-die router, resulting in congestion in the inter-chip network.
[0054] As can be seen from Figure 2 (b), different from the prior art, in step S2 proposed by the present invention, the inter-die network is generated guided by the communication requirements between dies. The inter-die network generated in this way can avoid redundant communication links while meeting the communication requirements between dies, thereby minimizing the communication overhead between dies as much as possible and avoiding the additional overhead caused by redundant numbers of routers between dies.
[0055] III. Step S3
[0056] In the step S3, based on the communication requirements between functional modules within each die and the inter-die network, the number of routers within the die is determined, and based on the inter-die network and the number of routers within the die, a preset second strategy is adopted to determine the internal communication path of each die to obtain the on-chip network; wherein, the internal communication path of each die includes communication links between die interfaces within the die, communication links between functional modules, and communication links between functional modules and die interfaces within the die. Although the on-chip network can be obtained by executing step S3 once, the on-chip network obtained by executing step S3 once is not necessarily the on-chip network with the minimum communication overhead. Therefore, the present invention proposes to iteratively execute step S3 multiple times to obtain the on-chip network with the minimum communication overhead, and the on-chip network obtained from the last iteration, the die layout obtained by iteratively executing step S1 multiple times, and the inter-die network obtained by iteratively executing step S2 multiple times form the die interconnection network.
[0057] According to an embodiment of the present invention, the step S3 includes steps S31 - S35, and the sub-steps in step S3 will be introduced below. [[ID=I4]]
[0058] In step S31, based on the inter-die network, it is determined whether there are communication requirements between any two die interfaces within each die and between die interfaces and functional modules. As can be seen from the foregoing step S2, the communication between die interfaces on different dies in the inter-die network may need to be relayed by one or more die interfaces on other dies. Therefore, it is necessary to analyze the possible communication requirements between any two die interfaces within each die; and the communication between die interfaces on different dies actually belongs to the communication between functional modules on different dies completed through die interfaces. Therefore, it is also necessary to analyze the possible communication requirements between functional modules and die interfaces within each die based on the inter-die network.
[0059] In step S32, based on the communication requirements between die interfaces within each die, between die interfaces and functional modules, and between functional modules within the die, the number of routers within the die is set starting from zero until the number of routers within the die can meet the communication requirements between die interfaces within the die, between die interfaces and functional modules, and between functional modules within the die.
[0060] Similar to step S23, the reason for setting up the intra-die router from scratch is that it can accurately analyze the number of intra-die routers required for the on-chip network and avoid wasting resources. Moreover, the routing resources within the die are relatively scarce. If all the connections between the die interfaces within the die, between the die interfaces and the functional modules are point-to-point connections using interconnect lines, there may be a shortage of space or congestion. Therefore, the communication between the die interfaces within the die, between the die interfaces and the functional modules, and between the functional modules is all forwarded through the intra-die router.
[0061] In step S33, based on the positions of each functional module and die interface within the die, the minimum cut algorithm is used to divide all the die interfaces and all the functional modules within each die into multiple communication groups. Among them, each communication group includes multiple die interfaces and / or multiple functional modules, and all the die interfaces and all the functional modules in each communication group are connected to the same intra-die router. It can be seen that when grouping, the grouping is carried out according to the positions of the functional modules and die interfaces within the die. Such a grouping method can give priority to allocating the communication of die interfaces and functional modules with similar positions within each die to the same intra-die router for forwarding, thereby saving routing resources.
[0062] In step S34, obtain the communication rates of all the die interfaces and all the functional modules connected to the corresponding intra-die router in each communication group, and calculate the weighted sum of the product of the communication rate of all the die interfaces in each communication group and their abscissas and the product of the communication rate of all the functional modules and their abscissas respectively, and the weighted sum of the product of the communication rate of all the die interfaces in each communication group and their ordinates and the product of the communication rate of all the functional modules and their ordinates respectively. Then, divide the weighted sum of the abscissas and the weighted sum of the ordinates by the sum of the communication rates of all the die interfaces and all the functional modules in this communication group to obtain the abscissa and ordinate of the intra-die router connected to each communication group, and determine the position of the intra-die router on the die based on the abscissa and ordinate. Among them, the abscissa and ordinate of the functional module represent the position of the functional module on the die where it is located. Among them, the solution process of the abscissa and ordinate of the intra-die router corresponding to each communication group can be expressed as follows.
[0063]
[0064]
[0065] Among them, represents the abscissa of the intra-die router, represents the ordinate of the intra-die router, represents the communication rate of the first die interface in the communication group, and respectively represent the horizontal and vertical coordinates of the first die interface in the communication group; represents the communication rate of the second die interface in the communication group, and respectively represent the horizontal and vertical coordinates of the second die interface in the communication group; represents the communication rate of the (n + 1)-th die interface in the communication group, and respectively represent the horizontal and vertical coordinates of the (n + 1)-th die interface in the communication group; represents the communication rate of the first functional module in the communication group, and respectively represent the horizontal and vertical coordinates of the first functional module in the communication group; represents the communication rate of the second functional module in the communication group, and respectively represent the horizontal and vertical coordinates of the second functional module in the communication group; represents the communication rate of the (n + 1)-th functional module in the communication group, and respectively represent the horizontal and vertical coordinates of the (n + 1)-th functional module in the communication group.
[0066] In step S35, construct an on-die router cost graph, and calculate the communication links between die interfaces within each die, between die interfaces and functional modules, and between functional modules within the die by using the minimum cost maximum flow algorithm based on the preset on-die path constraint conditions; wherein, the on-die router cost graph includes multiple nodes and multiple edges connecting two nodes, the nodes represent die interfaces, on-die routers or functional modules, the edges represent that there is a communication link between two nodes, and each edge takes the communication cost between the two connected nodes as the weight.
[0067] It should be noted that the types of each die interface on each die have been determined in step S2. Therefore, an on-die router cost graph can be directly constructed to calculate all communication links, and when calculating, it should be ensured as much as possible that the cost of all communication links is minimized. It should also be noted that the preset on-die path constraint conditions are determined by actual requirements, and the present invention does not make special restrictions. For example, the preset on-die path constraint condition can be that the communication data stream forwarded by the on-die router does not exceed its own limit.
[0068] It can be seen from step S3 that step S3 generates an on-chip network based on the inter-die network generated in step S2. Since the die interfaces of any intermediate die used for inter-die communication are already determined, when generating the on-chip network, the load brought by inter-die communication can be fully considered to avoid causing congestion in the chip system.
[0069] To better understand the differences between the present invention and the prior art, the following will be described with reference to Figure 3 the content shown. In Figure (3), "core" represents the functional modules within the die, the black rectangles on the die boundary are die interfaces, "R" represents the router (within the die), the green lines represent the connections between the die interfaces and the routers, and the red lines represent the connections between the functional modules and the routers as well as between the routers. Among them, Figure 3 Figure (a) shows the network-on-chip generated according to the prior art; Figure 3 Figure (b) shows the network-on-chip generated according to step S3 of the present invention.
[0070] As can be seen from Figure 3 Figure (a), in the network-on-chip generated according to the prior art, the functional modules within the die are connected to a router within the die, and the routers within the die are interconnected by a mesh. In this way, the number of routers within the die in the generated network-on-chip may be redundant, and if the communication distribution from the die interface to the router within the die is not reasonably designed, it may exceed the load of the router within the die, resulting in congestion in the network-on-chip.
[0071] As can be seen from Figure 3 Figure (b), different from the prior art, in step S3 of the present invention, the network-on-chip is generated guided by the communication requirements between the functional modules within the die, between the die interfaces, and between the die interface and the functional module. The network-on-chip generated in this way can avoid redundant communication links while meeting the communication requirements between the functional modules within the die, between the die interfaces, and between the die interface and the functional module, thereby minimizing the communication overhead between the functional modules within the die, between the die interfaces, and between the die interface and the functional module, and avoiding the additional overhead caused by the redundancy of the number of routers within the die. Moreover, since the inter-die network has been determined before the generation of the network-on-chip, the load brought by inter-die communication can be fully considered when generating the network-on-chip, avoiding congestion in the network-on-chip.
[0072] Based on the die interconnect network generation method proposed in the foregoing embodiments, the present invention further provides a die interconnect network generation system, which includes: a layout generation module, configured to determine the die layout based on a plurality of dies required for constructing a chip and layout constraint conditions by using a preset first strategy; wherein, one or more functional modules are integrated in each die, and the communication requirements between the functional modules are known; the die layout represents the positional arrangement relationship between the dies; an inter-die network generation module, configured to determine the number of die interfaces and the number of inter-die routers on each die based on the communication requirements between the functional modules in each die and the functional modules in other dies, and determine the communication paths between each die and other dies based on the number of die interfaces, the number of inter-die routers, and the die layout of each die by using a preset second strategy to obtain an inter-die network; wherein, the inter-die routers are responsible for the communication between the dies; the communication path between each die and other dies includes the communication links between all the die interfaces on this die and the corresponding die interfaces on other dies; an on-die network generation module, configured to determine the number of routers within each die based on the communication requirements between the functional modules within each die and the inter-die network, and determine the internal communication paths of each die based on the inter-die network and the number of routers within each die by using a preset second strategy to obtain an on-die network; wherein, the internal communication path of each die includes the communication links between the die interfaces within the die, the communication links between the functional modules, and the communication links between the functional modules and the die interfaces within the die; wherein, the die layout, the inter-die network, and the on-die network constitute the die interconnect network.
[0073] It should be noted that the die interconnect network generation method and the die interconnect network generation system proposed by the present invention are applicable to both active processes and passive processes.
[0074] Based on the foregoing embodiments, it can be seen that, compared with the prior art, the present invention sequentially generates a chip layout, an inter-die network, and an on-die network according to specific execution sequences based on the inter-die communication requirements, the intra-die communication requirements, and the layout constraint conditions, so as to rationalize the resource allocation in the chip system, is not limited by the traditional regular topology, and also avoids the waste of cost and bandwidth caused by using the regular topology for a dedicated chip system. At the same time, the introduced constraint conditions can pay attention to the physical implementation challenges faced by the chip system during the actual manufacturing process.
[0075] The beneficial effects of the present invention are as follows: (1) The die layout, the inter-die network, and the on-die network are sequentially generated in a specific execution sequence, avoiding the congestion of the on-die network caused by not considering the inter-die communication; (2) When generating the inter-die network and the on-die network, the corresponding inter-die network and on-die network are generated guided by the communication requirements, so as to avoid the redundancy of the number of inter-die routers and intra-die routers, and further avoid the redundant communication links on the inter-die network and the on-die network, reducing the communication overhead of the inter-die network and the on-die network.
[0076] It should be noted that although the above steps are described in a specific order, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently or even in a different order, as long as the required functions can be achieved.
[0077] The present invention may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present invention.
[0078] The computer-readable storage medium may be a tangible device that retains and stores instructions for use by an instruction execution device. The computer-readable storage medium may include, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in a groove having instructions stored thereon, and any suitable combination of the foregoing.
[0079] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of technologies in the market, or to enable other ordinary skill in the technical field to understand the embodiments disclosed herein.
Claims
1. A method for generating a die interconnect network, which is used to generate a die interconnect network required to form a complete chip based on multiple dies, and is characterized in that, The method includes: Step S1: Determine the die layout by using a preset first strategy based on multiple dies required for constructing a chip and layout constraint conditions; wherein, one or more functional modules are integrated in each die, and the communication requirements between the functional modules are known; the die layout represents the positional arrangement relationship between the dies; Step S2: Determine the number of die interfaces on each die and the number of inter-die routers based on the communication requirements between the functional modules in each die and the functional modules in other dies, and determine the communication paths between each die and other dies by using a preset second strategy based on the number of die interfaces, the number of inter-die routers, and the die layout of each die to obtain an inter-chip network; wherein, the inter-die routers are responsible for the communication between the dies; the communication path between each die and other dies includes the communication links between all the die interfaces on this die and the corresponding die interfaces on other dies; Step S3: Determine the number of intra-die routers based on the communication requirements between the functional modules in each die and the inter-chip network, and determine the internal communication path of each die by using a preset second strategy based on the inter-chip network and the number of intra-die routers to obtain an intra-chip network; wherein, the internal communication path of each die includes the communication links between the die interfaces within the die, the communication links between the functional modules, and the communication links between the functional modules and the die interfaces within the die; Wherein, the die layout, the inter-chip network, and the intra-chip network form a die interconnect network.
2. The method according to claim 1, wherein The preset second strategy is the minimum cut algorithm and the minimum cost maximum flow algorithm.
3. The method according to claim 2, wherein The die interface adopts a serial interface or a parallel interface.
4. The method according to claim 3, wherein The step S2 includes: Step S21: Determine the total input / output bandwidth of each die based on the communication requirements between the functional modules in each die and the functional modules in other dies, and divide the total input / output bandwidth of each die by the preset single die interface bandwidth to determine the number of die interfaces on each die, and evenly distribute the die interfaces on the boundary of this die according to the number of die interfaces on each die; wherein, the horizontal and vertical coordinates represent the positions of each die interface on the die where it is located; Step S22: Determine one or more target die interfaces corresponding to each die interface on each die and the die interfaces on other dies based on the number of die interfaces of each die and the communication requirements between the functional modules in each die and the functional modules in other dies; wherein, a die interface and a corresponding target die interface on other dies form an interface pair; Step S23: With the total number of die interfaces on all dies as a constraint, start setting the number of inter-die routers from zero until the number of inter-die routers can support the communication between all the functional modules in each die and the functional modules in other dies; Step S24: Divide all the die interfaces into multiple interface groups by using the minimum cut algorithm based on the positions of each die interface on all dies; wherein, each interface group includes multiple die interfaces, and all the die interfaces in each interface group are connected to the same inter-die router; Step S25: Obtain the communication rates of all die interfaces in each interface group, calculate the weighted sum of the product of the communication rate of each die interface in each interface group and its own abscissa and the weighted sum of the product of the communication rate of each die interface in each interface group and its own ordinate respectively, and divide the weighted sum of the product of the abscissa and the weighted sum of the product of the ordinate by the sum of the communication rates of all die interfaces in the interface group to obtain the abscissa and ordinate of the inter-die router connected to each interface group, and determine the position of the inter-die router on the die layout based on the abscissa and ordinate; Step S26: Determine the positions of the die where the two die interfaces in each interface pair are located based on the die layout, and select a matching interface type based on the positions of the die where the two die interfaces in each interface pair are located to obtain multiple different interface insertion schemes; among them, each interface insertion scheme represents the interface types adopted by all die interfaces, and the two die interfaces in each interface pair adopt the same interface type; Step S27: Select the interface insertion scheme with the lowest power consumption and latency from multiple different interface insertion schemes, construct an inter-die router cost graph based on the selected interface insertion scheme, and calculate the communication link between each interface pair using the minimum cost maximum flow algorithm based on the preset inter-die path constraint conditions; among them, the inter-die router cost graph includes multiple nodes and multiple edges connecting two nodes, the nodes represent die interfaces or inter-die routers, the edges represent the existence of a communication link between two nodes, and the weight of each edge is the communication cost between the two connected nodes.
5. The method according to claim 4, wherein The said Step S3 includes: Step S31: Based on the inter-chip network, determine whether there is a communication requirement between any two die interfaces within each die, between a die interface and a functional module; Step S32: Based on the communication requirements between die interfaces within each die, between a die interface and a functional module, and between functional modules within the die, start setting the number of intra-die routers from zero until the number of intra-die routers can meet the communication requirements between die interfaces within the die, between a die interface and a functional module, and between functional modules within the die; Step S33: Based on the positions of the functional modules and die interfaces within each die, use the minimum cut algorithm to divide all die interfaces and all functional modules within each die into multiple communication groups; among them, each communication group includes multiple die interfaces and / or multiple functional modules, and all die interfaces and all functional modules in each communication group are connected to the same intra-die router; Step S34: Obtain the communication rates of all die interfaces and all functional modules connected to the in-die router in each communication group. Calculate the product of the communication rate of all die interfaces in each communication group and its own abscissa, and the weighted sum of the product of the communication rate of all functional modules and its own abscissa, as well as the product of the communication rate of all die interfaces in each communication group and its own ordinate, and the weighted sum of the product of the communication rate of all functional modules and its own ordinate. Then, divide the weighted sum of the abscissa and the weighted sum of the ordinate by the sum of the communication rates of all die interfaces and all functional modules in this communication group to obtain the abscissa and ordinate of the in-die router connected to each communication group, and determine the position of the in-die router on the die based on the abscissa and ordinate; wherein, the abscissa and ordinate of the functional module represent the position of the functional module on its own die. Step S35: Construct an in-die router overhead graph, and calculate the communication links between die interfaces within each die, between die interfaces and functional modules, and between functional modules within the die using the minimum cost maximum flow algorithm based on the preset in-die path constraint conditions; wherein, the in-die router overhead graph includes multiple nodes and multiple edges connecting two nodes. The nodes represent die interfaces, in-die routers, or functional modules, and the edges represent the existence of a communication link between two nodes, and the weight of each edge is the communication overhead between the two connected nodes.
6. The method according to claim 5, wherein The preset first strategy is the simulated annealing algorithm, genetic algorithm, or particle swarm algorithm.
7. The method according to claim 6, characterized in that, The method can iteratively execute steps S1, S2, and S3 multiple times to obtain the die layout, inter-die network, and on-die network; wherein, in step S2, an inter-die network is generated based on the die layout obtained by executing step S1 for the last time; in step S3, an on-die network is generated based on the inter-die network obtained by executing step S2 for the last time.
8. A die interconnect network generation system based on the method according to any one of claims 1-7, characterized in that, The system includes: A layout generation module, configured to determine the die layout based on multiple dies required for constructing the chip and layout constraint conditions using a preset first strategy; wherein, one or more functional modules are integrated within each die, and the communication requirements between the functional modules are known; the die layout represents the positional arrangement relationship between the dies. An inter-die network generation module, configured to determine the number of die interfaces on each die and the number of inter-die routers based on the communication requirements between functional modules within each die and functional modules within other dies, and determine the communication paths between each die and other dies using a preset second strategy based on the number of die interfaces of each die, the number of inter-die routers, and the die layout to obtain an inter-die network; wherein, the inter-die router is responsible for communication between dies; the communication path between each die and other dies includes the communication links between all die interfaces on this die and the corresponding die interfaces on other dies. An on-chip network generation module is configured to determine the number of routers within each die based on the communication requirements between functional modules within each die and the inter-die network, and determine the internal communication paths of each die by using a preset second strategy based on the inter-die network and the number of routers within the die to obtain an on-chip network; wherein, the internal communication paths of each die include communication links between die interfaces within the die, communication links between functional modules, and communication links between functional modules and die interfaces within the die. Among them, the die layout, the inter-die network, and the on-chip network constitute the die interconnection network.
9. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and the computer program can be executed by a processor to implement the steps of the method according to any one of claims 1-7.
10. An electronic device, characterized in that, Comprising: One or more processors, and a memory, wherein the memory is used to store executable instructions; The one or more processors are configured to implement the steps of the method according to any one of claims 1-7 by executing the executable instructions.