Cross-chip processing system and its routing method

By designing efficient interconnection ports and routing solutions in cross-chip processing systems, the problems of large latency and poor performance of traditional server clusters are solved, and system connections with low latency, high reliability and high bandwidth utilization are achieved.

CN113825202BActive Publication Date: 2025-06-10VIA ALLIANCE SEMICON CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111141627.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-28
Publication Date
2025-06-10
Estimated Expiration
2041-09-28

AI Technical Summary

Technical Problem

The server cluster built with traditional Ethernet networks has large latency, poor system performance and high cost, making it difficult to achieve system connections with low latency, high reliability and high bandwidth utilization.

Method used

A cross-chip processing system is designed to form a system that communicates with each other and flexibly calls resources by providing interconnection ports between chips and packages, and to specifically disclose the routing schemes for each node in the system. The system uses efficient interconnect ports, such as ZPI and ZDI interconnect ports, combined with routing registers and routing tables, to achieve efficient routing of packets.

Benefits of technology

System connections with low latency, high reliability and high bandwidth utilization are achieved, improving system efficiency and reducing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113825202B_ABST
    Figure CN113825202B_ABST
Patent Text Reader

Abstract

The present invention discloses a cross-chip processing system and its routing design. A chip serving as a source node includes an interconnect bus and a microprocessor. The interconnect bus is provided with a routing register. When the microprocessor requests the source node to output a packet for transmission to a target node in the cross-chip processing system, a routing information carrying a routing path from the source node to the target node is stored in the routing register, and then carried to the header of the packet from the routing register and output from the source node along with the packet to guide the transmission of the packet in the cross-chip processing system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a cross-chip processing system, and particularly to a routing scheme between chips. Background Art

[0002] The traditional technology uses an Ethernet network to connect multiple systems. The server group implemented by the Ethernet network can provide powerful computing capabilities. However, the server group constructed by the Ethernet network has large latency, poor system performance, and high costs.

[0003] This technical field requires a system connection technology with low latency, high reliability, and high bandwidth utilization. Summary of the Invention

[0004] The present invention discloses a cross-chip processing system, which provides an interconnection interface between chips and between packages, forming a system that can communicate with each other and flexibly call resources. The present invention particularly discloses the routing scheme of each node in the system.

[0005] A cross-chip processing system implemented according to an embodiment of the present application includes a first chip. The first chip has a first interconnection bus and a first microprocessor coupled to the first interconnection bus. The first interconnection bus has a first routing register. When the first microprocessor requests to use the first chip as a source node to output a packet for transmission in the cross-chip processing system and deliver it to a target node, a routing information carrying a routing path from the source node to the target node is stored in the first routing register, and then carried from the first routing register to the header of the packet and output from the source node along with the packet to guide the transmission of the packet in the cross-chip processing system.

[0006] In one embodiment, the first interconnection bus further has a storage space for storing a routing table listing the routing paths from the source node to other nodes in the cross-chip processing system. In one embodiment, the routing table is burned (fixed) by the manufacturer according to the preset architecture of the cross-chip processing system. In one embodiment, the routing table is established (dynamically configured) by scanning the current architecture of the cross-chip processing system when the cross-chip processing system is started. In one embodiment, the routing information is marked with a first bit indicating whether the carried routing path is fixed or dynamically configured.

[0007] In one embodiment, the routing information is marked with a path valid bit number indicating the valid bit number of the routing path carried by the routing information. The path valid bit number is adjusted as the packet is transmitted in the cross-chip processing system to represent the progress of the journey.

[0008] In one embodiment, when the first microprocessor requests that the packet output by the source node be used as a request, it further evaluates whether a completion notification of the request is to be returned along the original path and marks it with a second bit of the routing information. The destination node decides whether to return the completion notification to the source node along the original path according to the second bit.

[0009] In one embodiment, the routing path carried by the routing information includes output interface information of intermediate nodes, which is used to count clockwise or counterclockwise from an input interface of a relative intermediate node to obtain an output interface. A third bit of the routing information indicates whether clockwise or counterclockwise counting is used. When the destination node identifies that the second bit indicates returning the completion notification to the source node along the original path, it further reverses the third bit. The routing path carried by the routing information includes the interface number of an output interface of the source node. The routing path carried by the routing information includes the interface number of an input interface of the destination node.

[0010] One embodiment is to implement the above technology as a routing method for a cross-chip processing system.

[0011] Specific embodiments are given below and in conjunction with the accompanying drawings, the content of the present invention is described in detail. Description of the Drawings

[0012] Figure 1 For an embodiment of the ZPI interconnect interface, where two packages socket0 and socket1 are connected by the ZPI interconnect interface (labeled ZPI in the figure);

[0013] Figures 2A - 2C Diagrammatic illustration of other planar interconnect embodiments implemented by the interconnect interface ZPI between packages;

[0014] Figure 3A 、 3B Diagrammatic illustration of a three-dimensional (3D) interconnect embodiment implemented by the interconnect interface ZPI between packages;

[0015] Figure 4 For an embodiment of the ZDI interconnect interface;

[0016] Figures 5A - 5C Diagrammatic illustration of a planar interconnect embodiment, where the packages are connected by the interconnect interface ZPI, and the chips inside the packages are connected by the interconnect interface ZDI;

[0017] Figure 6A 、 6B Diagrammatic illustration of a three-dimensional interconnect embodiment, where the packages are connected by the interconnect interface ZPI, and the chips inside the packages are connected by the interconnect interface ZDI;

[0018] Figure 7Illustrate a package 700, which includes a chipset 702 and other chips (such as, computing nodes, coprocessors, and accelerators);

[0019] Figure 8 Illustrate a routing design according to an embodiment of the present application;

[0020] Figure 9 Describe the transmission of routing information between nodes according to an embodiment of the present application;

[0021] Figure 10 Illustrate a format 1000 of routing information according to an embodiment of the present application, which includes 32 bits;

[0022] Figure 11 Take Figure 3B the routing path of the three-dimensional interconnection architecture socket0 → socket4 → socket5 → socket7 as an example, list the changes of routing information on the journey; and

[0023] Figure 12 Illustrate the editing technology and usage technology of routing information of a source node according to an embodiment of the present application. Detailed Embodiments

[0024] The following description lists various embodiments of the present invention. The following description introduces the basic concepts of the present invention and is not intended to limit the content of the present invention. The actual scope of the invention should be defined by the claims.

[0025] The present invention discloses a cross-chip processing system, which provides an interconnection interface between chips and between packages, forming a system that can communicate with each other and flexibly call resources. The present invention particularly discloses the routing scheme of each node in the system.

[0026] The present invention applies high-performance interconnection interfaces, including the interconnection interface between packages (socket-to-socket interconnect interface) and the interconnection interface between chips within each package (die-to-die interconnect interface).

[0027] First, introduce the interconnection interface between packages (socket-to-socket interconnect interface), which will be named ZPI interconnection interface hereinafter.

[0028] Figure 1An implementation of the ZPI interconnect interface, in which two packages, socket0 and socket1, are connected by the ZPI interconnect interface (labeled ZPI in the figure). Each package is shown to include two clusters, labeled cluster0 and cluster1: other implementations may have other numbers of clusters. Each cluster includes a number of central processing unit (CPU) cores. Each package may have a last level cache (labeled LLC), an interconnect bus 102, and various components (such as an input / output controller 104, a clock module 106, a power consumption module 108, etc.). Each package may also be connected to a dual in-line memory module (labeled DIMM).

[0029] Through the interconnect interface ZPI, the packages socket0 and socket1 form a system, in which the central processing unit cores and input / output resources of all clusters can be uniformly scheduled, and the memory owned by the packages socket0 and socket1 can be uniformly used.

[0030] For example, through the interconnect interface ZPI, the packets (flits) of different package caches have a consistent format. In other embodiments, the packets may also be referred to as packets or packet data packets. In this way, any central processing unit core or input / output device in the system formed by these packages can access any memory resource in the system.

[0031] Figures 2A - 2C Diagram of other planar interconnect embodiments implemented by the interconnect interface ZPI between packages. Figure 2A 、 2B Let the packages form a ring connection through the interconnect interface ZPI. Figure 2A It is a three-package ring connection. Figure 2B It is a four-package ring connection. Compared with Figure 2B , Figure 2C The four packages have more interconnect interfaces ZPI to ensure the shortest communication path between packages. The number of packages can be extended to a larger value.

[0032] Figure 3A 、 3B Diagram of a three-dimensional (3D) interconnect embodiment implemented by the interconnect interface ZPI between packages. Figure 3A Three-layer interconnect is realized; the packages socket0 to socket3 are in one plane (belonging to the same layer), and the front and rear layers are the packages socket4 and socket5 respectively. In addition to being connected to the middle layer plane by the interconnect interface ZPI, the front layer package socket4 is also connected to the rear layer package socket5 by the interconnect interface ZPI to form a ring connection. Figure 3BAchieve two - layer interconnection. The packages socket0 to socket3 on the first - layer plane are connected one - to - one with the packages socket4 to socket7 on the second - layer plane through the interconnection interface ZPI. There can be more layers for three - dimensional interconnection. The number of packages on each plane can be extended to a larger value.

[0033] In addition, the die - to - die interconnect interface between chips is introduced as follows; it will be named the ZDI interconnect interface hereinafter.

[0034] Figure 4 This is an implementation of the ZDI interconnect interface. Two chips Die0 and Die1 within a package 400 are connected through the ZDI interconnect interface (labeled ZDI in the figure). Other implementations can have a larger number of chips packaged together. Each chip can include multiple clusters. Each chip can have a last - level cache LLC, an interconnect bus 402, and various components (such as an input / output controller 404, a clock module 406, a power - consumption module 408, etc.), without limitation.

[0035] The above interconnection interfaces ZPI and ZDI can be used jointly so that chips in multiple packages can communicate with each other.

[0036] Figures 5A - 5C Illustrate an example of planar interconnection, where packages are connected through the interconnection interface ZPI, and chips within the packages are connected through the interconnection interface ZDI. Figure 5A Illustrate three packages with a ring - shaped connection of the interconnection interface ZPI, and chips within each package are connected through the interconnection interface ZDI; in this way, six chips form a system and resources can be shared. Figure 5B Illustrate four packages with a ring - shaped connection of the interconnection interface ZPI, and chips within each package are connected through the interconnection interface ZDI; in this way, eight chips form a system and resources can be shared. Compared with Figure 5B , Figure 5C Having more interconnection interfaces ZPI makes the communication path between chips in different packages the shortest. The number of chips in each package can be variable.

[0037] Figure 6A , 6B Illustrate an example of three - dimensional interconnection, where packages are connected through the interconnection interface ZPI, and chips within the packages are connected through the interconnection interface ZDI. Figure 6A Illustrate three - layer interconnection. Each layer can be a single - package or multi - package plane, and each package can include multiple chips (such as D0, D1); the three - layer planes are connected in a ring through the interconnection interface ZPI, supplemented by the interconnection interface ZDI. Each chip is a node of the system and can control the resources of other nodes. Figure 6BIllustrate a double - layer interconnection. On the double - layer plane, each package includes multiple chips; each chip on the double - layer architecture is a node of the system and can control the resources of other nodes. The above - mentioned three - dimensional interconnection can be extended to more layers, and there can be any number of chips in each package. In the application of a chipset, the interconnection interfaces ZPI and ZDI of the present application can be used as follows.

[0038] Figure 7 Illustrate a package 700, which includes a chipset 702 (a chip) and other chips (such as, computing nodes, coprocessors, and accelerators). The chipset 702 is connected to other chips (such as, computing nodes, coprocessors, and accelerators) through the interconnection interface ZDI. To form a larger system, multiple chipset packages can be connected through the interconnection interface ZPI to form Figures 2A - 2C a planar interconnection architecture, or Figure 3A 、 3B a three - dimensional interconnection architecture. In one implementation, a single package can include multiple chipsets; to form a larger system, such multiple packages can be connected through the interconnection interface ZPI to form Figures 5A - 5C a planar interconnection architecture, or Figure 6A 、 6B a three - dimensional interconnection architecture. The chipset 702 can also be linked to a dual - inline memory module (labeled DIMM)

[0039] The foregoing is various architectures of a cross - chip processing system.

[0040] The following discusses the routing design between nodes in a cross - chip processing system.

[0041] Figure 8 Illustrate a routing design according to an implementation of the present application. The chip 800 includes an interconnection bus 802 and a microprocessor 804 coupled to the interconnection bus 802. The interconnection bus 802 stores a routing table 806. The routing table 806 lists the routing paths from the chip 800 to different nodes. When the chip 800 is the source node of a transmission, it queries the routing table 806 according to the destination node of the transmission, obtains the routing path from the source node to the destination node, forms a routing message and fills it into a routing register 808, and then encodes it into the header of a packet for transmission. As shown in the figure, the packet 810 transmitted by the chip 800 to the next node contains routing information. The routing information is transmitted along with the packet 810 between nodes, guiding the transmission of the packet 810 in the cross - chip processing system.

[0042] Figure 9 Describe the transmission of routing information between nodes according to an implementation of the present application.

[0043] Figure 9The figure shows a source node 902, two intermediate nodes 904, 906, and a destination node 908. The source node 902 transmits a packet 910 containing routing information to the intermediate node 904, and the input routing information of the intermediate node 904 is cached in a routing register 912 (provided by an interconnect bus of a chip that serves as the intermediate node 904). The microprocessor 914 of the intermediate node 904 modifies the routing information in the routing register 912, notes that it has passed through the intermediate node 904, and then passes the modified routing information to the intermediate node 906 along with the packet 916. The routing information carried by the packet 916 is cached in a routing register 918 (provided by an interconnect bus of a chip that serves as the intermediate node 906) as the input routing information of the intermediate node 906. The microprocessor 920 of the intermediate node 906 modifies the routing information in the routing register 918, notes that it has passed through the intermediate node 906, and then passes the modified routing information to the destination node 908 along with the packet 922. The routing information carried by the packet 922 is cached in a routing register 924 (provided by an interconnect bus of a chip that serves as the destination node 908) as the input routing information of the destination node 908.

[0044] The corrections made to the routing information during the transmission process include noting the progress of the journey until the input destination node. In addition, if the transmitted packet is a request, its completion element will require a return. The return of the completion notice can be optionally returned along the original path. The microprocessor 926 of the destination node 908 can query the routing path when the request arrives based on the routing information stored in the routing register 924, and return the completion notice to the source node 902 along the original path accordingly. Figure 9 The concept of routing information maintenance can be extended to examples with other numbers of intermediate nodes.

[0045] In particular, the routing table 806 can be fixedly burned into the interconnect bus 802 by the manufacturer according to the preset architecture of the cross-chip processing system, or can be formed in a more flexible dynamic configuration manner.

[0046] In one implementation, a cross-chip processing system including multiple chips, or even multiple packages, can be completed by the manufacturer's architecture. In this way, the manufacturer can customize the routing paths of each node and burn a fixed routing table 806 into the interconnect bus 802.

[0047] A fixed routing path from a source node to a destination node can be formed by following the following rules: first cross the plane; then take the shortest path on the same plane; and when there are multiple shortest path candidates on the same plane, take the next node in the clockwise direction. Figure 3BTaking the three-dimensional interconnection architecture as an example, when the source node is the encapsulated socket0 and the target node is the encapsulated socket7, the routing path can be socket0 → socket4 → socket5 → socket7. The path socket0 → socket4 is the concept of cross-plane priority. The path socket4 → socket5 is the concept of taking the next node in the clockwise direction. The path socket5 → socket7 is the concept of the shortest path. The routing path formation rule can also have other implementation manners. For example, when there are multiple shortest path candidates in the same plane, the next node can be taken in the counterclockwise direction.

[0048] In one implementation manner, the cross-chip processing system scans the current architecture of node interconnection in a software manner, and then generates a routing path for each node accordingly, which is stored in the routing table 806 of the interconnection bus 802. For example, when the cross-chip processing system is started, the basic input / output system (BIOS) can be responsible for node scanning and then establish a routing path between different nodes according to the current node architecture. Such a routing path is called a dynamically configured routing path. The software can also adopt the aforementioned routing path formation rule to form a routing path from a source node to a target node: first cross the plane; then take the shortest path in the same plane; and when there are multiple shortest path candidates in the same plane, take the next node in the clockwise (or counterclockwise) direction.

[0049] Figure 10 According to an implementation manner of the present application, a format 1000 of routing information is illustrated, which includes 32 bits. Bit

[31] indicates whether this routing information is fixed routing information or is generated in a dynamically configured manner. Bit

[30] indicates whether the output interface of the intermediate node is found by clockwise counting or counterclockwise counting relative to the input interface. Bit

[29] indicates whether to start the return of the completion notification along the original path. Bits [15:0] record the routing path. Bits [28:25] indicate the valid number of bits of bits [15:0], which is the valid number of bits of the path. During transmission, each node corrects bits [28:25] and decrements the valid bits to mark the progress of the journey.

[0050] In one implementation manner, bits [15:0] of format 1000 indicate the path in the following manner. Regarding the intermediate node, the corresponding bits in bits [15:0] form a count value, which is used to count clockwise or counterclockwise from an input interface of the intermediate node to obtain an output interface. Regarding the source node, the corresponding bits in bits [15:0] form the interface number of the output interface. Regarding the target node, the corresponding bits in bits [15:0] form the interface number of the input interface.

[0051] Bits [15:0] of format 1000 can also be uniformly edited by interface numbers. In this way, bit

[30] can be a reserved bit without clockwise or counterclockwise information. In addition, the definitions and controls of each bit of the routing information can be adjusted.

[0052] Figure 11 Take Figure 3B the routing path of the three-dimensional interconnection architecture socket0 → socket4 → socket5 → socket7 as an example to list the changes of the routing information along the way.

[0053] The encapsulation of socket0 is the source node, and the encapsulation of socket7 is the destination node. The encapsulation of socket0 queries the routing table corresponding to the encapsulation of socket7 to find the routing path, so as to establish the routing information 1102 in the routing register of the source node socket0. The routing information 1102 uses the "1" in bit

[31] to indicate that this routing information 1102 is generated by software dynamic configuration, not provided by the manufacturer. The routing information 1102 uses "0111" (value 7) in bits [28:25] to indicate that for the encapsulation of socket0, only bits [7:0] in bits [15:0] are meaningful routing paths. Since the encapsulation of socket0 is the source node, the routing path information related to the encapsulation of socket0 is interpreted by the interface number of the encapsulation of socket0. The encapsulation of socket0 outputs the packet through interface 0 (connected to the encapsulation of socket4) marked by "00" in bits [7:6]. In particular, the encapsulation of socket0 modifies the routing information 1102 into the routing information 1104 and then inserts it into the packet and outputs it to the encapsulation of socket4. Compared with the routing information 1102, bit

[29] of the routing information 1104 is modified to "1" to indicate the need for a return along the original path, and bits [28:25] are modified to "0101", indicating that for the encapsulation of socket4, only bits [5:0] in bits [15:0] are meaningful routing paths.

[0054] The encapsulation of socket4 is an intermediate node, and bits [5:4] of the path information 1104 are used as a count value. The "0" marked by bit

[30] of the path information 1104 represents clockwise counting for the corresponding packet input interface, and then the packet output interface is found. "00" in bits [5:4] points to interface 0 (connected to the encapsulation of socket5) counted clockwise from the input interface. In particular, the encapsulation of socket4 modifies the routing information 1104 into the routing information 1106 and then inserts it into the packet and outputs it to the encapsulation of socket5. Compared with the routing information 1104, bits [28:25] of the routing information 1106 are modified to "0011", indicating that for the encapsulation of socket5, only bits [3:0] in bits [15:0] are meaningful routing paths.

[0055] The encapsulated socket5 is an intermediate node that uses bits [3:2] of the path information 1106 as a count value. The "0" indicated by bit

[30] of the path information 1106 represents clockwise counting for the corresponding packet input interface, and then the packet output interface is found. The "01" in bits [3:2] points to interface number 1 counted clockwise from the input interface (connected to the encapsulated socket7). In particular, the encapsulated socket5 modifies the routing information 1106 into the routing information 1108 and then inserts it into the packet and outputs it to the encapsulated socket7. Compared with the routing information 1106, bits [28:25] of the routing information 1108 are modified to "0001", indicating that for the encapsulated socket7, only bits [1:0] in bits [15:0] are meaningful routing paths.

[0056] The encapsulated socket7 is the destination node that interprets bits [1:0] of the path information 1108 as the interface number for packet input. The "01" in bits [1:0] of the routing information 1108 indicates that the encapsulated socket7 receives packets through interface number 1, and this interface number 1 is exactly the one connected to the encapsulated socket5. The encapsulated socket7 does not change the routing information 1108 and caches the same value of the routing information 1110 in the routing register for use when a completion notification is returned along the original path.

[0057] When a completion notification is returned along the original path, according to the routing information 1110, the encapsulated socket7 makes the completion notification output from interface number 1 (bits [1:0]) and received by the encapsulated socket5. In particular, the encapsulated socket7 modifies the cached routing information 1110 into the routing information 1112 and then outputs it to the encapsulated socket5 along with the completion notification. Bit

[30] of the routing information 1112 is "1", which is used to make the interface interpretation of subsequent intermediate nodes change to counterclockwise. Bits [7:0] are modified to bits [28:25] being "0011", indicating that for the encapsulated socket5, only bits [3:0] in bits [15:0] are meaningful routing paths.

[0058] According to the routing information 1112, the intermediate node socket5 uses bits [3:2] of the path information 1112 as a count value. The "1" indicated by bit

[30] of the path information 1112 represents counterclockwise counting for the corresponding packet input interface, and then the packet output interface is found. The "01" in bits [3:2] points to interface number 1 counted counterclockwise from the input interface (connected to the encapsulated socket4). In particular, the encapsulated socket5 modifies the routing information 1112 into the routing information 1114 and then inserts it into the packet and outputs it to the encapsulated socket4. Compared with the routing information 1112, bits [28:25] of the routing information 1114 are modified to "0101", indicating that for the encapsulated socket4, only bits [5:0] in bits [15:0] are meaningful routing paths.

[0059] According to routing information 1114, intermediate node socket 4 uses bit [5:4] of path information 1114 as a count value. The "1" indicated by bit

[30] of path information 1114 represents that the corresponding packet input interface is counted counterclockwise, and then the packet output interface is found. The "00" of bit [5:4] refers to the numbered 0 interface (connected to encapsulation socket 0) counted counterclockwise from the input interface. In particular, encapsulation socket 4 corrects routing information 1114 to routing information 1116 and then inserts the packet and outputs it to encapsulation socket 0. Compared with routing information 1114, bits [28:25] of routing information 1116 are corrected to "0111", indicating that for encapsulation socket 0, bits [15:0] and only bits [7:0] are meaningful routing paths.

[0060] Encapsulation socket 0 is the target node of the completion notification, and interprets bit [7:6] of path information 1116 as the interface number of the packet input. The "00" of bit [7:6] of routing information 1116 indicates that encapsulation socket 0 receives the packet through interface number 0, and this interface number 0 is connected to encapsulation socket 4. The completion notification is successfully returned along the original path of socket 7 → socket 5 → socket 4 → socket 4.

[0061] The above example may have other variations. The definition of each bit of routing information and its control may be adjusted.

[0062] Figure 12 According to one embodiment of the present application, the 32-bit routing information editing and use technology of a source node are illustrated.

[0063] In step S1202, the source node obtains information of the target node (which may be an address or an ID).

[0064] Step S1204 determines whether to set (pull up to "1") the bit

[31] of the routing information. If not, the process proceeds to step S1206 to set the bits [15:0] of the routing information with the fixed routing path preset by the manufacturer. Step S1208, based on this routing information, the packet is transmitted from the source node to the destination node according to the fixed routing path configured by the manufacturer.

[0065] If bit

[31] of the routing information is set, the process switches to dynamic configuration of the routing path design. Step S1210 determines whether the source node is now transmitting a completion element. If so, the process proceeds to step S1212 to identify whether bit

[29] of the routing information of the original request received is set (to "1"). If so, it means that the completion notification needs to be returned along the original path. The process proceeds to step S1214 to modify the routing information of the original request received, so that bit

[30] is set (raised to "1", switching the direction used for identifying the intermediate node exit, e.g., from clockwise to counterclockwise) and then used as the routing information for the completion notification. Then, the process proceeds to step S1208, and according to this routing information, the completion notification is returned along the original path.

[0066] If step S1212 identifies that bit

[29] of the routing information of the original request is not set (to "0"), the completion notification does not need to be returned along the original path. Regarding the routing information of this completion notification, the process sets bit

[30] to 0 (no need to reverse the exit interpretation direction) in step S1216, and sets bit

[29] to 0 (closing the original path return) in step S1218. Step S1220 sets bits [15:0] with the routing path configured by software. In step S1208, according to this routing information, the completion notification is transmitted from the source node to the target node along this routing path configured by software.

[0067] If step S1210 determines that the source node is now transmitting not a completion element but a request, the process forms the routing information of the request in steps S1222, S1224, and S1220. Step S1222 sets bit

[30] to 0 (no need to reverse the exit interpretation direction), and step S1224 determines whether to set bit

[29] (original path return flag) according to a hardware-customized algorithm. Step S1220 sets bits [15:0] with the routing path configured by software. In step S1208, according to this routing information, the request is transmitted from the source node to the target node along this routing path configured by software.

[0068] One implementation realizes the above technology as a routing method for a cross-chip processing system. Any technology that transmits the foregoing routing information between nodes and manages the routing information with the interconnection bus of each node falls within the technical scope of this application.

[0069] Although the present invention has been disclosed above with preferred embodiments, it is not intended to limit the present invention. Those skilled in the art can make some modifications and refinements without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the scope defined by the appended claims.

Claims

1. A cross-chip processing system for providing an interconnection interface between chips and between packages to enable communication with each other and flexible resource invocation. The system comprises: a first chip having a first interconnection bus and a first microprocessor coupled to the first interconnection bus, wherein the first interconnection bus includes a first routing register, and a second chip serving as a target node having a second interconnection bus and a second microprocessor coupled to the second interconnection bus, wherein the second interconnection bus includes a second routing register, wherein when the first microprocessor outputs a packet with the first chip as the source node for transmission in the cross-chip processing system and delivers it to the target node, the routing information carrying the routing path from the source node to the target node is stored in the first routing register, and is then carried from the first routing register to the header of the packet and output from the source node with the packet to guide the transmission of the packet in the cross-chip processing system; wherein the second routing register stores the routing information carried in the header of the packet received by the target node.

2. The cross-chip processing system according to claim 1, wherein: the first interconnection bus further has a storage space for storing a routing table, and the routing table includes routing paths from the source node to other nodes in the cross-chip processing system.

3. The cross-chip processing system according to claim 2, wherein: the routing table is burned according to a preset architecture of the cross-chip processing system.

4. The cross-chip processing system according to claim 2, wherein: the routing table is established by scanning the architecture of the cross-chip processing system when the cross-chip processing system is started.

5. The cross-chip processing system according to claim 2, wherein: the routing information is marked with a first bit indicating whether the carried routing path is fixed or dynamically configured; according to the preset architecture of the cross-chip processing system, the fixed routing path is burned in the routing table; and the dynamically configured routing path is established in the routing table by scanning the architecture of the cross-chip processing system when the cross-chip processing system is started.

6. The cross-chip processing system according to claim 1, wherein: the routing information is marked with a path valid bit number indicating the valid bit number of the routing path carried by the routing information; and the path valid bit number is adjusted as the packet is transmitted in the cross-chip processing system as a progress of the journey.

7. The cross-chip processing system according to claim 1, wherein: when the first microprocessor requests that the packet output by the source node be used as a request, it further determines whether the completion notification of the request needs to be returned along the original path and marks it with a second bit of the routing information; and the target node decides whether to return the completion notification to the source node along the original path according to the second bit.

8. The cross-chip processing system according to claim 7, wherein: the routing path carried by the routing information includes output interface information of intermediate nodes for counting the output interface in a clockwise or counterclockwise direction relative to the input interface of the intermediate node; and a third bit of the routing information indicates whether clockwise or counterclockwise counting is used.

9. The cross-chip processing system according to claim 8, wherein: When the target node identifies that the second bit indicates a return of the completion notice to the source node along the original path, it further reverses the third bit.

10. The cross-chip processing system according to claim 9, wherein: The routing path carried in the routing information includes the interface number of the output interface of the source node; and The routing path carried in the routing information includes the interface number of the input interface of the target node.

11. The cross-chip processing system according to claim 1, wherein: When the packet received by the target node is used as a request, the routing information stored in the second routing register indicates whether to return the completion notice to the source node along the original path; and When the second microprocessor returns the completion notice, if the routing information stored in the second routing register indicates a return along the original path, the completion notice is returned to the source node according to the routing path in the routing information stored in the second routing register.

12. The cross-chip processing system according to claim 1, further comprising: A third chip, serving as an intermediate node between the source node and the target node, having a third interconnection bus and a third microprocessor coupled to the third interconnection bus, wherein: The packet sent by the source node is transmitted to the target node through the intermediate node; The third interconnection bus includes a third routing register that stores the routing information carried in the header of the packet received by the intermediate node; The third microprocessor corrects the number of valid bits of the path carried in the routing information on the third routing register to indicate the progress of the journey, and updates the corrected routing information from the third routing register to the header of the packet and outputs the packet from the intermediate node along with the packet.

13. A routing method for a cross-chip processing system, used to provide an interconnection interface between chips and between packages for communication with each other and flexible resource invocation. The cross-chip processing system comprises: A first chip, serving as a source node, having a first interconnection bus and a first microprocessor coupled to the first interconnection bus. Among them, the first interconnection bus includes a first routing register, and a second chip, serving as a target node, having a second interconnection bus and a second microprocessor coupled to the second interconnection bus. Among them, the second interconnection bus includes a second routing register. The method includes: Managing the first routing register on the first interconnection bus of the first chip; When requesting the first chip to output a packet as a source node for transmission in the cross-chip processing system and deliver it to the target node, storing the routing information carrying the routing path from the source node to the target node in the first routing register, and loading it from the first routing register to the header of the packet, and outputting the packet from the source node along with the packet to guide the transmission of the packet in the cross-chip processing system; and The second routing register stores the routing information carried in the header of the packet received by the target node.

14. The routing method for a cross-chip processing system according to claim 13, wherein: The routing information is marked with a first bit to indicate that the carried routing path is fixed or dynamically configured; According to the preset architecture of the cross-chip processing system, the fixed routing path is burned into the routing table; The dynamically configured routing path is established in the routing table by scanning the current architecture of the cross-chip processing system when the cross-chip processing system starts up; and The first interconnect bus further has storage space for storing the routing table, and the routing table includes the routing paths from the source node to other nodes of the cross-chip processing system.

15. The routing method of the cross-chip processing system according to claim 13,[[]]END]] wherein: The routing information is marked with the number of valid bits of the path, indicating the number of valid bits of the routing path carried by the routing information; and The number of valid bits of the path is adjusted according to the transfer of the packet in the cross-chip processing system as the progress of the journey.

16. The routing method of the cross-chip processing system according to claim 13, further comprising: When the packet required to be output by the source node is used as a request, determining whether the completion notice of the request needs to be returned along the original path and marking it with the second bit of the routing information; wherein, the destination node decides whether to return the completion notice to the source node along the original path according to the second bit.

17. The routing method of the cross-chip processing system according to claim 16,[[]]END]] wherein: The routing path carried by the routing information includes the output interface information of the intermediate node, which is used to count the output interface in the clockwise or counterclockwise direction relative to the input interface of the intermediate node; and The third bit of the routing information indicates whether the clockwise or counterclockwise direction is used for counting.

18. The routing method of the cross-chip processing system according to claim 17,[[]]END]] wherein: When the destination node identifies that the second bit indicates that the completion notice is to be returned to the source node along the original path, it further reverses the third bit.

19. The cross-chip processing system according to claim 18,[[]]END]] wherein: The routing path carried by the routing information includes the interface number of the output interface of the source node; and The routing path carried by the routing information includes the interface number of the input interface of the destination node.

Citation Information

Patent Citations

  • Memory access processing method and system based on interconnection of memory chips, and memory chips

    CN103902472A

  • Method for routing data packet, node and communication system

    CN105814850A