Processing system and communication method for inter-chip communication

By employing high-bandwidth chip interconnect networks and multipath communication in the chip interconnect network, the problems of slow inter-chip communication speed and insufficient bandwidth are solved, improving data transmission rate and system performance, especially in neural network and artificial intelligence applications.

CN116701292BActive Publication Date: 2026-08-25T-HEAD (SHANGHAI) SEMICON CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210173584.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-24
Publication Date
2026-08-25
Estimated Expiration
2042-02-24

AI Technical Summary

Technical Problem

In chip interconnect networks, existing technologies suffer from slow communication speeds and insufficient bandwidth, especially in applications that process large amounts of information, such as neural network and artificial intelligence workloads. The bandwidth limitations and insufficient speed of the PCIe bus lead to limited system performance.

Method used

Employing a high-bandwidth chip interconnect network (ICN), and through a statically pre-defined routing table and multi-path communication, parallel processing units are allowed to communicate without using the PCIe bus, enabling the parallel and overlapping transmission of multiple communication data packets and bypassing the bandwidth limitations of the PCIe bus.

Benefits of technology

It improves the data transmission rate and bandwidth of inter-chip communication, enhances the system's execution speed and efficiency, and enables more efficient information processing and communication, especially in neural network and artificial intelligence applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116701292B_ABST
    Figure CN116701292B_ABST
Patent Text Reader

Abstract

The present disclosure provides a processing system and a communication method for inter-chip communication, the processing system comprising: a plurality of parallel processing units, the parallel processing units comprised in a first chip, each of the parallel processing units comprising: a plurality of processing cores; a plurality of memories, a first memory group of the plurality of memories coupled with a first processing core group of the plurality of processing cores; a plurality of interconnects located in a chip interconnect network, the plurality of interconnects configured to communicatively couple the plurality of parallel processing units, wherein each of the parallel processing units is configured to communicate through the chip interconnect network according to a respective routing table, the respective routing table stored and resident in a register of the respective parallel processing unit, and the respective routing table comprising information of a plurality of paths to any other given parallel processing unit. The present disclosure enables efficient and effective network communication across chips.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of information processing and communication in chip interconnect networks, and more particularly to processing systems and communication methods for inter-chip communication. Background Technology

[0002] Many electronic technologies, such as digital computers, calculators, audio devices, video devices, and telephone systems, contribute to increased productivity and reduced costs in analyzing and communicating data and information across most sectors of business, science, education, and entertainment. Electronic components are used in many important applications (e.g., medical procedures, vehicle assistance systems, financial applications, etc.) that frequently involve processing and storing large amounts of information. To handle the sheer volume of processing, a system can include many interconnected processing chips. In many applications, it is crucial that the system processes information quickly and accurately. The ability to process information quickly and accurately often depends on communication between the processing chips. Establishing fast and reliable information communication within a network of interconnected chips can present problems and difficulties.

[0003] Figure 1 This is a block diagram illustrating an example conventional system 100 that can be used to accelerate neural networks. Typically, system 100 includes multiple servers, and each server includes multiple parallel computing units (PPUs). Figure 1 In the example, system 100 includes server 101 and server 102. Server 101 includes multiple parallel processing units PPU_0 to PPU_n connected to a Peripheral Component Interconnect Express (PCIe) bus 111, and server 102 includes a similar array of PPUs connected to a PCIe bus 112. Each PPU includes elements such as a processing core and memory (not shown). In one embodiment, the PPU may be a neural network processing unit (NPU). In an exemplary embodiment, multiple NPUs are arranged in a parallel configuration. Each server in system 100 includes a main central processing unit (CPU), and as... Figure 1 As shown, each server in system 100 is connected to network 130 via a corresponding network interface controller card (NIC).

[0004] System 100 includes a unified memory addressing space, which employs, for example, a partitioned global address space (PGAS) programming model. Therefore, in Figure 1 In the example, each PPU of server 101 can read data from or write data to the memory of any other PPU on server 101 or server 102, and vice versa. For example, to write data from PPU_0 to PPU_n on server 101, the data is sent from PPU_0 to PPU_n via PCIe bus 111; and to write data from PPU_0 on server 101 to the memory of PPU_m on server 102, the data is sent from PPU_0 to NIC 121 via PCIe bus 111, then to NIC 122 via network 130, and then to PPU_m via PCIe bus 112.

[0005] System 100 can be used in applications such as, but not limited to, graph analytics and graph neural networks. More specifically, System 100 can be used in applications such as, but not limited to, online shopping engines, social networks, recommendation engines, mapping engines, fault analysis, network management, and search engines. These applications perform a large number of memory access requests (e.g., read and write requests), and therefore also transfer (e.g., read and write) large amounts of data for processing. Although PCIe has relatively high bandwidth and data transfer rates, it still limits these applications. For these applications, PCIe is generally considered too slow and its bandwidth too narrow. Summary of the Invention

[0006] This disclosure provides solutions to the aforementioned problems. It improves the functionality of general-purpose computing systems and applications such as, but not limited to, neural networks and artificial intelligence (AI) workloads. More specifically, the methods, systems, and programming models provided in this disclosure improve the execution speed of applications such as neural networks and AI workloads by increasing the speed at which memory access requests (e.g., read and write requests) are sent and received between system components and the resulting data transfers are completed. The systems, methods, and programming models of this disclosure allow processing units in a system to communicate without using traditional networks (e.g., Ethernet) via relatively narrow and slow Peripheral Component Interconnect Express (PCIe) buses.

[0007] In some embodiments, the system includes a high-bandwidth inter-chip network (ICN) that allows communication between parallel processing units within the system. For example, an ICN allows a PPU to communicate with other PPUs on the same compute node or server, as well as with PPUs on other compute nodes or servers. In some embodiments, communication can be at the command level (e.g., at the direct memory access level) and the instruction level (e.g., at the finer-grained load / store instruction level). The ICN allows PPUs in the system to communicate without using the PCIe bus, thus avoiding its bandwidth limitations and relative speed limitations. The corresponding routing table includes information on multiple paths to any other given PPU.

[0008] According to a method embodiment, a setup operation including creating a static pre-defined routing table can be performed. Communication data packets can be forwarded from a source parallel processing unit, wherein the communication data packets are formed and forwarded according to the static pre-defined routing table. Communication data packets can be received at a destination parallel processing unit. The respective parallel processing units among the multiple parallel processing units in the network include a source PPU and a destination PPU. A first processing core group of multiple processing cores is included in a first chip, a second processing core group of multiple processing cores is included in a second chip, and the multiple processing units communicate in parallel through multiple interconnects, configuring the corresponding communication according to the static pre-defined routing table.

[0009] According to another embodiment of this disclosure, the system includes a first parallel processing unit group in a first computing node. Corresponding parallel processing units in the first parallel processing unit group are included in separate corresponding chips. The system also includes a second parallel processing unit group in a second computing node. Corresponding parallel processing units in the second parallel processing unit group are included in separate corresponding chips. The system further includes multiple interconnects in a chip interconnect network (ICN) configured to communicatively couple the first and second parallel processing unit groups. Parallel processing units in the first and second parallel processing unit groups communicate through the multiple interconnects, and the corresponding communication is configured according to routing tables residing in the storage characteristics of the respective parallel processing units in the first and second parallel processing unit groups. The multiple interconnects are configured to couple parallel communication from the first parallel processing unit in the first computing node to the second parallel processing unit in the second computing node via at least two paths.

[0010] The above scheme enables communication between neural network processing units in the system without using the PCIe bus, thus avoiding bandwidth limitations and relative speed deficiencies, and improving the data transmission rate and bandwidth of inter-chip communication. Parallel processing units communicate via the chip interconnect network according to corresponding routing tables, which include information on multiple paths to any other given parallel processing unit. The source parallel processing unit can forward at least two of the multiple communication data packets in parallel and / or overlapping, further improving the data transmission rate and bandwidth of inter-chip communication.

[0011] Those skilled in the art will recognize these and other objects and advantages of the various embodiments of this disclosure upon reading the following detailed description of the embodiments shown in the accompanying drawings. Attached Figure Description

[0012] The accompanying drawings, which are included in and form a part of this specification, illustrate embodiments of the present disclosure and, together with the detailed description, serve to explain the principles of the present disclosure, wherein similar numbers describe similar elements.

[0013] Figure 1 An example conventional system is shown.

[0014] Figure 2A A block diagram of an example system according to an embodiment of the present disclosure is shown.

[0015] Figure 2B A block diagram of an example ICN topology according to an embodiment of this disclosure is shown.

[0016] Figure 2C A block diagram of an example parallel processing unit according to an embodiment of the present disclosure is shown.

[0017] Figure 3 A block diagram of an example unified memory addressing space according to an embodiment of the present disclosure is shown.

[0018] Figure 4 A block diagram of an example scaling hierarchy according to an embodiment of the present disclosure is shown.

[0019] Figure 5A , Figure 5B , Figure 5C and Figure 5D A block diagram of an example portion of a communication network according to an embodiment of this disclosure is shown.

[0020] Figure 6 A block diagram of an example portion of a communication network according to an embodiment of this disclosure is shown.

[0021] Figure 7 A block diagram of an example portion of a communication network according to an embodiment of this disclosure is shown.

[0022] Figure 8 Examples of different workload balancing based on different numbers of physical address bits interleaved are shown according to embodiments of this disclosure.

[0023] Figure 9 A flowchart of an example communication method according to an embodiment of this disclosure is shown.

[0024] Figure 10 A flowchart of an example parallel communication method according to an embodiment of this disclosure is shown. Specific Implementation

[0025] Reference will now be made in detail to various embodiments of this disclosure, examples of which are illustrated in the accompanying drawings. Although described in conjunction with these embodiments, it will be understood that they are not intended to limit the disclosure to these embodiments. Rather, the disclosure is intended to cover alternatives, modifications, and equivalents that may be included within the spirit and scope of the disclosure as defined in the appended claims. Furthermore, in the following detailed description of this disclosure, numerous specific details are set forth in order to provide a thorough understanding of the disclosure. However, it will be understood that the disclosure may be practiced without these specific details. In other instances, well-known methods, processes, components, and circuits have not been described in detail in order to avoid unnecessarily obscuring aspects of the disclosure.

[0026] Certain portions of the following detailed description are presented in the form of procedures, logic blocks, processes, and other symbolic representations of operations on data bits in computer memory. These descriptions and representations are means used by those skilled in the art of data processing to most effectively convey the substance of their work to others skilled in the art. In this disclosure, procedures, logic blocks, processes, etc., are considered as a self-consistent sequence of steps or instructions that lead to a desired result. These steps are those that utilize physical operations on physical quantities. Typically, although not essential, these quantities take the form of electrical or magnetic signals that can be stored, transmitted, combined, compared, and otherwise manipulated in a computing system. It has proven convenient, primarily for common usage reasons, to sometimes refer to these signals as transactions, bits, values, elements, symbols, characters, samples, pixels, etc.

[0027] However, it should be remembered that all these and similar terms are associated with appropriate physical quantities and are merely convenient labels applied to those quantities. Unless explicitly stated in the discussion below, it will be understood that throughout this disclosure, discussions using terms such as “access,” “allocate,” “store,” “receive,” “send,” “write,” “read,” “emit,” “load,” “push,” “pull,” “process,” “cache,” “routing,” “determine,” “select,” “request,” “synchronize,” “copy,” “map,” “update,” “convert,” “generate,” “allocate,” etc., refer to a device or computing system (e.g., Figure 7 , 8 Methods 9 and 10) or similar electronic computing devices, systems, or networks (e.g., Figure 2A The actions and processes of a system and its components and elements. A computing system or similar electronic computing device manipulates and transforms data represented as physical (electronic) quantities in memory, registers or other such information storage, transmission or display devices.

[0028] Some of the elements or embodiments described herein can be discussed in the general context of computer-executable instructions residing on some form of computer-readable storage medium (e.g., program modules) and executed by one or more computers or other devices. By way of example and not limitation, computer-readable storage media may include non-transitory computer storage media and communication media. Typically, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. In various embodiments, the functionality of program modules may be combined or allocated as needed.

[0029] The term "non-transitory computer-readable medium" should be interpreted as excluding only those types of transient computer-readable media that, according to 35 USC §101in In re Nuijten, 500F.3d 1346, 1356-57 (Fed. Cir. 2007), are found to be outside the scope of patentable subject matter. The use of this term should be understood as removing the propagation of a transient signal itself from the scope of the claims, without waiving the rights to all standard computer-readable media that not only propagate transient signals themselves.

[0030] Computer storage media include volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer storage media include, but are not limited to, processor registers, double data rate (DDR) memory, random access memory (RAM), static RAM (SRAM) or dynamic RAM (DRAM), read-only memory (ROM), electrically erasable programmable ROM (EEPROM), flash memory (e.g., SSD) or other memory technologies, CD-ROM miniature optical disc ROM, digital versatile optical disc or other optical storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed to obtain that information.

[0031] Communication media may contain computer-executable instructions, data structures, and / or program modules, as well as any information transmission medium. By way of example and not limitation, communication media includes wired media such as wired networks or direct wired connections, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media. Any combination of the foregoing may also be included within the scope of computer-readable media.

[0032] The systems and methods disclosed herein are configured to efficiently and effectively implement parallel data flow routing schemes. In some embodiments, the parallel data flow routing schemes are implemented in an interconnected chip network (ICN). In some embodiments, the systems and methods provide balanced workloads, guaranteed dependencies, and guaranteed access ordering (e.g., for consistency, etc.). Balanced workloads can be provided through multiple flows via minimal links and interleaving with flow IDs and source parallel processing unit (PPU) IDs. Guaranteed dependencies can be provided by using physical addresses (PAs) interleaved with flow IDs and hashed at the source parallel processing unit (PPU). Guaranteed access ordering can be provided by using flow IDs along the route and splittable remote-fences. The systems and methods may also support workload balancing. It should be understood that various routing schemes may exist. Routing schemes may include routing tables. Routing schemes can be used to point to routes at source PPUs, relay PPUs, etc. In one exemplary embodiment, a simplified multipath routing scheme may be used.

[0033] Figure 2AA block diagram of an example system 200 according to an embodiment of this disclosure is shown. System 200 can be used for various types of workloads (e.g., general processing, graphics processing, neural network data processing, etc.). In some embodiments, system 200 can be used for neural network and artificial intelligence (AI) workloads. Typically, system 200 can be used for any parallel computing, including massive data parallel processing.

[0034] Typically, system 200 includes multiple computing nodes (e.g., servers), and each computing node or server includes multiple parallel computing units or chips (e.g., PPUs). Figure 2A In the example, system 200 includes compute nodes (e.g., servers, etc.) 201 and compute node 202. Although Figure 2 includes two compute nodes, it is understood that the number of compute nodes may be different.

[0035] exist Figure 2A In some embodiments, compute node 201 includes a main central processing unit (CPU) 205 and is connected to network 240 via a network interface controller or card (NIC) 206. Compute node 201 may include elements and components other than those described herein. Parallel computing units of compute node 201 include network processing units (PPUs) PPU_0a to PPU_na, which are connected to a peripheral component interconnect high-speed (PCIe) bus 208, which in turn is connected to the NIC 206.

[0036] In some embodiments, computing node 202 includes elements similar to those of computing node 201 (although 'm' may or may not be equal to 'n'). Other computing nodes in system 200 may be constructed similarly. In one exemplary embodiment, computing node 201 and computing node 202 may have the same structure, at least to the extent described herein.

[0037] The PPUs on compute node 201 can communicate with each other via bus 208 (communication coupling). The PPUs on compute node 201 can communicate with the PPUs on compute node 202 via network 240 through buses 208 and 209 and NICs 206 and 207.

[0038] Figure 2ASystem 200 includes a high-bandwidth chip interconnect network (ICN) 250, which allows communication between PPUs in system 200. That is, the PPUs in system 200 are communicatively coupled to each other via ICN 250. For example, ICN 250 allows PPU_0a to communicate with other PPUs on compute node 201, and also allows PPU_0a to communicate with PPUs on other compute nodes (e.g., compute node 202). Figure 2A In the example, ICN 250 includes interconnects (e.g., interconnects 252 and 254) that directly connect two PPUs and allow bidirectional communication between the two connected PPUs. The interconnect can be a half-duplex link, on which only one PPU can transmit data at a time, or the interconnect can be a full-duplex link, on which data can be transmitted simultaneously in both directions. In one embodiment, the interconnects (e.g., interconnects 252 and 254) are lines or cables based on or utilizing serializer / deserializer (SerDes) functionality.

[0039] exist Figure 2A In one example embodiment, interconnect 252 is a hardwired or cable connection that directly connects PPU_0a on compute node 201 to PPU_na, and interconnect 254 is a hardwired or cable connection that directly connects PPU_na on compute node 201 to PPU_0b on compute node 202. That is, for example, one end of interconnect 252 is connected to PPU_0a on compute node 201, and the other end is connected to PPU_na. More specifically, one end of interconnect 252 is inserted into switch 234 (…). Figure 2C The other end of the interconnect 252 is inserted into a port on the PPU_na switch or coupled to a port on the PPU_na switch.

[0040] Understandably, the actual connection topology (which PPU each PPU connects to outside itself) can be different. Figure 2BThis is a block diagram illustrating an example ICN topology. In some embodiments, each compute node has three PPUs. It is understood that the number of PPUs per compute node can be different. In one exemplary embodiment, PPU_0c on compute node 291 is connected to PPU_xc on compute node 291, and PPU_xc on compute node 291 is then connected to PPU_yc on compute node 291. PPU_0c on compute node 291 can be connected to a PPU on another compute node (not shown). PPU_yc on compute node 291 is connected to PPU_0d on compute node 292. In one embodiment, a PPU can be connected to its immediate neighbor on a compute node or its immediate neighbor on an adjacent compute node. It should be understood that the connections of PPUs can be different. Therefore, in Figure 2B In the example, PPU_0d on compute node 292 is connected to PPU_yd on compute node 292, and PPU_yd on compute node 292 is in turn connected to PPU_xd on compute node 292. PPU_xd on compute node 292 can be connected to a PPU on another compute node (not shown). Interconnections connecting PPUs on the same compute node are called intra-chip interconnects, and interconnections connecting PPUs on different compute nodes are called inter-chip interconnects.

[0041] Communication between any two PPUs (e.g., the two PPUs may be on the same compute node or on different compute nodes, etc.) can be direct or indirect. In some embodiments, direct communication occurs via a single link between the two PPUs, and indirect communication occurs when information from one PPU is relayed to another PPU via one or more intermediate PPUs. For example, in Figure 2A In the illustrated configuration, PPU_0a on compute node 201 can communicate directly with PPU_na on compute node 201 via interconnect 252, and PPU_0a on compute node 201 can communicate indirectly with PPU_0b on compute node 202 via interconnect 252 to PPU_na and interconnect 254 on compute node 202 from PPU_na to PPU_0b. Communication between PPUs may include the transmission of memory access requests (e.g., read requests and write requests) and data transfers in response to these requests.

[0042] Communication between PPUs can be at the command level (e.g., DMA copy) and the instruction level (e.g., direct load or store). The ICN250 allows compute nodes and PPUs in system 200 to communicate without using the PCIe bus 208, thus avoiding its bandwidth limitations and relative speed inadequacies.

[0043] A PPU can also be implemented using a neural PPU, or it can be called a neural PPU. A PPU can also be implemented as or use various different processing units, including general-purpose processing units, graphics processing units, neural network data processing units, etc.

[0044] In some embodiments, the PPUs on compute node 201 can communicate with each other via bus 208 in addition to communicating via ICN 250 (communication coupling). Similarly, the PPUs on compute node 201 can also communicate with the PPUs on compute node 202 via network 240 via buses 208 and 209 and NICs 206 and 207.

[0045] System 200 and PPU (e.g., PPU_0a, etc.) may include elements or components other than those shown and described below, and these elements or components may be arranged as shown or in different ways. Some blocks in the example system 200 and PPU can be described according to the functions they perform. The elements and components of the system are described and shown as separate blocks, but this disclosure is not limited thereto; that is, for example, combinations of blocks / functions may be integrated into a single block that performs multiple functions. System 200 may be extended to include additional PPUs, and system 200 is compatible with different scaling schemes, including hierarchical scaling schemes and flat scaling schemes.

[0046] Typically, each PPU on compute node 201 includes elements such as processing cores and memory. Figure 2C This is a block diagram illustrating an example parallel processing unit PPU_0a according to an embodiment of the present disclosure. Figure 2C The PPU_0a shown includes a network-on-a-chip (NoC) 210 coupled to one or more computing elements or processing cores (e.g., core 212a, core 212b, core 212c, core 212d, etc.) and one or more caches (e.g., cache 214a, cache 214b, cache 214c, cache 214d, etc.). The PPU_0a also includes, for example, one or more high-bandwidth memories (HBMs) (e.g., HBM 216a, HBM 216b, HBM 216c, HBM 216d, etc.) coupled to the NoC 210. Figure 2CThe processing core, cache, and HBM can also be collectively referred to herein as core 212, cache 214, and HBM 216, respectively. In one exemplary embodiment, cache 214 is the last level of cache between HBM 216 and NoC 210. Compute node 201 may include other levels of cache (e.g., L1, L2, etc.; not shown). Memory space in HBM 216 may be declared or allocated (e.g., at runtime) as buffers (e.g., ping-pong buffers). Figure 2C (Not shown in the image).

[0047] PPU_0a may also include other functional blocks or components (not shown), such as an command processor, a direct memory access (DMA) block, and a PCIe block facilitating communication with the PCIe bus 208. PPU_0a may include the components described herein or... Figure 2C Components and components other than those shown.

[0048] In some embodiments, system 200 includes a unified memory addressing space, for example, using a Partitioned Global Address Space (PGAS) programming model. Therefore, memory space in system 200 can be globally allocated such that, for example, the HBM216 on PPU_0a can be accessed by other PPUs on that compute node and by PPUs on other compute nodes in system 200. Similarly, PPU_0a can access HBMs on other PPUs / compute nodes in the system. Therefore, in Figure 2A and Figure 2B In the example, a PPU can read data from or write data to another PPU in system 200, wherein the two PPUs can be on the same compute node or different compute nodes, and the read or write can occur directly or indirectly.

[0049] Computing node 201 communicates with ICN subsystem 230 (e.g. Figure 2CThe ICN subsystem 230 is coupled to ICN 250, and ICN subsystem 230 is coupled to NoC 210. In some embodiments, ICN subsystem 230 includes ICN communication control block (communication controller) 232, switch 234, and one or more inter-communication links (ICLs) (e.g., ICL 236; collectively referred to as ICL 236). ICL 236 may be coupled to switch 234, or may be a component of switch 234. ICL 236 constitutes or includes ports. In one example embodiment, there are 7 ICLs. Each ICL 236 is connected to a corresponding interconnect (e.g., interconnect 252). One end of interconnect 252 is coupled or inserted into ICL (port) 236 on PPU_0a, and the other end of interconnect is coupled or inserted into another ICL / port on another PPU.

[0050] exist Figure 2C In one configuration, for example, a memory access request (e.g., a read request or a write request) for PPU_0a is sent from NoC 210 to ICN communication control block 232. The memory access request includes an address identifying which compute node / PPU / HBM is the target of the memory access request. ICN communication control block 232 uses this address to determine which ICN 236 (directly or indirectly) is communication-coupled to the compute node / PPU / HBM identified by that address. The memory access request is then routed by switch 234 to the selected ICL 236, and then by ICN 250 to the compute node / PPU / HBM identified by that address. At the receiving end, the memory access request is received at the ICL of the target PPU, provided to the ICN communication control block and then to the NoC of that PPU, and finally to the HBM on that PPU addressed by the memory access request. If the memory access request is a write request, the data is stored at the address in the HBM on the target PPU. If the memory access request is a read request, the data at the address in the HBM on the target PPU is returned to PPU_0a. In this way, the high-bandwidth ICN250 is used to quickly complete inter-chip communication, bypassing the PCIe bus 208, thus avoiding its bandwidth limitations and relative speed inadequacy.

[0051] The PPU may include a computation command ring coupled between core 212 and ICN subsystem 230. The computation command ring may be implemented as multiple buffers. A one-to-one correspondence may exist between the core and the computation command ring. Commands from processes executing on the core are pushed into the headers of the corresponding computation command rings in the order they are issued or to be executed.

[0052] ICN subsystem 230 may also include multiple chip-to-chip (C2C) DMA units coupled to the command scheduling block and instruction scheduling block. The DMA units are also coupled to NoC 210 via a C2C fabric and a network interface unit (NIU), and are further coupled to switch 234, which in turn is coupled to ICL 236 coupled to ICN 250.

[0053] In one embodiment, there are 16 communication command rings and 7 DMA units. A one-to-one correspondence may exist between DMA units and ICLs. The command scheduling block maps the communication command rings to DMA units, and thus to ICL 236. The command scheduling block, instruction scheduling block, and DMA units may each include buffers such as first-in-first-out (FIFO) buffers (not shown).

[0054] The ICN communication control block 232 maps the output memory access request to the ICL 236 selected based on the address in the request. The ICN communication control block 232 forwards the memory access request to the DMA unit corresponding to the selected ICL 236. The switch 234 then routes the request from the DMA unit to the selected ICL.

[0055] System 200 and PPU are for implementing methods such as those disclosed herein (e.g., Figure 9 Examples of systems and processing units (such as those described in the text).

[0056] Figure 3 A block diagram of an example unified memory addressing space according to an embodiment of this disclosure is shown. The unified memory addressing space enables a programming model of the Partitioned Global Address Space (PGAS) type. Communication between programs can flow at different levels. At the command level, communication may include direct memory access (DMA) copy operations. At the instruction level, communication may include direct load / store operations. In some embodiments, a variable VAR, considered a portion of shared space, may be stored locally in a physical memory. In an exemplary embodiment, the value of VAR may be written to a first process on a first PPU, and the value of VAR may be read by a second process on a second PPU.

[0057] Figure 4A block diagram of an example system 400 according to an embodiment of the present disclosure is shown. System 400 is an example scaling hierarchy or topology method according to one embodiment of the present disclosure. In an exemplary embodiment, multiple PPUs are communicatively coupled to each other via ICN connections (e.g., 401, 402, 403, 404, 405, 409, 411, 412, 421, 422, 423, 441, 442, 443, etc.).

[0058] In some embodiments, multiple chips are communicatively networked together in some topology, and a routing scheme enables the system to route communication to the appropriate target chip for each request. In one exemplary embodiment, the chip includes a communication processing component. The communication processing component can be considered as a parallel processing unit (PPU). In some embodiments, routing tables and some routing rules instruct the hardware how to control the data flow. In one exemplary embodiment, the routing scheme can be implemented in a neural network configuration.

[0059] It is understood that communication networks can have various configurations. Generally, information / data communication flows from a source to a destination. In one embodiment, a communication routing path can be identified by indications of the source and destination associated with the communication. The network may include more destination components than the egress ports included in the PPU, in which case each PPU is not directly coupled to every other PPU in the network. In one embodiment, the network includes intermediate PPUs in addition to the source and destination PPUs to help route data between the source and the destination. In an exemplary embodiment, an intermediate PPU is considered a relay. The network can be configured with intermediate or relay components between the source and the destination (e.g., in a “mesh” topology, etc.). Generally, the source is the starting / first communication PPU in the path, the relay is an intermediate PPU (e.g., which receives communication from an upstream / previous PPU and forwards that communication to a downstream / next PPU in the path, etc.), and the destination is the final / last communication PPU in the path. The network may include multiple paths or routes between the source and the destination.

[0060] Data can be communicated or flow along a communication "path" in a network formed by communication "links" between communication components (e.g., PPUs, etc.). Data can be segmented / separated into data packets. Different data packets can flow along the same link or different links. Additional descriptions of communication paths and links are presented in other parts of this specification.

[0061] When communicating over a network, several key considerations typically need to be addressed. One consideration is how to handle communication when multiple network paths exist from the source to the destination. Achieving a proper balance between resource utilization and workload distribution across a multi-path network can be difficult, as consistently utilizing only a single link essentially idles other available resources, while attempting to always fully utilize all links can introduce unrealistic coordination complexities. In many applications, accuracy is often critical, and the timing of memory access operations is often essential to maintaining accuracy.

[0062] In complex communication networks, various problems and conditions may exist that could potentially affect the timing of memory access operations (e.g., reads, writes, etc.). In one exemplary embodiment, there may be multiple write operations associated with transmitting information (e.g., different writes may involve / transmit different parts of data, etc.), and different parts of the write operations may potentially communicate via different network paths. One path may transmit information faster than another. If proper write operation timing cannot be guaranteed (e.g., including considerations of communication duration, etc.), a second (e.g., "later," subsequent, etc.) write request may overtake a first (e.g., earlier), "earlier," etc., write request via another, faster communication path, and the second write operation may occur earlier than the first write operation in a manner that violates proper operation timing. Typically, timing or write operations can be associated with memory / data consistency. An application may require a first data segment and a second data segment to be written to the same address / location or different addresses / locations. In one embodiment, if the two different data segments are written to the same address, timing is considered to be related to data dependency, and if the two different data segments are written to different addresses, timing is considered to be related to access order.

[0063] Figure 5AA block diagram of an example portion of a communication network according to an embodiment of the present disclosure is shown. Communication network portion 510 includes a source PPU 501, a relay PPU 502, and a target PPU 503. The source PPU 501 is communicatively coupled to the relay PPU 502 via communication link 511, while the relay PPU 502 is communicatively coupled to the target PPU 503 via communication links 512 and 513. In one embodiment, bandwidth is constrained by the narrowest path / link between the source and the target. For example, although there are two links / paths (i.e., communication link 512 and communication link 513) in the network segment between PPU 502 and PPU 503, there is only one link / path 511 in the network segment between PPU 501 and PPU 502. In an exemplary embodiment, even though links 512 and 513 exist (for a total of two path widths between the second PPU 502 and the third PPU 503, etc.), the throughput is determined by link 511. In one exemplary embodiment, a link (e.g., link 511, etc.) is a limitation on the total throughput, and thus it is the basis of the minimum link concept. From the perspective of the chip including PPU 501, the route to the target chip including PPU 503 has a minimum link (Minlink#) 1.

[0064] Figure 5B A block diagram of an example portion of a communication network according to an embodiment of this disclosure is shown. Communication network portion 520 includes a source PPU 521, a relay PPU 522, and a target PPU 523. The source PPU 521 is communicatively coupled to the relay PPU 522 via communication links 531 and 532, and the relay PPU 522 is communicatively coupled to the target PPU 523 via communication link 533. Communication network portion 520 has a different configuration than communication network portion 510, but still has a minimum link 1. Similar to communication portion 510, in communication portion 520, bandwidth is constrained by the narrowest path / link between the source and the target. For example, although there are two links / paths (i.e., communication link 531 and communication link 532) in the network segment between PPU 521 and PPU 522, there is only one link / path 533 in the network segment between PPU 522 and PPU 523. In one exemplary embodiment, even though links 531 and 532 exist (for a total of two path widths between the first PPU 521 and the second PPU 522, etc.), the throughput is determined by link 533. In one exemplary embodiment, again, a link (e.g., link 533, etc.) is a constraint on the total throughput, and therefore it forms the basis of the minimum link concept. From the perspective of the chip including PPU 521, the route to the target chip including PPU 523 has a minimum link (Minlink#) 1.

[0065] In a multi-stream system, the relay chip can operate in a way that balances the flow between exits. The minimum link from PPU 501 to PPU 503 is 1, therefore, in the network segment between PPU 502 and PPU 503, one of link 512 or link 513 is used to pass information from PPU 501 to PPU 503. It is understood that if one of the links, such as link 512, is used for communication from PPU 501 to PPU 503, another link, such as link 513, can be used for other communications. In an exemplary embodiment, another PPU (not shown) similar to PPU 501 is coupled to PPU 502, and communication from this other PPU is forwarded from PPU 502 to PPU 503 via link 513. When there are multiple links between two PPUs, it is generally desirable to balance the data flow on multiple links. In this case, the systems and methods of this disclosure balance the PPU workload in the communication channel / path.

[0066] Figure 5C A block diagram of an example portion of a communication network according to an embodiment of the present disclosure is shown. The communication network portion 540 includes a source PPU 541, a relay PPU 542, and a target PPU 543. The source PPU 541 is communicatively coupled to the relay PPU 542 via communication links 551 and 552, while the relay PPU 542 is communicatively coupled to the target PPU 543 via communication links 553 and 554. Link 551 transmits data 571 from PPU 541 to PPU 542, and link 552 transmits data 572 from PPU 541 to PPU 542. Link 553 transmits both data 571 and data 572 from PPU 542 to PPU 543. In one exemplary embodiment, the workload communication network portion 540 is not optimally balanced by the relay PPU 542. The communication network section 540 has two links / paths (i.e., communication link 551 and communication link 552) in the network segment between PPU 541 and PPU 542, and two links / paths (i.e., communication link 553 and communication link 554) in the network segment between PPU 542 and PPU 543. Therefore, the minimum number of links (Minlink#) is 2. The source PPU 541 balances the data traffic by sending data 571 on link 551 and data 572 on link 552. Unfortunately, the data traffic from PPU 542 to PPU 543 on links 553 and 554 is unbalanced (e.g., relay PPU 542 sends both data 571 and data 572 on link 553, leaving link 554 empty, etc.).

[0067] Figure 5D A block diagram of an example portion of a communication network according to an embodiment of the present disclosure is shown. The communication network portion 580 includes a source PPU 581, a relay PPU 582, and a target PPU 583. The source PPU 581 is communicatively coupled to the relay PPU 582 via communication links 591 and 592, and the relay PPU 582 is communicatively coupled to the target PPU 583 via communication links 593 and 594. Link 591 transmits data 577 from PPU 581 to PPU 582, and link 592 transmits data 578 from PPU 581 to PPU 582. Link 593 transmits data 578 from PPU 582 to PPU 583, and link 594 transmits data 577 from PPU 582 to PPU 583. In one exemplary embodiment, Figure 5D The workload communication network part of the 580 is compared to Figure 5C The communication network section 540 is further optimized for load balancing. Communication network section 580 has two links / paths (i.e., communication link 591 and communication link 592) in the network segment between PPU 581 and PPU 582, and two links / paths (i.e., communication link 593 and communication link 594) in the network segment between PPU 582 and PPU 583. Therefore, the minimum number of links (Minlink#) is 2. Source PPU 581 balances data transmission volume by transmitting data 577 on link 591 and data 578 on link 592. Relay PPU 582 also balances data transmission volume by transmitting data 578 on link 593 and data 577 on link 594.

[0068] In one embodiment, the basic routing scheme is simple and implemented in hardware. In one exemplary embodiment, statically pre-defined routes are used and set before runtime. In one exemplary embodiment, once the routing table is set, it cannot be changed by the hardware at runtime. Routing information can be programmed / reprogrammed into hardware registers before runtime. The hardware can reconfigure routes the next time routing is reset, but once routes are set, they cannot be changed at runtime. In one embodiment, the system uses a basic XY routing scheme plus some simple rules to develop the routing table.

[0069] In one embodiment, the PPU includes a routing table. In one embodiment, the routing table may have the following configuration:

[0070] Table 1

[0071]

[0072] In one embodiment, each chip / PPU has 7 ports, and bits are set to indicate whether a port is available / should be used for routing packets. In one embodiment, a bit is set to logic 1 to indicate that an egress port is available, and a bit is set to logic 0 to indicate that an egress port is unavailable. In an exemplary embodiment, the system can support up to 1024 chips / PPUs, and the routing table can have a corresponding number of entries from each source PPU. This means that up to 1023 target chips / PPUs can be reached from a source PPU. Each entry corresponds to communication from an associated source PPU to a specific target PPU. For the example table above, the routing table indicates which of the 7 egress ports is available / should be used to communicate with PPU 4 and PPU 5. The routing table may include a field indicating the minimum link for communication associated with the indicated target PPU. In one embodiment, although multiple paths with multiple hops may exist, the minimum number of links represents the narrowest path. To reach chip / PPU number 4, there is a minimum link 2. In this scenario, another hop in the entire route / path between the source and destination can have 3 links, but the hop from the PPU associated with the search table to PPU 4 has 2 links.

[0073] Figure 6 A block diagram of an example portion of a communication network 600 according to an embodiment of the present disclosure is shown. The communication network 600 includes PPUs 610 to 625 communicatively coupled via links 631 to 693.

[0074] In one exemplary embodiment, data is sent from PPU 610 to PPU 615. In this example, PPU 610 is the source PPU, and PPU 615 is the destination PPU. PPUs 611 and 614 can act as relay PPUs. Links 631 and 632 can be used to forward data from PPU 610 to PPU 611, and link 671 can be used to forward data from PPU 611 to PPU 615. Link 661 can be used to forward data from PPU 610 to PPU 614, and links 633 and 634 can be used to forward data from PPU 614 to PPU 615. While links 631, 632, 633, 634, 661, and 671 are generally used to pass information between some PPUs, some links can be turned off or disabled. Additional instructions on turning off or disabling PPUs are presented in other parts of this detailed specification.

[0075] In a scenario where PPU 610 is the source PPU and PPU 615 is the target PPU, when examining the first path via relay PPU 614 and the second path via relay PPU 611, the minimum number of links is determined to be 1. For example, for the first path via relay PPU 614, even though the path between PPU 614 and PPU 615 may have two links (e.g., links 633 and 634), the path between PPU 610 and PPU 614 has only one link (e.g., link 661). Therefore, the minimum number of links is 1 (e.g., corresponding to link 661, etc.). Similarly, the second path via relay PPU 611 has a minimum of 1 link (e.g., corresponding to link 671). In some embodiments, the driver uses the minimum link information and disables link 631 or link 632.

[0076] In one exemplary embodiment, PPU 610 is not concerned with which relay PPU (e.g., PPU 611, PPU 614, etc.) is used; it is only concerned with instructing the target PPU 615 and from which egress port the communication will be forwarded. In one exemplary embodiment, information is sent to relay PPU 611, and the first hop in the path from source PPU 610 to target PPU 615 is through either egress port 1 or egress port 2 of PPU 610. If egress port 1 of PPU 610 is used, then egress port 2 can be disabled / closed, and vice versa. Since PPU 614 is not used, links 633 and 634 can remain enabled / open to handle other communications (e.g., servicing communication from PPU 618 to PPU 615, etc.). In one exemplary embodiment, while PPU 610 may not necessarily be concerned with which relay PPU (e.g., PPU 611, PPU 614, etc.) is used, the overall system may be concerned with which relay PPU is used, and this may be important. Additional notes on considerations for path / link / PPU selection (e.g., impact on workload balancing, data consistency, etc.) are presented in other parts of this specification.

[0077] In one exemplary embodiment, data is sent from PPU 610 to PPU 620. In this example, PPU 610 acts as the source PPU, and PPU 620 acts as the destination PPU. PPUs 611, 612, 614, 615, 616, 618, and 619 can act as relay PPUs. Although there are more potential paths from PPU 610 to PPU 620 than from PPU 610 to PPU 615, the routing decisions of concern to PPU 610 are the same because in both scenarios PPU 610 forwards information to either PPU 611 or PPU 614, and PPU 610 has no control over further downstream routing decisions on the path from PPU 610 to PPU 620. In one exemplary embodiment, the only flow that a PPU is concerned with is the next hop leaving that PPU, not other hops in the path or flow.

[0078] In one embodiment, a flow ID (Identifier) ​​is used in the routing decision. The flow ID can correspond to a unique path between the source and destination locations. As described above, the data flow from PPU 610 to the destination PPU 615 can be split into two data flows (e.g., via PPU 611 and PPU 614). In an exemplary embodiment, the flow ID can be created by hashing bits selected from the physical address of the destination PPU. Furthermore, the flow ID function can include a minimum link indicator. In an exemplary embodiment, the flow ID determination operation can be represented as:

[0079] Flow_id=gen_flow_id(hashing(selected_bits(&PA)),#MinLink)

[0080] Similarly, the indication of the exit port selection can be expressed as:

[0081] ePort_id=mapping(Flow_id%#MinLink)

[0082] As previously mentioned, there can be multiple paths for a source PPU and destination PPU pair (e.g., srcPUU-dstPPU, etc.). In some embodiments, the source PPU uses only the minimum number of links (e.g., #MinLink, etc.). In an exemplary embodiment, within each program flow (e.g., process, stream, etc.), data flowing from the local PPU (or source PPU) to the destination PPU is evenly distributed along multiple paths. In some embodiments, paths are split based on physical addresses (PAs). Hash functions can be applied to better balance accesses with strides.

[0083] Figure 7 A block diagram of an example portion of a communication network 700 according to an embodiment of the present disclosure is shown. The communication network 700 includes PPUs 701, 702, 703, 704, 705, and 706. PPU 701 is communicatively coupled to PPU 705 via links 711 and 712. PPU 702 is communicatively coupled to PPU 704 via links 713 and 714. PPU 703 is communicatively coupled to PPU 704 via links 717 and 719. PPU 704 is communicatively coupled to PPU 705 via path 721. PPU 705 is communicatively connected to PPU 706 via links 731, 732, and 733. Links 711, 712, 713, 714, 717, and 719 transmit data 751, 752, 753, 754, 757, and 759, respectively. Link 731 transmits data 751 and 759. Link 732 transmits data 753 and 752. Link 733 transmits data 757 and 754.

[0084] In some embodiments, relay PPU 705 receives RR ingress data from PPU 701, PPU 702, PPU 703, and PPU 704 and forwards the data to PPU 706. Regarding forwarding information to PPU 706, PPU 705 looks up entries for the target PPU 706 to determine which egress ports are available. In some embodiments, the routing table includes information about the connection topology between chips or PPUs. Table 2 is a block diagram of an example portion of the routing table used by PPU 705 for communication to the target PPU 706.

[0085] Table 2

[0086]

[0087] The routing table indicates that egress ports 2, 3, and 5 are available to forward information to PPU 706. After determining the available ports, PPU 705 then uses "smart" features to determine which egress port to send data to the target PPU 706.

[0088] In some embodiments, a unique flow ID exists, including a target information ID (e.g., PPU 706, etc.). The flow ID is used to determine on which available egress port to forward communication packets. In one exemplary embodiment, the flow ID is locally unique. In one exemplary embodiment, a globally unique ID is established by adding an indication of the source PPU to the flow ID. In some embodiments, a first source uses a first link and a third link, a second source uses a second link and a third link, and a third source uses a third link and a first link (e.g., see...). Figure 7 Links 731, 732, and 733 transmit information from the first source 701, the second source 702, and the third source 703, etc. In one exemplary embodiment, the six traffic flows are generally balanced.

[0089] Table 3 is not a routing table, but merely an exemplary illustration of the association between a data flow and an exit port after the relay routing algorithm selects an exit port for a particular data flow.

[0090] Table 3

[0091]

[0092] Refer again Figure 7 Data 751 and 759 are forwarded on link 731 via egress port 2, where data 751 is associated with flow ID "flow-A" and data 759 is associated with flow ID "flow-F". Data 752 and 753 are forwarded on link 732 via egress port 3, where data 752 is associated with flow ID "flow-B" and data 753 is associated with flow ID "flow-C". Data 754 and 757 are forwarded on link 733 via egress port 5, where data 754 is associated with flow ID "flow-D" and data 757 is associated with flow ID "flow-E". In some embodiments, by adding the source PPU IDs, the system assigns an offset to the flow IDs so that different source flows will start from different links. However, the workload may not be well balanced.

[0093] In one embodiment, using more PA bits for interleaving can improve workload distribution and balancing. Figure 8Examples of different workload balancing based on different numbers of PA bits interleaving according to one embodiment are shown. Tables 810 and 820 show different workload distributions between 2-bit interleaving and 3-bit interleaving on three communication links (e.g., A, B, C, etc.). In some embodiments, the interleaved bits are the least significant bits of the address. The bottom table displays the same information in a different tabular format, providing a more intuitive indication of the better workload balancing of 3-bit interleaving compared to 2-bit interleaving. The interleaving pattern repetition in Table 810 is indicated by reference numbers 811A, 812A, and 813A in the top table and corresponding reference numbers 811B, 812B, and 813B in the bottom table. The interleaving pattern repetition in Table 820 is indicated by reference number 821A in the top table and corresponding reference number 821B in the bottom table.

[0094] In some embodiments, the routing scheme utilizes routing tables and destination physical addresses for interleaving. The routing scheme at the source PPU may include allocating flow IDs based on a minimum link count indication for interleaving physical addresses (e.g., PA-inter intoMinLink => flowid, etc.). The routing scheme at the relay PPU may include interleaving flow IDs and source IDs (e.g., using flowid=srcid interleaving, etc.).

[0095] In some embodiments, the PPU includes a routing table shared by two different modules of the PPU's ICN subsystem (e.g., 230, etc.). In one exemplary embodiment, when the PPU acts as a source PPU, the ICN subsystem communication control block (e.g., 232, etc.) uses the routing table to forward messages. In one exemplary embodiment, when the PPU acts as a relay PPU, the ICN subsystem switch (e.g., 234, etc.) uses the routing table to forward messages. In one exemplary embodiment, the ICN subsystem communication control block (e.g., 232, etc.) implements a first algorithm for forwarding messages using the routing table. In one exemplary embodiment, when the PPU acts as a relay PPU, the ICN subsystem switch (e.g., 234, etc.) implements a second algorithm for forwarding messages using the routing table. The first algorithm and the second algorithm can be different. The first algorithm may include:

[0096] ePort_id=mapping(Flow_id%#MinLink)

[0097] And the second algorithm may include:

[0098] ePort=mapping[(src_PPU_ID+flow_ID)%num_possible_ePort]

[0099] In one exemplary embodiment, when the PPU is used as a relay PPU, the indication of the relay's exit port selection can be represented by a second algorithm.

[0100] Figure 9 This is a flowchart of an example communication method 900 according to one embodiment.

[0101] In block 910, a setup operation including creating a routing table is performed. The routing table may include a statically predefined routing table. In one exemplary embodiment, during setup, the driver traverses the topology and collects information included in the routing table. The routing table may include an indication of the minimum number of links in the path to the target PPU. In some embodiments, some links in the routing links are disabled, such that each target PPU entry in the table has a number of available egress ports equal to the minimum number of links in the communication path (e.g., #ePort = #minLink, etc.).

[0102] In block 920, communication data packets are forwarded from the source parallel processing unit (PPU). In some embodiments, communication data packets are formed and forwarded according to a statically predetermined routing table.

[0103] In some embodiments, communication data packets are forwarded based on a routing scheme. In one exemplary embodiment, a routing scheme at the source PPU is determined by the following steps: creating a flow ID associated with a unique communication path via interconnection, wherein the flow ID is established by multiple bits selected in a hashed physical address; determining the minimum link path to the destination using an appropriate routing table; and establishing a route selection based on the flow ID and the minimum link path. In one exemplary embodiment, the routing scheme at the relay PPU includes selecting an egress port, wherein selecting an egress port includes: creating a flow ID associated with a unique communication path via interconnection, wherein the flow ID is established by multiple bits selected in a hashed physical address; mapping the source PPU ID and the flow ID; determining multiple possible egress ports available based on the mapping; determining the minimum link path to the destination using an appropriate routing table; and establishing a route selection based on the flow ID, the multiple possible egress ports, and the minimum link path.

[0104] In block 930, communication data packets are received at the target parallel processing unit (PPU). In some embodiments, the source PPU and the target PPU are included in corresponding processing units among a plurality of processing units in a network, wherein a first processing core group of a plurality of processing cores is included in a first chip, a second processing core group of a plurality of processing cores is included in a second chip, the plurality of processing units communicate through a plurality of interconnects, and the respective communication is configured according to a static predetermined routing table.

[0105] In some embodiments, the example communication method 900 further includes balanced forwarding of communication packets, which includes distributing communication packets via physical address-based interleaving.

[0106] Improvements to the parallel data flow routing described above may include including information about the next PPU in the routing table. The next PPU may be a preferred intermediate PPU, selected for routing to the final target PPU via one or more intermediate PPUs.

[0107] Refer again Figure 6 There are two exemplary paths from PPU 610 to PPU 615. The first path is from PPU 610 to PPU 615 via PPU 614. The second path is from PPU 610 to PPU 611 and then to PPU 615. Typically, there can be multiple paths. Furthermore, there are multiple available links between PPU 610 and PPU 611, such as links 631 and 632, and multiple available links between PPU 614 and PPU 615, such as links 633 and 634. In some embodiments, the PPU includes a routing table that may include additional information about the next PPU. In some embodiments, the routing table may have the following configuration:

[0108] Table 4

[0109]

[0110] It is understandable that the minimum number of links between PPU 610 and PPU 615 is two. For example, there are two paths from PPU 610 to PPU 615, such as PPU 610 to PPU 614 and then to PPU 615, and PPU 610 to PPU 611 and then to PPU 615. However, if links 631 and 632 are both used to send data from PPU 610 to PPU 615, then because link 671 from PPU 611 to PPU 615 is a single link and cannot carry all the data from both links 631 and 632, the entire data flow will be blocked at PPU 611.

[0111] Therefore, the routing algorithm should also have information on how many available links there are from the source to the destination (e.g., from PPU 610 to PPU 615 in this example).

[0112] In some embodiments, each chip / PPU has 7 ports, and bits are set to indicate whether a port is available / should be used to route packets. In some embodiments, bits are set to logic 1 to indicate that an egress port is available, and bits are set to logic 0 to indicate that an egress port is unavailable. In one exemplary embodiment, the system can support up to 1024 chips / PPUs, and the routing table can have multiple entries corresponding to each source PPU. For example, a source PPU can implement up to 1023 target chips / PPUs. Each entry corresponds to communication from the associated source PPU to a specific target PPU. For the example table above, the routing table indicates which of the 7 egress ports is available / which egress port should be used to communicate with PPU 4 and PPU 5. The routing table may include a minimum link indication field associated with communication to the indicated target PPU. In some embodiments, although multiple paths with multiple hops may exist, the minimum number of links represents the narrowest path.

[0113] For example, in the example routing table 4 above, ports 1 and 3 are enabled for target PPU 5, as shown by routing table entry "1". Since more than one port can potentially be used to send information, a "BranchID" entry is added to the routing table.

[0114] In addition, minimum links, such as the entry "minlink", are added to the routing table to identify the minimum link in a source-destination pair. For example, the path from PPU 610 to PPU 615 has a minimum link of 1. Similarly, the path from PPU 610 via PPU 611 to PPU 615 must pass through link 617 with a width of 1.

[0115] In some embodiments of this disclosure, a link (e.g., link 632) may be indicated as unavailable / disabled. For example, since the minimum link required for a transfer from PPU 610 to PPU 615 via PPU 611 is 1, link 632 can be considered redundant and is indicated as disabled. As previously stated, using both links 631 and 632 may cause congestion because link 671 cannot transmit the amount of data provided by the combination of links 621 and 632.

[0116] According to embodiments of this disclosure, links with no path in the routing table are permanently marked as unavailable. For example, links 631 and 632 can both be enabled for data transmission from PPU 610 to PPU 615 via PPU 611. The routing algorithm can dynamically select one of links 631 or 632 for a given data packet on its way to PPU 615. Due to the minimum link length from PPU 610 to PPU 615, only one of links 631 or 632 will be selected. Transmission of a second data packet from PPU 610 to PPU 615 can again select either link 631 or 632.

[0117] For example, a first data packet is sent from PPU 610 to PPU 615 via PPU 611 through link 631. A second data packet is sent from PPU 610 to PPU 615 via PPU 611 through link 632. Compared to embodiments indicating that the links are unavailable / inactive, the first and second data packets can be sent in parallel and / or overlapped, thereby maximizing the available bandwidth of links 631 and 632. Such embodiments allow multiple processes (e.g., running on PPU 610) to fully utilize links 631 and 632, rather than waiting for a single process (e.g., at a rate less than the available bandwidth of both links 631 and 632) to create data packets.

[0118] Below is exemplary pseudocode for implementing a multi-branch routing system:

[0119] / / 0.Generates num_MinLink_Src2Dst flow_ids according to PA

[0120] num_MinLink_Src2Dst=sum(num_MinLink in each src dest branch);

[0121] Flow_id=gen_flow_id(PA_bits),num_MinLink_Src2Dst

[0122] / / 1.for each branch,select#minLink ePorts out of#ePorts_thisBranch;merge into a candidate_ePort list

[0123] / / 2.select an ePort from candicate_ePort_list

[0124] ePort_id_local=SrcPPU_id+flow_id)%size_candicate_ePort_list;; / / AddSrcPPU_id iff not SrcPPU

[0125] ePort_id=the id of the"ePort_id_local th in thisBranch's ePortList

[0126] Figure 10 A flowchart of an exemplary parallel communication method 1000 according to an embodiment of the present disclosure is shown.

[0127] In block 1010, a setup operation including creating a routing table is performed. The routing table may include a statically predefined routing table. In one exemplary implementation, during setup, the driver traverses the topology to collect information included in the routing table. The routing table may include an indication of the minimum number of links in the path to the target PPU. In some embodiments, no routing links are disabled. In some embodiments, the PPU includes a routing table that includes additional information about the next PPU (e.g., in the forwarding chain). In some embodiments, the routing table includes information about how many available links are from the source to the destination. In some embodiments, the routing table includes information about multiple paths from the source PPU to the target PPU, including, for example, multiple intermediate PPUs, one or more of which may be identified as the "next PPU" in the routing table.

[0128] In block 1020, multiple communication data packets are forwarded from a source parallel processing unit (PPU). According to embodiments of this disclosure, at least two of the multiple communication data packets are forwarded from different ports of the source PPU. In some embodiments, at least two of the multiple communication data packets are forwarded on different paths to a target PPU. In some embodiments, at least two of the multiple communication data packets are received and forwarded by different relay PPUs en route to the target PPU. In one embodiment, each of the multiple communication data packets is formed and forwarded according to a statically predetermined routing table.

[0129] In one embodiment, multiple communication packets are forwarded based on a routing scheme. In an exemplary embodiment, a routing scheme at the source PPU is determined by the following steps: creating a flow ID associated with multiple interconnected communication paths. In some embodiments, one or more communication packets from the multiple communication packets advance from the source PPU to a relay PPU.

[0130] In one exemplary embodiment, the routing scheme at the relay PPU includes selecting an egress port for each of a plurality of communication packets, wherein selecting the egress port includes: creating a flow ID associated with a unique communication path via interconnection, wherein the flow ID is established by hashing a plurality of bits selected in a physical address; mapping the source PPU ID and the flow ID; determining a plurality of possible egress ports available based on the mapping; determining the minimum link path to the destination using an appropriate routing table; and establishing a route selection based on the flow ID, the plurality of possible egress ports, and the minimum link path.

[0131] In block 1030, multiple communication data packets are received at the target parallel processing unit (PPU). In one embodiment, the source PPU and the target PPU are included in corresponding processing units among multiple processing units in the network, wherein a first processing core group of multiple processing cores is included in a first chip, a second processing core group of multiple processing cores is included in a second chip, and the multiple processing units communicate through multiple interconnects and the corresponding communication is configured according to a static predetermined routing table.

[0132] In some embodiments, at least two of the multiple communication data packets travel through different paths from the source PPU to the target PPU. In some embodiments, at least two of the multiple communication data packets are received by different relay PPUs, and at least two of the multiple communication data packets are forwarded to the target PPU by different relay PPUs.

[0133] In summary, embodiments of this disclosure provide improvements to the functionality of conventional computing systems and applications such as neural networks and AI workloads executed on such systems. More specifically, embodiments of this disclosure provide methods, programming models, and systems that improve the operational speed of applications such as neural networks and AI workloads by increasing the transmission speed of memory access requests (e.g., read requests and write requests) and the completion speed of result data transfers between system elements.

[0134] While the foregoing disclosure has illustrated various embodiments using specific block diagrams, flowcharts, and examples, each block diagram component, flowchart step, operation, and / or component described and / or illustrated herein can be implemented individually and / or collectively using a wide range of configurations. Furthermore, because many other architectures can be implemented to achieve the same functionality, any disclosure of a component contained within other components should be considered as an example.

[0135] Although the subject matter has been described in language specific to structural features and / or methodological behavior, it should be understood that the subject matter defined in this disclosure is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are disclosed as exemplary forms for implementing this disclosure.

[0136] This is a description of embodiments according to the present disclosure. Although the present disclosure has been described in particular embodiments, it should not be construed as being limited to these embodiments, but rather as being interpreted in accordance with the foregoing claims.

Claims

1. A processing system for inter-chip communication, comprising: Multiple parallel processing units, each included in a first chip, comprising: Multiple processing cores; Multiple memories, wherein a first memory group of the multiple memories is coupled to a first processing core group of the multiple processing cores; Multiple interconnects, located in a chip interconnect network, are configured to communicatively couple the multiple parallel processing units, wherein each of the parallel processing units is configured to communicate through the chip interconnect network according to a corresponding routing table, the corresponding routing table being stored and residing in a register of the corresponding parallel processing unit, and the corresponding routing table including information on multiple paths to any other given parallel processing unit; Specifically, when the parallel processing unit acts as a source parallel processing unit, it selects the exit port based on the flow identifier established according to the physical address hash and the minimum number of links in the routing table according to the first routing algorithm; and when the parallel processing unit acts as a relay parallel processing unit, it selects the exit port based on the source parallel processing unit identifier, the flow identifier and the routing table according to the second routing algorithm, wherein the second routing algorithm is different from the first routing algorithm.

2. The processing system according to claim 1, wherein, The routing table is stored in registers associated with the plurality of parallel processing units.

3. The processing system according to claim 2, wherein, The routing table is reconfigurable.

4. The processing system according to claim 1, wherein, The routing table is compatible with the basic XY routing scheme.

5. The processing system according to claim 1, wherein, The corresponding routing table is static and is preset during the configuration and setting operations of the communication capabilities of the chip interconnect network and the plurality of parallel processing units; and The corresponding routing table is part of the chip interconnect network configuration and setup operation for the communication capabilities of the chip interconnect network and the plurality of parallel processing units, and can be reconfigured. The setup operation includes loading and storing the corresponding routing table in the registers of the chip interconnect network subsystem of the parallel processing unit.

6. The processing system according to claim 5, wherein, The corresponding routing table includes information about the next parallel processing unit in the routing sequence.

7. The processing system according to claim 1, wherein, The parallel processing unit is included in a first parallel processing unit group of the plurality of parallel processing units, the first parallel processing unit group is included in a first computing node, and the second parallel processing unit group of the plurality of parallel processing units is included in a second computing node of the plurality of parallel processing units.

8. The processing system according to claim 5, wherein, The corresponding parallel processing unit among the plurality of parallel processing units includes a corresponding routing table.

9. The processing system according to claim 5, wherein, When communication is initiated by a corresponding parallel processing unit among multiple parallel processing units, the corresponding parallel processing unit among multiple parallel processing units is regarded as the source parallel processing unit; When communication is conducted through a corresponding parallel processing unit among multiple parallel processing units, the corresponding parallel processing unit among the multiple parallel processing units is regarded as a relay parallel processing unit; or when communication ends at a corresponding parallel processing unit among multiple parallel processing units, the corresponding parallel processing unit among the multiple parallel processing units is regarded as a target parallel processing unit.

10. The processing system according to claim 9, wherein, The respective interconnects of the plurality of interconnects are configured for multi-stream equalization, wherein the source parallel processing unit supports parallel communication streams of up to a maximum number of links between the source parallel processing unit and the target parallel processing unit.

11. The processing system according to claim 9, wherein, The corresponding interconnects among the plurality of interconnects are configured for multi-flow balancing, wherein the relay parallel processing unit runs routes to balance the flow between egress points.

12. The processing system according to claim 1, wherein, Communication is handled according to a routing scheme that provides balanced workloads, guaranteed dependencies, and guaranteed access order.

13. A communication method, comprising: Perform setup operations, including creating a static predefined routing table; The communication data packets are forwarded from the source parallel processing unit, wherein the communication data packets are formed and forwarded according to the static predetermined routing table; The communication data packet is received in the target parallel processing unit; The source parallel processing unit and the target parallel processing unit are included in corresponding parallel processing units among multiple parallel processing units in the network. A first processing core group of multiple processing cores is included in a first chip, a second processing core group of multiple processing cores is included in a second chip, and the multiple parallel processing units communicate through multiple interconnects and are configured with corresponding communication according to the static predetermined routing table. The source parallel processing unit selects the exit port based on the flow identifier established by the physical address hash and the minimum number of links in the routing table according to the first routing algorithm; the relay parallel processing unit for forwarding selects the exit port based on the source parallel processing unit identifier, the flow identifier and the routing table according to the second routing algorithm, which is different from the first routing algorithm.

14. The communication method according to claim 13, wherein, The routing scheme for the source parallel processing unit is determined by the following steps: Create a flow identifier associated with a unique communication path through the interconnection, wherein the flow identifier is established by hashing a selected number of bits in a physical address; The minimum link path to the target is determined using the corresponding routing table; as well as Routing is established based on the flow identifier and the minimum link path.

15. The communication method according to claim 13, wherein, The routing scheme at the relay parallel processing unit includes selecting an exit port, wherein the selected exit port includes: Create a flow identifier associated with a unique communication path through the interconnection, wherein the flow identifier is established by a selected number of bits in a hashed physical address; Mapping source parallel processing unit identifier and the stream identifier; Based on the mapping, multiple possible available egress ports are determined; The minimum link path to the target is determined using the corresponding routing table; and Multiple routing options are established based on the flow identifier, the multiple possible exit ports, and the multiple paths.

16. The communication method according to claim 13, wherein, The communication method further includes: Balance the forwarding of the communication data packets. The equalization of the forwarding of the communication data packets includes: The communication data packets are distributed via interleaving based on physical addresses.

17. A processing system for inter-chip communication, comprising: A first parallel processing unit group, the first parallel processing unit group being included in a first computing node, and the corresponding parallel processing units in the first parallel processing unit group being included in separate corresponding chips; The second parallel processing unit group is included in the second computing node, and the corresponding parallel processing units in the second parallel processing unit group are included in separate corresponding chips. as well as Multiple interconnects, located in a chip interconnect network, are configured to communicatively couple a first parallel processing unit group and a second parallel processing unit group, wherein parallel processing units in the first and second parallel processing unit groups communicate through the multiple interconnects, and the communication is configured according to routing tables in the storage characteristics of the respective parallel processing units residing in the first and second parallel processing unit groups, wherein the multiple interconnects are configured to couple parallel communication from a first parallel processing unit in the first computing node to a second parallel processing unit in the second computing node through at least two paths; Specifically, when the parallel processing unit acts as a source parallel processing unit, it selects the exit port based on the flow identifier established according to the physical address hash and the minimum number of links in the routing table according to the first routing algorithm; and when the parallel processing unit acts as a relay parallel processing unit, it selects the exit port based on the source parallel processing unit identifier, the flow identifier and the routing table according to the second routing algorithm, wherein the second routing algorithm is different from the first routing algorithm.

18. The processing system according to claim 17, wherein, The corresponding parallel processing units in the first parallel processing unit group and the second parallel processing unit group include: Multiple processing cores, wherein a corresponding group of processing cores is included in the corresponding parallel processing unit; and Multiple memories, wherein corresponding memory groups are communicatively coupled to the corresponding processing core groups and are included in the corresponding parallel processing units.

19. The processing system according to claim 17, wherein, Multiple processing units communicate through the multiple interconnections and configure corresponding communications according to a routing table, which is static and predetermined. Before normal processing operations are performed, the routing table is loaded as part of the settings of the multiple processing units into registers associated with the multiple processing units. The routing table includes indications of multiple links between sources and destinations.

20. The processing system according to claim 17, wherein, A balanced workload is provided by multiple streams through the minimum link, as well as interleaved stream IDs and source parallel processing unit IDs.

Citation Information

Patent Citations

  • Data processing system, method and interconnect fabric for improved communication in a data processing system

    US20060176890A1