Parallel processing units, processing systems, and related methods

By employing parallel processing units and external units in the heterogeneous chip interconnect network, individual addressing of each memory space is achieved, solving the problem of long writing or reading time for multi-terabyte memory and improving data transmission efficiency.

CN116226023BActive Publication Date: 2026-07-31T-HEAD (SHANGHAI) SEMICON CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
T-HEAD (SHANGHAI) SEMICON CO LTD
Filing Date
2021-12-02
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In heterogeneous chip interconnect networks, writing to or reading shared memory rows of multi-terabyte memory can take a significant amount of time.

Method used

By employing parallel processing units and external units, and addressing each memory space individually, data transmission latency in the inter-chip network is reduced.

Benefits of technology

By addressing each memory space individually, the time for reading and writing from multi-TB memory chips is significantly reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116226023B_ABST
    Figure CN116226023B_ABST
Patent Text Reader

Abstract

This disclosure provides a parallel processing unit, a processing system, and related methods. The parallel processing unit includes: a core; a cache coupled to the core; a processor memory coupled to the core; and an inter-chip network controller coupled to the core. The inter-chip network controller includes a routing table containing multiple addresses that identify multiple memory spaces directly and indirectly coupled to multiple devices of the parallel processing unit. The multiple memory spaces in the multiple devices have the same size, and each address is used to identify a memory row within the memory space. This disclosure reduces the time required to read from and write to multi-TB memory chips in an inter-chip network by dividing each memory chip into multiple memory spaces and then addressing each memory space individually using an address identifying the memory space, a row number within the memory space, and multiple transmit / receive ports for accessing the row number.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to routing schemes for chip interconnect networks, and more particularly to routing schemes for heterogeneous chip interconnect networks using distributed shared memory. Background Technology

[0002] A homogeneous interconnected-chip network is a network that couples together multiple essentially identical chips, while a heterogeneous interconnected-chip network couples together multiple different chips. For example, a large-scale processing system may include a large number of parallel processing chips and a large number of multi-terabyte memory chips that are directly and indirectly coupled together through a heterogeneous interconnected-chip network.

[0003] In one example of a traditional routing scheme, each chip in a heterogeneous chip interconnect network includes a routing table that uses memory addresses to identify another chip in the network, a shared line of memory within the chip, and one or more ports to be used when transferring data to or from the shared line of memory.

[0004] In heterogeneous chip interconnect networks, one drawback of this type of routing scheme is that writing to or reading from shared memory rows of multi-terabyte memory is time-consuming. Therefore, a routing scheme that consumes less time for writing to or reading from multi-terabyte memory rows is needed. Summary of the Invention

[0005] This disclosure provides a routing scheme that reduces the latency associated with transferring data to and from multi-terabyte memory.

[0006] This disclosure includes a parallel processing unit comprising: a core; a cache coupled to the core; and processor memory coupled to the core. The parallel processing unit also includes an inter-chip network (ICN) controller coupled to the core. The ICN controller includes a routing table. The routing table includes multiple addresses that identify multiple memory spaces in multiple devices directly and indirectly coupled to the parallel processing unit. The multiple memory spaces in the devices have the same size. Each address identifies a memory row in the memory space.

[0007] This disclosure also includes a processing system comprising multiple parallel processing units. Each parallel processing unit includes: a core; a cache coupled to the core; and processor memory coupled to the core. Each parallel processing unit also includes an inter-chip network (ICN) controller coupled to the core. The ICN controller includes a first routing table. The first routing table includes multiple addresses that identify multiple memory spaces in multiple devices directly and indirectly coupled to the parallel processing units. The multiple memory spaces in the devices have the same size. Each address identifies a memory row in the memory space. The processing system also includes multiple external units directly and indirectly coupled to the multiple parallel processing units. Each external unit includes extended memory, such that the multiple external units have multiple extended memories. The extended memory in each external unit includes multiple memory spaces. The multiple memory spaces in the extended memory have the same size. The extended memory includes a second routing table. The second routing table includes multiple first addresses that identify multiple memory spaces in each extended memory. Each address identifies a memory row in the memory space.

[0008] This disclosure also includes a method for forming a parallel processing unit. The method includes: forming a core; forming a cache coupled to the core; and forming processor memory coupled to the core. The method further includes forming a routing table including multiple addresses that identify multiple memory spaces in multiple devices directly and indirectly coupled to the parallel processing unit. The multiple memory spaces in the devices have the same size. Each address identifies a memory row in the memory space.

[0009] The above scheme enables individual addressing of each memory space, thereby reducing the time required to read from and write to multi-TB memory chips in an inter-chip network.

[0010] A better understanding of the features and advantages of this disclosure will be obtained by referring to the following detailed description and accompanying drawings, which illustrate illustrative embodiments in which the principles of this disclosure are utilized. Attached Figure Description

[0011] The accompanying drawings described herein are provided to further understand this disclosure and constitute a part of this disclosure. The exemplary embodiments of this disclosure and their description are used to explain this disclosure and do not constitute a limitation thereof.

[0012] Figures 1A to 1D Block diagrams of four example topologies of a processing system 100 according to embodiments of the present disclosure are shown.

[0013] Figure 2 A block diagram of an example parallel processing unit (PPU) 200 according to an embodiment of the present disclosure is shown.

[0014] Figure 3 A block diagram of an example external unit (SMX) 300 according to an embodiment of the present disclosure is shown.

[0015] Figure 4 A block diagram of an example processing system 400 according to an embodiment of the present disclosure is shown.

[0016] Figure 5 An example memory address 500 according to an embodiment of this disclosure is shown.

[0017] Figure 6 An example address allocation 600 according to an embodiment of this disclosure is shown. Specific Implementation

[0018] Figures 1A to 1D Block diagrams of four example topologies of a processing system 100 according to embodiments of the present disclosure are shown. As described in more detail below, the processing system 100 addresses each shared memory space individually to reduce the latency associated with transferring data to and from memory.

[0019] like Figures 1A to 1D As shown, the processing system 100 includes multiple parallel processing units (PPUs) 110 and multiple external units (SMXs) 112 directly and indirectly coupled to the PPUs 110. Further, as... Figures 1A to 1D As shown, PPU 110 and SMX112 can be coupled together in a variety of different ways.

[0020] Figure 2 A block diagram of an example parallel processing unit (PPU) 200 according to an embodiment of the present disclosure is shown. In this example, each PPU 110 can be implemented using PPU 200. Figure 2 As shown, the PPU 200 includes multiple cores 210 (four in this example) and multiple local caches 212, which are coupled to the cores 210 such that each core 210 has a corresponding local cache 212.

[0021] Furthermore, such as Figure 2As shown, the PPU 200 also includes multiple processor memories (HBM) 214 coupled to core 210, such that each core 210 has a corresponding HBM 214. HBMs can be implemented in various ways. In one example, the HBM is implemented as a high-bandwidth memory comprising multiple dynamic random access memory (DRAM) dies stacked vertically on top of each other to provide a large storage capacity with a small form factor, and the high-bandwidth memory also includes two 128-bit data channels per die to provide high bandwidth. For example, the maximum size of the HBM could be 4GB, 24GB, and 64GB.

[0022] Each HBM 214 is divided into a first address range and a second address range, such that the first address range can only be accessed by the core 210 associated with the HBM 214, while the second address range, as a shared address range, can be accessed by all cores 210 on the PPU 200 and cores in other PPUs in the processing system (e.g., other PPUs 110 in the processing system 100).

[0023] In this example, the second address range shared across the PPU addresses 256GB of memory space, although only a small portion (e.g., 64GB) is actually available. In other words, each PPU has a shared memory space addressable over 256GB, but the usable space in that shared memory space may be less than 256GB. Furthermore, each addressable 256GB memory space represents a virtual device ID.

[0024] PPU 200 also includes a network-on-chip (NoC) 216, which couples core 210 and HBM 214 together to provide a high-bandwidth, high-speed communication path between core 210 and HBM 214. PPU 200 also includes an inter-chip network (ICN) controller 220, which is coupled to core 210 via NoC 216 and directly coupled to multiple other PPUs and SMXs, such as other PPUs 110 and SMX 112, via the inter-chip network.

[0025] Accordingly, the ICN controller 220 includes communication control circuitry 222, a switch 226, and multiple inter-chip ports 228. The communication control circuitry 222 is coupled to the core 210 via NoC 216, and the multiple inter-chip ports 228 are coupled to the switch 226 and directly coupled to other PPUs and SMXs. In this example, seven inter-chip ports 228 are used, namely inter-chip ports P0 to P6. Ports 228 include transmit and receive circuitry, and the control circuitry 222 controls the incoming and outgoing data flow of the inter-chip ports 228 via the switch 226.

[0026] Furthermore, as described in more detail below, PPU 200 includes a routing table 230 generated and subsequently stored in ICN controller 220. Initially, routing table 230 is stored in host memory, but is subsequently programmed by the driver into a routing register (hardware table) in ICN controller 220. There are two routing registers in ICN controller 220: one in control circuitry 222 and another in switch 226. In addition to generating and responding to read and write requests, each PPU 200 also functions as a forwarding device, forwarding read / write requests and data from one device to another.

[0027] In other words, routing table 230 identifies ports 228 in ICN controller 220 used for communication with directly and indirectly connected devices. Furthermore, although not shown for simplicity, PPU 200 may also include additional circuitry typically included in the processing chip.

[0028] Figure 3 A block diagram of an example external unit (SMX) 300 according to an embodiment of the present disclosure is shown. In this example, each SMX 112 can be implemented using an SMX 300. Figure 3 As shown, the SMX 300 includes an extended memory 310, a memory control circuit 312, and an inter-chip network (ICN) controller 314. The memory control circuit 312 is coupled to the extended memory 310, and the inter-chip network (ICN) controller 314 is coupled to the extended memory 310, the memory control circuit 312, and multiple other directly coupled PPUs and / or inter-chip networks in the SMX.

[0029] ICN controller 314 can be implemented in a manner similar to ICN controller 220, including communication control circuitry 322 coupled to extended memory 310 and control circuitry 312, switch 324 coupled to control circuitry 322, and multiple inter-chip ports 326 coupled to switch 324 and one or more directly coupled devices.

[0030] In this example, seven inter-chip ports 326 are used, namely inter-chip ports P0 to P6. Although seven ports are shown, one port is sufficient (as a slave device) to provide basic access. Additional ports provide additional bandwidth. Furthermore, switch 324 may not be necessary. If switch 324 is present, the SMX can further function as a bridge for forwarding packets; otherwise, the SMX can simply act as an end device providing accessibility. In short, the number of ports on the SMX is independent, and the switch is optional and determined by the SMX's IP designer. Port 326 includes transmit and receive circuitry, while control circuitry 322 controls the incoming and outgoing data flow through inter-chip ports 326 via switch 324.

[0031] The extended memory 310 in the SMX 300 includes multiple memory spaces of equal size. For example, each memory space can be 256GB in size. Furthermore, each 256GB memory space represents a virtual device ID. In this example, the address space of the memory in each extended memory 310 is a shared memory space accessible by each parallel processing unit.

[0032] In addition, the SMX 300 includes a routing table 316 generated and subsequently stored in the ICN controller 314. Similar to routing table 230, routing table 316 is programmed by the driver into a routing register (hardware table) in the ICN controller 314. When control circuitry 322 and switch 324 are present, there are two routing registers in the ICN controller 314, located in control circuitry 322 and switch 324 respectively.

[0033] In addition to receiving read and write requests from the PPU in the system, each SMX 300 also functions as a forwarding device. Routing table 316 identifies ports in the ICN controller 314, which are used for directly and indirectly connected devices in response to requests. Ports are also used to forward read / write requests and data from one device to another. Furthermore, although not shown for simplicity, the SMX 300 may also include additional circuitry typically included in the memory chip.

[0034] Figure 4 A block diagram of an example processing system 400 according to an embodiment of the present disclosure is shown. In this example, processing system 400 represents a portion of processing system 100. Figure 4 As shown, the processing system 400 includes three PPUs (i.e., PPU 410, PPU 412 and PPU 414) and two SMXs (i.e., SMX 416 and SMX 418).

[0035] Furthermore, in this example, PPU 410 is coupled to SMX 416 via ports 420-0, 420-1, and 420-2, to PPU 414 via ports 420-3 and 420-4, and to SMX 418 via ports 420-5 and 420-6. Correspondingly, PPU 412 is coupled to SMX 416 via ports 430-3 and 430-4, to PPU 414 via port 430-6, and to SMX 418 via port 430-5. PPU 414 is coupled to PPU 412 via port 440-0 and to PPU 410 via ports 440-2 and 440-3.

[0036] In practice, when transmitting data, large amounts of data can be divided into multiple parts and transmitted in parallel, thus significantly reducing the time required for data transmission. The number of parts into which the data is divided is defined by the minimum number of links between the source and destination devices.

[0037] exist Figure 4 In the example, PPU 410 is coupled to SMX 416 via three links and ports, while PPU 412 is coupled to SMX 416 via two links and ports, making the minimum number of links between PPU 410 and PPU 412 2. In this scenario, the most efficient data transfer occurs when the data is split into two parts based on the minimum number of links of 2. This is because, although data can be transferred from PPU 410 to SMX 416 in a shorter time when split into three parts, the conversion from three input streams to two outputs to be forwarded to PPU 412 creates a bottleneck at SMX 416, which takes longer.

[0038] Figure 5 An example memory address 500 according to an embodiment of this disclosure is shown. As an example, such as Figure 5 As shown, memory address 500 has a 48-bit address, where the first 10 bits (B47-B38) identify 1024 256GB memory spaces for 1024 virtual devices. For example, SMX 416 may include 16 256GB memory spaces representing 16 virtual devices 416-0 to 416-15. In this example, the first ten bits may identify the third 256GB memory space / virtual device 416-2 in SMX 416, while the remaining 38 bits (B37-B0) identify the rows within that third 256GB memory space / virtual device 416-2.

[0039] Table 1 below shows an example routing table according to an embodiment of the present disclosure. Routing table 230 is implemented as shown in FIG1, and routing table 316 is implemented in a similar manner.

[0040] Table 1

[0041]

[0042]

[0043] As shown in Table 1, the routing table includes a destination ID (dest ID) column, a minimum link (minLink) column, and seven port columns P0-P6. Further, as shown in Table 1, each entry in the dest ID column represents a memory row in the memory space and a 48-bit physical address identifying the port used to transfer data from PPU 410 to the memory row. The first ten bits identify one of 1024 memory spaces / virtual devices in SMX416, PPU 412, PPU 414, and SMX 418.

[0044] In this example, each 256GB memory space represents a different virtual device ID, and each virtual device ID identifies the number of ports used to send, receive, and forward data to and from the virtual device. In this example, the SMX 416 has 16 256GB memory spaces, representing 16 virtual device IDs: SMX 416-0 to SMX 416-15. Additionally, the SMX 418 has 4 256GB memory spaces, representing 4 virtual device IDs: SMX 418-0 to SMX 418-3.

[0045] In addition, each address also identifies a memory row in the memory space and a port used to transfer data from one device to another. Furthermore, as shown in Table 1, the minimum link from PPU 410 to SMX 416 is 3, and ports P0-P2 are identified as 3 transmit / receive ports, while ports P3-P6 are not identified.

[0046] Furthermore, as shown in Table 1, the shared memory space in each PPU (e.g., PPU 412) is addressed as a 256GB memory space, although only a small amount of memory space is available. Therefore, the routing table has multiple addresses to identify the 256GB memory space in the PPU and SMX. The shared memory space in the PPU can be combined with the memory space in the SMX to form a global shared address range for parallel processing units. The highest 10 bits are the virtual device ID, which is used as the high-order bits of the global physical address space.

[0047] Figure 6 An example address allocation 600 according to an embodiment of this disclosure is shown. As an example, such as... Figure 6 As shown, addresses are allocated to 16 256GB memory spaces in the SMX such that the first 10 bits (the four least significant bits, B41-B38) identify the 16 256GB memory spaces. Additionally, addresses are allocated to 4 256GB memory spaces in another SMX such that the first 10 bits (the two least significant bits, B39-B38) identify the 4 256GB memory spaces.

[0048] One of the advantages of this disclosure is that by using SMX memory, which comprises multiple 256GB memory spaces, and by addressing each 256GB memory space individually, the time required to transmit and receive data can be significantly reduced. Although this disclosure is described in the context of 256GB memory space, other sizes of memory space, such as 512GB memory space, can also be utilized.

[0049] In this example, during system startup, routing tables 230 and 316 in the PPU and SMX are populated by assigning static identities to each 256GB of memory space. Identifiers can be assigned in various ways, such as by physical configuration or via a software-configured identifying switch, or via self-assigning software where an identity-generating token is sent to each device in the processing system, and then the self-assigning software self-assigns a unique identifier before sending the token to the next device in the processing system.

[0050] Subsequently, each PPU broadcasts a message to each directly connected PPU and SMX, which responds to the broadcast message by identifying itself, its device type (PPU, SMX-4GB, SMX-24GB, SMX-64GB), its device address, and the identifiers of its coupled PPUs and SMXs. The address can be used as an index as shown in Table 1.

[0051] After several rounds, each PPU and SMX determines one or more one-hop or multi-hop paths from each PPU and SMX to every other PPU and SMX in the processing system, including the most efficient paths. Subsequently, each PPU and SMX allocates one or more inter-chip ports, such as ports P0-P6, based on path efficiency, and then populates routing tables 230 and 316 in the PPU and SMX. Additional periodic broadcast messages are used to detect faulty or malfunctioning devices and to maintain the up-to-date routing tables.

[0052] Various embodiments of this disclosure have now been described in detail, examples of which are illustrated in the accompanying drawings. Although described in conjunction with various embodiments, it should be understood that these various embodiments are not intended to limit this disclosure. Rather, this disclosure is intended to cover alternatives, modifications, and equivalents that may be included within the scope of this disclosure as interpreted by the claims.

[0053] Furthermore, numerous specific details have been set forth in the foregoing detailed description of the various embodiments of this disclosure to provide a thorough understanding of the disclosure. However, those skilled in the art will recognize that this disclosure may be practiced without these specific details or with their equivalents. In other instances, well-known methods, processes, components, and circuits have not been described in detail to avoid unnecessarily obscuring aspects of the various embodiments of this disclosure.

[0054] It is important to note that although this paper may describe the methods as a sequence of operations for clarity, the described sequence of operations does not necessarily specify the order of operations. It should be understood that some operations may be skipped, executed in parallel, or performed without needing to maintain a strict sequence order.

[0055] The accompanying drawings illustrating various embodiments of this disclosure are semi-illustrative and not drawn to scale; in particular, some dimensions are enlarged for clarity of expression. Similarly, although views in the drawings are generally shown in similar orientations for ease of description, such descriptions in the drawings are arbitrary in most cases. Generally, the various embodiments of this disclosure can be operated in any orientation.

[0056] Some parts of this disclosure can be represented by other notations of programs, logic blocks, processes, and data bit operations in computer memory. Those skilled in the art of data processing use these descriptions and representations to effectively communicate the substance of their work to others skilled in the art.

[0057] In this disclosure, programs, logic blocks, procedures, etc., are considered as self-consistent sequences of operations or instructions that lead to desired results. These operations are physical operations that utilize physical quantities. Typically, although not strictly necessary, these quantities exist in the form of electrical or magnetic signals that can be stored, transmitted, combined, compared, and otherwise manipulated in a computing system. It has proven convenient, primarily for common usage reasons, to refer to these signals as transactions, bits, values, elements, symbols, characters, samples, pixels, etc.

[0058] However, it should be remembered that all these and similar terms are associated with appropriate physical quantities and are merely convenient labels applicable to those quantities. Unless explicitly stated otherwise, it should be understood from the following discussion that throughout this disclosure, discussions using terms such as “generate,” “determine,” “allocate,” “aggregate,” “utilize,” “virtualize,” “process,” “access,” “execute,” and “store” refer to the actions and processes of a computer system or similar electronic computing device or processor.

[0059] A processing system or similar electronic computing device or processor manipulates or transforms data represented as physical (electronic) quantities in computer system memory, registers, other such information storage, and / or other computer-readable media into other data similarly represented as physical quantities in computer system memory, or registers, or other such information storage, transmission or display devices.

[0060] The technical solutions in the embodiments of this disclosure have been clearly and completely described in the preceding chapters with reference to the accompanying drawings of the embodiments of this disclosure. It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific sequence or order. It should be understood that these numbers can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in an order other than that shown or described herein.

[0061] The functions described in this embodiment, if implemented as software functional units and sold or used as independent products, can be stored in a computing device-readable storage medium. Based on this understanding, a part of the embodiments or technical solutions of this disclosure that contribute to the prior art can be embodied in the form of a software product stored in a storage medium including a plurality of instructions for causing a computing device (which may be a personal computer, server, mobile computing device, or network device, etc.) to perform all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes: a USB drive, portable hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk, optical disk, etc., capable of storing program code.

[0062] The various embodiments in this disclosure are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical portions between embodiments may be referred to in another context. These embodiments are merely a portion of the embodiments, and not all embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without departing from the scope of the invention are within the scope of this disclosure.

[0063] The above embodiments are for illustrative purposes only and are not intended to limit the technical solutions of this disclosure. Although this disclosure has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions recorded in the above embodiments can still be modified, or some or all of the technical features can be equivalently substituted. These modifications or substitutions will not cause the substance of the corresponding technical solutions to deviate from the scope of the technical solutions in the embodiments of this disclosure.

[0064] It should be understood that the above description is an example of this disclosure, and various alternatives to this disclosure described herein can be used to implement this disclosure. Therefore, the foregoing claims are intended to define the scope of this disclosure, and structures and methods within the scope of these claims, as well as their equivalents, should be included.

Claims

1. A parallel processing unit, comprising: nuclear; The cache is coupled to the kernel; On-screen network; The processor memory is coupled to the core via the on-chip network; as well as An inter-chip network controller is coupled to the core via the on-chip network. The inter-chip network controller includes a routing table, which includes multiple addresses. Each address includes a space identifier portion and a row identifier portion. The space identifier portion is used to identify the memory space of a device, and the row identifier portion is used to identify the memory row in the memory space. Multiple memory spaces in multiple devices have the same size.

2. The parallel processing unit of claim 1, wherein, The processor memory is divided into a first address range and a second address range, such that the first address range is accessed only by the core, and the second address range is a shared address range accessed by each device directly or indirectly coupled to the parallel processing unit.

3. The parallel processing unit according to claim 1, wherein the inter-chip network controller is coupled to each device directly coupled to the parallel processing unit.

4. The parallel processing unit of claim 3, wherein, The inter-chip network controller includes multiple ports that are directly coupled to one or more of the multiple devices, and an address in a first routing table is used to identify the one or more ports associated with that address.

5. A processing system, comprising: Multiple parallel processing units, each of which includes: nuclear; The cache is coupled to the kernel; On-screen network; Processor memory, coupled to the core via the on-chip network; and A first inter-chip network controller, coupled to the core via the on-chip network, includes a first routing table comprising multiple addresses. Each address includes a space identifier portion and a row identifier portion. The space identifier portion identifies the memory space of a device, and the row identifier portion identifies a memory row within that memory space. Multiple memory spaces in multiple devices have the same size. Multiple external units are directly and indirectly coupled to the multiple parallel processing units. Each external unit includes an extended memory, such that the multiple external units include multiple extended memories. The extended memory in each external unit includes multiple memory spaces of the same size. The extended memory includes a second routing table, which includes multiple first addresses for identifying the multiple memory spaces in each extended memory. Each address is used to identify a memory row in a memory space.

6. The processing system of claim 5, wherein, Multiple memory spaces within each extended memory serve as a shared memory space, accessible by each parallel processing unit.

7. The processing system of claim 6, wherein, The processor memory of the parallel processing unit is divided into a first address range and a second address range, such that the first address range is accessed only by the core of the parallel processing unit, and the second address range is a shared address range accessed by each other parallel processing unit directly or indirectly coupled to the parallel processing unit; wherein, the first routing table includes addresses identifying the shared address range of each parallel processing unit, and the second routing table of the external unit includes multiple second addresses, which are used to identify the shared address range of each parallel processing unit.

8. The processing system of claim 5, wherein, The first inter-chip network controller is coupled to each of the other parallel processing units and external units directly coupled to the parallel processing units; the external units also include a second inter-chip network controller, which is coupled to the extended memory, the parallel processing units, and other external units directly coupled to the external units.

9. The processing system of claim 8, wherein, The first inter-chip network controller includes multiple ports that are directly coupled to one or more of the external units, and the addresses in the first routing table are used to identify one or more ports associated with those addresses.

10. The processing system of claim 5, wherein, Each address has a first number of bits and a second number of bits, the first number of bits identifying the memory space and the second number of bits identifying the row within the memory space.

11. A method for forming a parallel processing unit, the method comprising: Nucleus formation; A cache is formed that is coupled to the core; A processor memory is formed that is coupled to the core via an on-chip network; An inter-chip network controller is formed and coupled to the core via an on-chip network. The inter-chip network controller includes a routing table, which includes multiple addresses. Each address includes a space identifier portion and a row identifier portion. The space identifier portion is used to identify the memory space of a device, and the row identifier portion is used to identify the memory row in the memory space. Multiple memory spaces in multiple devices have the same size.

12. The method of forming of claim 11, wherein, The processor memory is divided into a first address range and a second address range, such that the first address range is accessed only by the core, and the second address range is a shared address range accessed by each device directly or indirectly coupled to the parallel processing unit.

13. The method of forming according to claim 12, wherein the inter-chip network controller is coupled to the core and to each device directly coupled to the parallel processing unit, the inter-chip network controller including a plurality of ports directly coupled to one or more of the plurality of devices, and an address in a first routing table is used to identify one or more ports associated with a virtual address.