Multi-core tile integrated processor architecture, method, processor and electronic device

By introducing multiple parallel chip-to-chip links and chip hubs into the processor architecture, data routing is optimized, solving the problem of limited communication bandwidth between computing chips and input/output chips, and achieving efficient cross-chip data transmission and increased throughput.

CN122045121BActive Publication Date: 2026-07-31SUZHOU YIZHU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SUZHOU YIZHU INTELLIGENT TECH CO LTD
Filing Date
2026-04-13
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In existing processor architectures, the communication bandwidth between computing chips and input/output chips is limited, resulting in bottlenecks and link congestion in cross-chip data transmission, which cannot meet the transmission requirements of high-throughput data streams.

Method used

The architecture employs a multi-core integrated processor architecture. By establishing multiple parallel core-to-core links between computing cores and input/output cores, and introducing a core hub within the computing core, parallel data transmission and distribution are achieved. This optimizes data routing logic and ensures that each block processing cluster can independently drive multiple links for data transmission.

Benefits of technology

It effectively expands the cross-core transmission bandwidth, reduces transmission latency, improves the throughput and data transmission efficiency of the processor architecture, and avoids link congestion and data accumulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122045121B_ABST
    Figure CN122045121B_ABST
Patent Text Reader

Abstract

This disclosure provides a multi-core integrated processor architecture, method, processor, and electronic device. The multi-core integrated processor architecture includes input / output cores and computing cores. Each computing core includes a core hub and multiple block processing clusters. Each block processing cluster is connected to the core hub via a cluster connection link. Each core hub is connected to the input / output cores via multiple parallel core-to-core links. A single block processing cluster simultaneously drives multiple core-to-core links between the computing cores and the input / output cores for data transmission through the core hub. In this disclosure, the internal bandwidth of the block processing cluster matches the external link bandwidth, which can improve the throughput and transmission efficiency of the processor architecture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of processor technology, and in particular to a multi-chip integrated processor architecture, method, processor, and electronic device. Background Technology

[0002] With the rapid growth in demand for artificial intelligence and high-performance computing, processors are now widely adopting chip-to-die integrated architectures to overcome the limitations of single-chip area. In existing technologies, compute dies and input / output (IO) dies are interconnected via die-to-die (D2D) links. Compute dies typically contain multiple block processing clusters to support parallel computing, but only a single die-to-die link is usually deployed between compute dies and IO dies, resulting in severely limited cross-die communication bandwidth. When block processing clusters generate high-throughput data streams, the current link architecture cannot provide sufficient transmission capacity, causing link congestion and data backlog. Summary of the Invention

[0003] This disclosure provides a multi-chip integrated processor architecture, method, processor, and electronic device that can match the internal bandwidth of the block processing cluster with the external link bandwidth, effectively solve the data transmission bottleneck problem, and improve the throughput and transmission efficiency of the processor architecture.

[0004] In a first aspect, embodiments of this disclosure provide a multi-core integrated processor architecture, including input / output cores and computing cores. The computing core includes a core hub and multiple block processing clusters. Each block processing cluster is connected to the core hub via a cluster connection link. Each core hub is connected to the input / output core via multiple parallel core-to-core links.

[0005] In this context, a single block processing cluster simultaneously drives multiple chip-to-chip links between the computing chip and the input / output chip to transmit data through the chip hub.

[0006] In the multi-chip integrated processor architecture of this disclosure embodiment, the chip hub includes multiple inter-cluster buses and multiple link ports that are connected one-to-one with the chip-to-chip links. The cluster connection links of multiple block processing clusters are interconnected within the chip hub through the inter-cluster buses, and each cluster connection link is connected to the inter-cluster buses and the link ports within the chip hub.

[0007] In the multi-chip integrated processor architecture of this disclosure embodiment, the chip hub structure includes multiple cluster ports, which are connected to multiple block processing clusters through multiple cluster connection links; the cluster ports connected to different block processing clusters are interconnected, and the link ports are respectively connected to each of the cluster ports.

[0008] In the multi-chip integrated processor architecture of this disclosure embodiment, the cluster connection links connected to each of the block processing clusters include a first cluster connection link and a second cluster connection link. The chip-to-chip link between the chip hub and the input / output chip includes a first chip-to-chip link and a second chip-to-chip link. The plurality of link ports include a first link port and a second link port. A first branch of the first cluster connection link is connected to the first inter-cluster bus, and a second branch is connected to the first link port. A first branch of the second cluster connection link is connected to the second inter-cluster bus, and a second branch is connected to the second link port. The first link port is connected to the first chip-to-chip link, and the second link port is connected to the second chip-to-chip link.

[0009] In the multi-chip integrated processor architecture of this disclosure embodiment, the total bandwidth of all cluster connection links connected to a single block processing cluster is a multiple of the bandwidth of a single chip-to-chip link.

[0010] In the multi-chip integrated processor architecture of this disclosure embodiment, the total bandwidth of all cluster connection links connected to a single block processing cluster is equal to the total bandwidth of multiple chip-to-chip links.

[0011] In the multi-chip integrated processor architecture of this disclosure embodiment, each block processing cluster includes a cluster hub, an intra-cluster shared memory, and multiple computing units. The intra-cluster shared memory and the multiple computing units are respectively connected to the cluster hub. The cluster hub is connected to the chip hub through multiple cluster connection links. The cluster hub is used to distribute data from the multiple computing units to the multiple cluster connection links connected to the cluster hub for parallel transmission.

[0012] In the multi-chip integrated processor architecture of this disclosure embodiment, the cluster hub is further configured to aggregate the data generated by the multiple computing units, and distribute the aggregated data to the cluster connection links for parallel transmission at a rate equal to the total bandwidth of all the cluster connection links connected to the cluster hub.

[0013] In the multi-chip integrated processor architecture of this disclosure embodiment, the data generated by a single computing unit is distributed through the cluster hub to multiple cluster connection links connected to the cluster hub for parallel transmission.

[0014] In the multi-chip integrated processor architecture of this disclosure embodiment, the transmission link bandwidth between the intra-cluster shared memory and the cluster hub is equal to the total bandwidth of all cluster connection links connected to a single cluster hub.

[0015] In the multi-chip integrated processor architecture of this disclosure embodiment, the number of block processing clusters connected to a single chip hub is a multiple of the number of chip-to-chip links between a single chip hub and the input / output chips.

[0016] In the multi-chip integrated processor architecture of this disclosure embodiment, the input / output chip further includes an external switching port for connection to an external network switching device, and a cross-chip transmission port for connection to other input / output chips; wherein, the signal types transmitted within the chip hub include:

[0017] Cross-cluster communication signal, wherein the transmission path of the cross-cluster communication signal is that any of the block processing clusters is transmitted to another block processing cluster within the same core hub via the inter-cluster bus;

[0018] And / or, cross-core communication signals, wherein the transmission path of the cross-core communication signals is that any of the block processing clusters is transmitted to the cross-core transmission port through the link port of the core hub;

[0019] And / or, external exchange communication signals, wherein the transmission path of the external exchange communication signals is that any of the block processing clusters is transmitted to the external exchange port through the link port of the core hub.

[0020] Secondly, this disclosure also proposes a data transmission method applied to the multi-chip integrated processor architecture described in the first aspect embodiment, the data transmission method comprising:

[0021] In response to the block processing cluster sending a cross-core communication signal, the cross-core communication signal is split into multiple cross-core communication sub-signals and sent in parallel to multiple link ports through all the cluster connection links connected by the core hub, so that multiple core-to-core links are simultaneously driven to transmit the cross-core communication signal.

[0022] Furthermore, embodiments of this disclosure also propose that, in response to the block processing cluster sending an inter-cluster communication signal, the inter-cluster communication signal is split into multiple inter-cluster communication sub-signals, and sent in parallel to multiple inter-cluster buses through all the cluster connection links connected by the core hub, so as to transmit the inter-cluster communication signal to the cluster connection links of other block processing clusters.

[0023] Thirdly, embodiments of this disclosure also provide a processor comprising the multi-chip integrated processor architecture as described in the first aspect embodiments.

[0024] Fourthly, embodiments of this disclosure also provide an electronic device, the electronic device including a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for implementing connection communication between the processor and the memory, wherein the program is executed by the processor to implement the data transmission method described in the second aspect above.

[0025] Fifthly, embodiments of this disclosure also provide a computer-readable storage medium storing one or more programs that can be executed by one or more processors to implement the data transmission method described in the second aspect above.

[0026] The multi-core integrated processor architecture, method, processor, and electronic device disclosed herein include an input / output core and a computing core. The computing core includes a core hub and multiple block processing clusters. Each block processing cluster is connected to the core hub through multiple cluster connection links. The core hub is connected to the input / output core through multiple parallel core-to-core links. This expands the transmission channel between the computing core and the input / output core from a single path to parallel multiple paths, allowing any processing cluster to independently and simultaneously drive multiple core-to-core links. Therefore, by optimizing the internal architecture of the processor, the embodiments of this disclosure can effectively expand the cross-core transmission bandwidth, reduce the cross-core transmission latency, effectively solve the data transmission bottleneck problem, and improve the throughput and transmission efficiency of the processor architecture.

[0027] Other features and advantages of this disclosure will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the disclosure. The objectives and other advantages of this disclosure may be realized and obtained by means of the structures particularly pointed out in the description, claims and drawings. Attached Figure Description

[0028] The accompanying drawings are provided to further understand the technical solutions of this disclosure and constitute a part of the specification. They are used together with the embodiments of this disclosure to explain the technical solutions of this disclosure and do not constitute a limitation on the technical solutions of this disclosure.

[0029] Figure 1 This is a schematic diagram of the multi-chip processor architecture provided in the embodiments of this disclosure;

[0030] Figure 2 This is a schematic diagram of the internal connections of the core hub provided in an embodiment of this disclosure;

[0031] Figure 3 This is a schematic diagram of the internal structure connection of the computing chip provided in the embodiments of this disclosure;

[0032] Figure 4This is a schematic diagram of the transmission of multiple types of communication signals provided in the embodiments of this disclosure;

[0033] Figure 5 This is a schematic flowchart of an optional data transmission method provided in an embodiment of this disclosure;

[0034] Figure 6 This is a schematic diagram of an optional structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0035] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the scope of this disclosure.

[0036] In this disclosure, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0037] To facilitate understanding of the technical solutions provided in the embodiments of this disclosure, some key terms used in the embodiments of this disclosure will be explained below:

[0038] A multi-chip integrated processor architecture is a processor system architecture in which multiple independently manufactured and functional semiconductor chips (including at least one computing chip and one input / output chip) are integrated into the same package through high-speed chip-to-chip interconnect technology. The chips work together to form a complete processor function.

[0039] An input / output die (IO-die) is a dedicated die in a multi-die integrated processor architecture. It is responsible for data exchange between the processor and external devices or systems, such as handling external network communication, memory interface management, and data transfer with other processors or accelerators, or vice versa, transferring external data to the computing die. It serves as a bridge between the processor architecture and the outside world.

[0040] A compute die is a compute unit in a multi-die integrated processor architecture, integrating compute resources. A compute die typically contains a die hub and multiple block processing clusters for performing compute tasks and communicating with input / output die dies across die dies via die-to-die links.

[0041] The Die Hub is a central interconnect and routing node located within the computing die. It connects various block processing clusters through cluster connection links and connects die to die links through link ports. It is responsible for data distribution, traffic scheduling, and protocol conversion between block processing clusters and between block processing clusters and input / output die.

[0042] In current multi-core integrated processor architectures, computing cores and input / output cores typically use only a single core-to-core link. However, the bandwidth of a single core-to-core link is not aligned with the bandwidth required within the computing core, resulting in limited cross-core communication bandwidth. The single-link structure cannot provide sufficient transmission capacity, causing link congestion and data pushing, which in turn affects data transmission efficiency.

[0043] For example, during the training of a large language model, the computational core contains multiple block processing clusters that perform matrix operations in parallel. Each block processing cluster outputs a large amount of feature data to the core hub through the cluster link. However, due to the limited bandwidth of the core-to-core link between the core hub and the input / output cores, it cannot carry all the transmission requirements, causing data to accumulate at the core hub entrance, forcing the computation task to be delayed and affecting the processor's throughput.

[0044] Based on this, embodiments of this disclosure propose a multi-core integrated processor architecture, method, processor, and electronic device. The multi-core integrated processor architecture includes input / output cores and computing cores. Each computing core includes a core hub and at least two block processing clusters. Each block processing cluster is connected to the core hub via an intra-core link. Each core hub is connected to the input / output cores via multiple parallel core links. This allows any processing cluster to independently and simultaneously drive multiple core-to-core links. Therefore, by optimizing the processor's internal architecture, embodiments of this disclosure can effectively expand cross-core transmission bandwidth, reduce cross-core transmission latency, effectively solve data transmission bottleneck problems, and improve the throughput and transmission efficiency of the processor architecture.

[0045] Firstly, referring to Figure 1 , Figure 1This is a schematic diagram of a multi-core processor architecture provided in this disclosure. As shown in the figure, the multi-core processor architecture includes input / output cores and computing cores. The input / output cores are responsible for handling data exchange with external devices, such as connecting to external networks or storage devices through external switching ports. The computing cores are used to perform complex computing tasks. The physical interface resources of the input / output cores are fixed and limited. For example, if an input / output core provides 8 physical interfaces, the total number of core-to-core links that can be established by that input / output core is 8. In related technologies, a single-link connection topology is used, meaning each computing core only occupies one physical interface of the input / output core to establish a single link. In this case, one input / output core can connect to a maximum of 8 computing cores, and each computing core obtains 1 / 8 of the total bandwidth, resulting in low interface resource utilization. In contrast, the solution proposed in this disclosure establishes multiple parallel links between each computing core and the input / output cores, such as... Figure 1 As shown, two parallel core-to-core links are established, increasing the number of interfaces occupied by a single computing core to two. The interface resources are upgraded from "quantity partitioning" to "bandwidth pooling". At this time, the eight interfaces of the input and output cores are no longer allocated in a static way of "one interface per core", but are reorganized in the mode of "one core with multiple interfaces". The bandwidth obtained by the input and output cores is increased to twice that of the relevant technical solutions. Without expanding the interface resources of the input and output cores, the high bandwidth transmission requirements of a single core can be met.

[0046] Understandably, these particle-to-particle links are configured as high-speed data transmission channels between compute particles and input / output particles. By deploying multiple parallel links, the total communication bandwidth between compute particles and input / output particles can be effectively increased. These parallel links can be used as independent physical channels, each with independent send and receive endpoints; or, these parallel links can share some physical resources but remain logically independent data transmission paths. Therefore, a single block processing cluster can simultaneously drive multiple particle-to-particle links for data transmission, thereby optimizing the efficiency of cross-particle data transmission. Specifically, when a single block processing cluster needs to send a large amount of data to input / output particles, the particle hub can coordinate the splitting of the data and distribute it simultaneously through multiple parallel particle-to-particle links. That is, the particle hub can include a data distribution module that can split data from the block processing cluster into multiple sub-data streams and assign each sub-data stream to a different particle-to-particle link. Alternatively, the particle hub can include multiple independent transmission controllers, each managing one particle-to-particle link and receiving data from the block processing cluster for transmission.

[0047] In a specific example, a computing chip in a multi-chip integrated processor architecture needs to transfer a large amount of computation results generated by a block processing cluster within it to an input / output chip so that the input / output chip can write these results to external memory. In related technologies, since only a single chip-to-chip link is typically deployed between the computing chip and the input / output chip, this single link becomes a bottleneck when the block processing cluster generates a high-throughput data flow, resulting in low data transmission efficiency and even link congestion and data backlog. The multi-chip integrated processor architecture provided in this embodiment includes an input / output chip and a computing chip, which contains a chip hub and multiple block processing clusters. When one of the block processing clusters (e.g., block processing cluster A) completes its computation task and generates a large amount of data, block processing cluster A sends the data to the chip hub through its cluster connection link. After receiving the data from block processing cluster A, the chip hub recognizes that the data needs to be transferred to the input / output chip. Since there are multiple parallel chip-to-chip links connecting the chip hub and the input / output chip, the chip hub initiates a data transmission process. Specifically, the core node splits the raw data from block processing cluster A into multiple smaller sub-data packets. Then, the core node distributes these sub-data packets in parallel to the multiple core-to-core links it is connected to. For example, if there are four core-to-core links, the core node can split the data into four parts and transmit them simultaneously through all four links.

[0048] Therefore, multiple kernel-to-kernel links are driven simultaneously by a single block processing cluster, transmitting data in parallel from the computation kernel to the input / output kernel. Upon receiving these parallelly transmitted sub-data packets, the input / output kernel reassembles them and processes them in the original data order, such as writing them to external memory. In this way, the high-throughput data stream generated by the single block processing cluster can fully utilize the bandwidth of multiple parallel links, thereby improving the overall transmission capacity of cross-kernel communication, avoiding the bottleneck effect of a single link, and effectively solving the problems of limited data transmission, link congestion, and data backlog.

[0049] In one possible implementation, such as Figure 1As shown, the computing core comprises a core hub and multiple block processing clusters. The core hub acts as the communication center within the computing core, coordinating data flow. Block processing clusters are the units that execute computational tasks; multiple block processing clusters enable parallel processing within the computing core, thereby improving overall computing power. For example, the core hub can be a high-speed router, while each block processing cluster can contain multiple computing units and independent computing modules with local caches. Specifically, each block processing cluster is connected to the core hub via cluster connection links. These cluster connection links serve as data transmission paths between the block processing clusters and the core hub. Through these cluster connection links, the block processing clusters can efficiently transmit data to the core hub. These cluster connection links can consist of a set of parallel data lines.

[0050] Within a single kernel hub, cluster connection links from different block processing clusters are interconnected to form multiple inter-cluster buses. These inter-cluster buses enable direct data exchange between block processing clusters within a computing kernel, eliminating the need to transmit data to and from input / output kernels and back, thus optimizing internal communication efficiency. Specifically, cluster connection links can be connected within the kernel hub via crossbar switches or multiplexers to form inter-cluster buses. Additionally, the kernel hub includes multiple link ports, each corresponding to a kernel link. These link ports serve as the physical interface between the kernel hub and external kernel-to-kernel links, with each link port dedicated to data transmission and reception for a specific kernel link. Furthermore, each cluster connection link within the kernel hub connects to both the inter-cluster bus and a link port, forming a dual-branch structure. Data can be routed to the inter-cluster bus for communication between internal block processing clusters, or it can be routed to a link port and then transmitted to the input / output kernel via a kernel-to-kernel link.

[0051] It is worth noting that the total bandwidth of all cluster connection links connected to a single block processing cluster is equal to the total bandwidth of multiple kernel-to-kernel links. When a block processing cluster can simultaneously drive multiple links with bandwidth alignment, the dynamic reorganization of interface resources can achieve an end-to-end bandwidth multiplication effect. Related technologies that only achieve physical multiplexing of the number of interfaces cannot solve the technical challenges of bandwidth alignment and single-kernel exclusive transmission. This disclosure ensures that when the block processing cluster within a computation kernel generates high-throughput data streams, the output bandwidth of the cluster connection links can match the transmission capacity of external kernel-to-kernel links. This effectively prevents data accumulation at the kernel hub and avoids link congestion. For example, if the kernel hub is connected to four kernel-to-kernel links, each with a bandwidth of X, the total bandwidth is 4X. In this case, the total bandwidth of the cluster connection links connected to a single block processing cluster is also 4X. Therefore, when a single processing cluster sends signals, it can fully utilize the bandwidth of the cluster connection links and the total bandwidth of the four kernel-to-kernel links, i.e., independently drive all external kernel-to-kernel links, effectively improving bandwidth utilization.

[0052] The following is a more concrete example illustrating this. A multi-core integrated processor architecture comprises one input / output core and multiple compute cores. Each compute core contains a core hub and four block processing clusters, labeled Block Processing Cluster A, Block Processing Cluster B, Block Processing Cluster C, and Block Processing Cluster D. Each block processing cluster is connected to the core hub via two cluster connection links. Simultaneously, the core hub is connected to the input / output core via two parallel core-to-core links, each with the same bandwidth. In a computing scenario, block processing cluster A needs to send a large amount of data to the input / output core, such as a high-resolution video stream or intermediate results from a large-scale machine learning model. In a traditional architecture, if there is only one core-to-core link between the compute core and the input / output core, the high-throughput data stream generated by block processing cluster A will quickly exceed the transmission capacity of this single link, causing data to accumulate at the core hub, resulting in link congestion and data transmission delays. In the architecture proposed in this disclosure, when block processing cluster A generates a high-throughput data stream, this data is first transmitted to the kernel hub via two connected cluster connection links. Since the total bandwidth of all cluster connection links connected to a single block processing cluster is designed to be equal to the total bandwidth of the two (i.e., all) kernel-to-kernel links, block processing cluster A can push data to the kernel hub at its maximum output rate. Inside the kernel hub, this data from block processing cluster A is routed to multiple link ports corresponding one-to-one with the two kernel-to-kernel links. Specifically, the routing logic of the kernel hub splits or distributes the data from block processing cluster A to these link ports, allowing two parallel kernel-to-kernel links to be driven simultaneously. For example, data can be divided into two sub-streams, each of which is transmitted in parallel to the input / output kernel via one link port and one kernel-to-kernel link. Therefore, a high-throughput data stream that might otherwise cause congestion on a single link can now be carried by two parallel links, effectively improving cross-kernel communication bandwidth. Simultaneously, if block processing cluster A needs to communicate with block processing cluster B within the computing core during the process of sending data to the input / output core, such as exchanging control information or intermediate calculation results, then the cluster connection link from block processing cluster A is also connected to the inter-cluster bus within the core hub. The routing logic of the core hub can guide the data to be sent to block processing cluster B to the inter-cluster bus, and then transmit it to the cluster connection link connected to block processing cluster B via the inter-cluster bus. This dual-branch approach allows cross-core communication and inter-cluster communication to proceed in parallel without interference, improving the communication efficiency and flexibility of the computing core. It is understood that the architecture of this disclosure embodiment can effectively utilize the total bandwidth provided by multiple parallel core-to-core links, ensuring that even the data throughput generated by a single block processing cluster can be efficiently transmitted to the input / output core, thereby solving the problems of limited cross-core communication bandwidth, link congestion, and data backlog in traditional architectures.

[0053] Understandably, in related technologies, only a single chip-to-chip link is typically deployed between computing chips and input / output chips. When a block processing cluster generates high-throughput data streams, the fixed bandwidth of this single link becomes a bottleneck, resulting in severely insufficient data transmission capacity, which in turn leads to link congestion and data backlog. The architecture provided in this disclosure, however, introduces multiple parallel chip-to-chip links, directly and linearly extending the communication bandwidth between computing chips and input / output chips. In the example above, the total bandwidth of the two parallel chip-to-chip links is twice that of a single link, enabling block processing cluster A to transmit data to input / output chips at a higher rate, effectively avoiding bandwidth bottlenecks. Since the cluster connection links in this embodiment are simultaneously connected to the inter-cluster bus and link ports within the chip hub, data streams can be flexibly routed as needed, enabling fast communication between block processing clusters within the computing chip and efficient data exchange with input / output chips. Compared to traditional architectures where data may require complex arbitration or queuing at the chip hub to determine the transmission path, the solution in this disclosure improves the efficiency and flexibility of data routing. It is worth noting that in this embodiment, the total bandwidth of all cluster connection links connected to a single block processing cluster is equal to the total bandwidth of multiple kernel-to-kernel links. By matching the internal output capability with the external transmission capability, a single block processing cluster can independently drive all external kernel-to-kernel links without waiting for signal patching from other block processing clusters, thereby ensuring smooth data transmission and mitigating the risk of link congestion. In summary, this embodiment, by optimizing the architecture of multi-kernel integrated processors, can effectively expand cross-kernel transmission bandwidth, matching the internal bandwidth of the block processing cluster with the external link bandwidth, reducing cross-kernel transmission latency, effectively solving the data transmission bottleneck problem, and improving the throughput and transmission efficiency of the processor architecture.

[0054] In one possible implementation, the number of block processing clusters connected to a single core node is proportional to the number of core-to-core links between the single core node and input / output cores. Specifically, the number of physical interfaces of the input / output cores is fixed, meaning the number of core-to-core links that can be connected is also fixed. For example, an input / output core has 8 physical interfaces and supports a maximum of 8 core-to-core links. In traditional architectures, each compute core occupies only 1 link, allowing the input / output core to connect to 8 compute cores. However, the architecture provided in this disclosure expands the bandwidth from the compute core to the input / output core by allocating multiple links (such as 2 or 4) to a single compute core, but this correspondingly reduces the total number of compute cores that the input / output core can connect to. For example, when each compute core occupies 2 core-to-core links, an input / output core with 8 physical interfaces can only connect to 4 compute cores. In this scenario, there is a multiple relationship between the number of block processing clusters connected to a single core node and the number of core-to-core links occupied by a single computing core. This multiple relationship reflects that, under the premise of a fixed total input and output core bandwidth, the ratio between computing resource density and communication bandwidth is coordinated through architecture planning. The number of block processing clusters is set to N times the number of core-to-core links (N is a positive integer, such as 2 times, 4 times, etc.). This allows the core node to evenly map the aggregated data stream onto the multiple parallel links it occupies according to the number of block processing clusters and the output bandwidth requirements. This ensures that the throughput of each block processing cluster can obtain matching transmission capacity across cores to suit different application scenarios or load requirements.

[0055] The following is a more concrete example. Assume an input / output kernel has 8 physical interfaces, supporting a maximum of 8 kernel-to-kernel links. Each compute kernel occupies 2 kernel-to-kernel links, meaning this input / output kernel can connect to 4 compute kernels. In this case, the kernel hub of a single compute kernel can connect to 8 block processing clusters. The number of block processing clusters connected to a single kernel hub (8) is 4 times the number of kernel-to-kernel links occupied by the kernel hub (2). When these 8 block processing clusters generate data streams simultaneously, the kernel hub can use priority arbitration to distribute the data generated by these 8 block processing clusters. Data is sequentially transmitted to two parallel particle-to-particle links; alternatively, each compute particle can occupy four particle-to-particle links, in which case the input / output particles connect only two compute particles, and the particle hub can connect eight block processing clusters. In this case, the number of block processing clusters is twice the number of links. However, it is worth noting that the total bandwidth of the links connecting a single particle hub to a single block processing cluster is still equal to the total bandwidth of the links connecting the particle hub to the input / output particles. This allows for flexible adjustment between the number of compute particles and the bandwidth of a single particle while ensuring that the total bandwidth of the particle-to-particle links is fully utilized.

[0056] In one possible implementation, refer to Figure 2 , Figure 2 This is a schematic diagram of the internal connections of a core node provided in an embodiment of this disclosure. The core node includes multiple cluster ports and multiple link ports. Cluster ports are used to connect to block processing clusters of computing cores via cluster connection links, while multiple link ports are used to connect to the same input / output core via parallel core-to-core links. Cluster ports connected to different block processing clusters are interconnected, and each link port is connected to each cluster port. The total bandwidth of all cluster ports connected to the same block processing cluster is equal to the total bandwidth of all link ports. By optimizing the structure of the core node, expanding multiple link ports within the core node, and then connecting the multiple cluster ports to the corresponding link ports, the internal bandwidth matches the external bandwidth, thereby effectively solving the link congestion problem caused by limited cross-core communication bandwidth, improving data transmission efficiency, and avoiding data accumulation.

[0057] Specifically, multiple cluster ports allow block processing clusters within the computing core to access the core hub in parallel via cluster connection links, ensuring sufficient internal bandwidth. Multiple link ports establish connections to the same input / output core via parallel core-to-core links, expanding the total bandwidth capacity of external communication channels and providing ample transmission capacity for high-throughput data streams. Cluster ports connected to different block processing clusters interconnect to form an internal communication network, enabling data exchange between block processing clusters to be completed directly within the computing core without passing through input / output cores, thereby reducing dependence on external links and alleviating cross-core traffic pressure. Link ports connect to individual cluster ports, ensuring each cluster port can independently access external links, improving data routing flexibility and transmission efficiency. Crucially, the total bandwidth of all cluster ports connected to the same processing cluster equals the total bandwidth of all link ports, meaning internal data output capacity matches external link transmission capacity, preventing congestion points during data transmission.

[0058] In a specific example, when a block processing cluster generates a high-throughput data stream, the data signal is transmitted to the core node structure through multiple cluster ports connected to that block processing cluster. Since the total bandwidth of the connected cluster ports is equal to the total bandwidth of the link ports, the data signal can be completely distributed to each link port and synchronously transmitted to the input and output cores through parallel core-to-core links. That is, the data signal initiated by a single block processing cluster can fill the total bandwidth of the link ports, thus transmitting it in one go without waiting for signal piecing together. During this process, if another processing cluster needs to communicate internally with other block processing clusters, the data can be directly routed to the target cluster port through the interconnected cluster ports without occupying link port resources, thereby achieving parallel processing of cross-core and inter-cluster communication. Therefore, the computing core can fully utilize the total bandwidth of the parallel core-to-core links, ensuring the smooth transmission of high-throughput data streams and completely avoiding data backlog.

[0059] In one possible implementation, the cluster connection links connecting each block processing cluster include a first cluster connection link and a second cluster connection link. The chip-to-chip link between the chip hub and the input / output chips includes a first chip-to-chip link and a second chip-to-chip link. Multiple link ports include a first link port and a second link port. A first branch of the first cluster connection link is connected to a first inter-cluster bus, and a second branch is connected to a first link port. Similarly, a first branch of the second cluster connection link is connected to a second inter-cluster bus, and a second branch is connected to a second link port. The first link port is connected to a first chip-to-chip link, and the second link port is connected to a second chip-to-chip link. The multiple cluster connection links subdivide the communication path between the block processing cluster and the chip hub, for example, as... Figure 2 As shown, two independent sets of parallel data lines can be used, or two independent logical channels can be divided on a high-speed serial link. Each block processing cluster is connected to cluster port 0 and cluster port 1 of the kernel hub through two cluster connection links. The inter-cluster bus within the kernel hub can interconnect the cluster ports connected to different block processing clusters. For example, inter-cluster bus 0 connects cluster port 0 of block processing cluster A to cluster port 0 of block processing cluster B within the kernel hub, while cluster port 1 of block processing cluster A and cluster port 1 of block processing cluster B are interconnected within the kernel hub to form inter-cluster bus 1. Multiple kernel-to-kernel links correspond one-to-one with the link ports of the kernel hub, with each link port connecting to one kernel-to-kernel link, such as... Figure 2As shown, link port 0 is connected to core-to-core link 0, and link port 1 is connected to core-to-core link 1. That is, the first link port is connected to the first core-to-core link, and the second link port is connected to the second core-to-core link, establishing a direct correspondence between the internal link ports of the core hub and the external core-to-core links. This ensures that data can be smoothly transmitted from the core hub to the input / output cores. Specifically, the first branch of the first cluster connection link is connected to the first inter-cluster bus, and the second branch is connected to the first link port. This branching structure allows the first cluster connection link to serve both inter-cluster communication within the core and inter-core communication. For example, a routing switch or multiplexer can be set up inside the core hub to direct data from the first cluster connection link to the first inter-cluster bus or the first link port based on the destination address or type of the data packet. The first branch of the second cluster connection link is connected to the second inter-cluster bus, and the second branch is connected to the second link port. This is similar to the branching structure of the first cluster connection link, expanding parallel processing capabilities and the flexibility of communication paths.

[0060] Understandably, multiple parallel and functionally separated data transmission paths are constructed by branching the cluster connection links between block processing clusters and the kernel hub, as well as the kernel-to-kernel links between the kernel hub and input / output kernels. Specifically, when a block processing cluster generates data, this data can be transmitted to the kernel hub in parallel via the first and second cluster connection links. Inside the kernel hub, these cluster connection links are designed with a branching structure: one branch (e.g., the first branch of the first cluster connection link and the first branch of the second cluster connection link) connects to the inter-cluster bus to handle communication between different block processing clusters within the same kernel hub; another branch (e.g., the second branch of the first cluster connection link and the second branch of the second cluster connection link) connects to dedicated link ports (the first link port and the second link port), which in turn connect to kernel-to-kernel links (the first kernel-to-kernel link and the second kernel-to-kernel link) outside the kernel hub for communication with input / output kernels. Therefore, the structure of the core node hub can simultaneously handle internal and external communication requests from the block processing cluster, and can divert different types of communication traffic to different physical or logical paths. For example, cross-core communication signals can be transmitted in parallel through the first cluster connection link and the first core-to-core link, as well as the second cluster connection link and the second core-to-core link. Cross-cluster communication signals can be transmitted in parallel through the first cluster connection link, the first cluster port and the inter-cluster bus, as well as the second cluster connection link, the second cluster port and the inter-cluster bus. The parallel and diversion mechanism effectively avoids a single communication path becoming a bottleneck, thereby making full use of the total bandwidth of multiple parallel core-to-core links between the core node hub and the input / output cores, as well as the total bandwidth of all cluster connection links connected to a single block processing cluster, thereby improving data transmission efficiency.

[0061] In one possible implementation, refer to Figure 3 , Figure 3 This is a schematic diagram of the internal structure of the computing chip provided in the embodiments of this disclosure. Each block processing cluster includes a cluster hub, an intra-cluster shared memory, and multiple computing units. The intra-cluster shared memory and multiple computing units are respectively connected to the cluster hub. The cluster hub is connected to the chip hub through multiple cluster connection links, so that the cluster hub can distribute data from multiple computing units to the multiple cluster connection links it is connected to for parallel transmission.

[0062] In one possible implementation, the bandwidth of the transmission link between the intra-cluster shared memory and the cluster hub is equal to the total bandwidth of all cluster connection links connected to a single cluster hub.

[0063] The cluster hub, acting as the central node for data flow within the block processing cluster, manages and routes data transmission between the internal computing units and the intra-cluster shared memory, and coordinates communication with external core hubs. It contains multiple input and output ports, enabling arbitrary input and output. The intra-cluster shared memory provides high-speed, low-latency data access to the computing units within the block processing cluster. This internal storage resource allows multiple computing units to share access, supporting parallel computing and data exchange. The intra-cluster shared memory can be a dynamic random access memory (DRAM) or an L1 / L2 cache integrated within the block processing cluster. The multiple computing units are the core processing units that execute actual computational tasks, responsible for executing instructions, processing data, and generating results. These units can include general-purpose CPU cores, dedicated GPU cores, digital signal processors, or accelerator units designed for specific applications (such as artificial intelligence inference). The intra-cluster shared memory and the multiple computing units are connected to the cluster hub, allowing all data flows within the block processing cluster to be centrally managed and scheduled through the cluster hub, achieving efficient internal communication. Cluster hubs are connected to kernel hubs via multiple cluster connection links, providing block processing clusters with high-bandwidth, parallel data transmission channels to support high-throughput data exchange. The bandwidth of the transmission link between the intra-cluster shared memory and the cluster hub is equal to the total bandwidth of all cluster connection links connected to a single cluster hub. Therefore, the data transmission capacity within the block processing cluster matches the transmission capacity of the external cluster connection links, avoiding internal bottlenecks. Specifically, this can be achieved by designing the interface width and operating frequency of the intra-cluster shared memory, as well as the routing and switching capabilities within the cluster hub, to ensure that the total bandwidth remains consistent with the aggregate bandwidth of all cluster connection links.

[0064] Understandably, the block processing cluster uses a cluster hub as the core intermediary node to centrally manage internal data flow. Shared memory and multiple computing units within the cluster are directly connected to the cluster hub, which in turn connects to the core hub via multiple cluster connection links. This supports parallel data transmission and fully utilizes external bandwidth resources. Data from multiple computing units is effectively distributed across the multiple cluster connection links they are connected to, enabling parallel transmission and maximizing the bandwidth of the cluster connection links. This avoids transmission bottlenecks. For example, when multiple computing units simultaneously generate large amounts of data that need to be transmitted to the core hub, the cluster hub can distribute this data across multiple cluster connection links, allowing the block processing cluster to fully utilize its internal computing resources and external parallel transmission channels, improving data processing efficiency and overall performance. Furthermore, the bandwidth of the transmission link between the shared memory within the cluster and the cluster hub is equal to the total bandwidth of all cluster connection links connected to a single cluster hub, ensuring consistency between internal link capacity and the maximum external transmission capacity. Therefore, when computing units generate high-throughput data, internal transmission will not become a bottleneck, effectively eliminating the risk of congestion. Furthermore, the bandwidth required for transmitting communication signals generated by a single computing unit can be less than or equal to the total bandwidth of all cluster connection links connected to a single cluster hub. This means that the communication signals generated by a single computing unit can be transmitted to the core hub in one go through the cluster connection links. For example, the communication signals generated by a single computing unit can directly fill both the cluster connection links and the core-to-core links, or the communication signals generated by multiple computing units can be combined to fill the gap. In this way, the block processing cluster can efficiently process and transmit data, fully utilizing the high-bandwidth cluster connection links provided by the core hub, enabling the entire multi-core integrated processor architecture to achieve higher data throughput and processing efficiency.

[0065] In one possible implementation, data generated by a single computing unit is distributed via a cluster hub to multiple cluster connection links connected to the cluster hub for parallel transmission. Specifically, the signal generated by a single computing unit is transmitted in parallel across all cluster connection links connected to a single cluster hub; that is, the transmission rate of the signal generated by a single computing unit can be equal to the total bandwidth of all cluster connection links connected to a single cluster hub, or the transmission rate of the signal generated by a single computing unit can be less than the total bandwidth of all cluster connection links connected to a single cluster hub.

[0066] Specifically, signals generated by a computing unit refer to data streams or control signals generated by a specific computing unit within a block processing cluster (e.g., a processing core, a dedicated accelerator, or a functional module). These signals can be computation results, intermediate data, memory access requests, or status updates. Instead of being confined to a single cluster connection link, these signals are broken down or duplicated and distributed to all cluster connection links connected to the cluster hub for simultaneous transmission. For example, a high-bandwidth data stream can be split into multiple sub-streams, each transmitted through an independent cluster connection link; or, for control signals that need to be broadcast, multiple copies can be made and sent simultaneously through multiple cluster connection links. This parallel transmission mechanism fully utilizes all available transmission resources, preventing a single link from becoming a bottleneck. The transmission rate of a signal generated by a single computing unit can be less than or equal to the total bandwidth of all cluster connection links connected to a single cluster hub. This means that the output capacity of the computing unit is precisely matched with the carrying capacity of the cluster connection links. If the transmission rate of a signal generated by a single computing unit is equal to the total bandwidth of all cluster connection links connected to a single cluster hub, then the data stream generated by a single computing unit can fill the cluster connection links from its block processing cluster to the core hub without waiting for data streams from other computing units to be pieced together. This minimizes transmission latency and improves the real-time performance of data transmission. For example, in an artificial intelligence inference scenario, when a dedicated accelerator unit within a block processing cluster generates a high-resolution feature map data stream, this data stream can be directly transmitted in parallel to the core hub through all cluster connection links. This avoids the accumulation of waiting time caused by segmented transmission, ensures low-latency execution of inference tasks, effectively improves the resource utilization of cluster connection links, avoids waste caused by idle link bandwidth, and allows the output capacity of each computing unit to be fully released.

[0067] It is understandable that the problem of low transmission efficiency of high-bandwidth signals from a single computing unit is solved by transmitting the signal generated by a single computing unit in parallel across all cluster connection links connected to a single cluster hub, and ensuring that the transmission rate of the signal is less than or equal to the total bandwidth of all cluster connection links. In the multi-core integrated processor architecture described above, the cluster hub is connected to the core hub through multiple cluster connection links, and the transmission link bandwidth between the shared memory within the cluster and the cluster hub is equal to the total bandwidth of all cluster connection links connected to a single cluster hub. Therefore, the high-bandwidth data stream generated by a single computing unit can be effectively distributed across multiple cluster connection links through parallel transmission, avoiding data accumulation at the cluster hub. At the same time, the output of a single computing unit can fully utilize the bandwidth of the cluster connection links, preventing link idleness or data overflow. This allows even a single computing unit to fully utilize the high-bandwidth transmission capability within the block processing cluster, thereby improving the data throughput efficiency and response speed of the entire processor architecture.

[0068] In one possible implementation, a cluster hub can aggregate data generated by multiple computing units and distribute the aggregated data to the cluster connection links for parallel transmission at a rate equal to the total bandwidth of all cluster connection links connected to the cluster hub. That is, signals generated by multiple computing units are aggregated within the block processing cluster and then distributed to multiple cluster connection links connected to the cluster hub for parallel transmission, with the transmission rate of the aggregated signals equal to the total bandwidth of all cluster connection links connected to a single cluster hub. Specifically, the signals generated by multiple computing units can refer to data streams, control information, or status updates generated after multiple computing units within the block processing cluster execute computational tasks in parallel. These signals typically have high throughput and real-time requirements, such as intermediate results or final outputs generated in artificial intelligence inference or high-performance computing. These signals can exist in the form of data packets, data blocks, or continuous data streams.

[0069] In this context, aggregation within a block processing cluster refers to the process by which the cluster hub integrates, merges, or packages multiple signals from different computing units. This aggregation can be implemented in various ways. For example, the cluster hub may include a dedicated aggregation logic unit that merges multiple small data packets into a large data packet, or a shared buffer may be set up where multiple computing units write their respective signals. The buffer controller is then responsible for reading these signals sequentially or by priority and integrating them into a unified data stream.

[0070] Distributing aggregated data at a rate equal to the total bandwidth of all cluster connection links connected to the cluster hub can mean that the data transmission output rate matches the sum of the theoretical maximum transmission capacity of all cluster connection links connected to the cluster hub. For example, the transmission rate of the aggregated signal equals the total bandwidth of all cluster connection links connected to a single cluster hub. This means that the aggregation logic and distribution mechanism are designed to fully utilize the transmission capacity of all available cluster connection links, and the aggregated data output rate matches the total transmission capacity of all parallel cluster connection links, thereby avoiding link congestion or bandwidth waste. For instance, the aggregation module can dynamically adjust the aggregation rate or packet size based on the total bandwidth of the cluster connection links to ensure that the data stream can pass through the cluster connection links with maximum efficiency; that is, the signal aggregated and pieced together by multiple computing units can fill the cluster connection links connected to that processing cluster.

[0071] The following concrete example illustrates this. In a block processing cluster, assume four computing units execute tasks in parallel, each generating a data stream with a bandwidth of X. To efficiently transmit this data to the core node, an aggregation controller is implemented within the block processing cluster. This aggregation controller receives data signal streams from any two of the four computing units and integrates them into a single, high-bandwidth data signal stream with a bandwidth of 2X. This aggregated data stream is then sent to the cluster node. The cluster node is connected by two cluster connection links, each with a bandwidth of X. The distribution logic within the cluster node splits the aggregated 2X bandwidth data stream into two X bandwidth sub-data streams, which are then transmitted in parallel to the core node via the two cluster connection links. In this way, the 2X bandwidth of the aggregated signal matches the total bandwidth of the eight cluster connection links, 2X, ensuring maximum data transmission efficiency. Therefore, when multiple computing units generate high-throughput signals simultaneously, by aggregating the signals of multiple computing units within the block processing cluster and using multiple cluster connection links for parallel transmission, the aggregated signals can fill the total bandwidth of the cluster connection links, effectively improving data transmission efficiency, avoiding data accumulation and link congestion, thereby improving the high performance and high throughput of the entire multi-core integrated processor architecture.

[0072] In one possible implementation, the number of block processing clusters connected to a single core hub is proportional to the number of core-to-core links between the single core hub and input / output cores. The number of block processing clusters can be 2 or 4 times the number of core-to-core links, or the ratio can be dynamically adjusted according to configuration or actual operational needs.

[0073] Understandably, when the number of block processing clusters is a multiple of the number of kernel-to-kernel links, the kernel hub can efficiently manage and schedule data flows. For example, if the number of block processing clusters is N times the number of kernel-to-kernel links, the kernel hub can be designed to aggregate the data flows of N block processing clusters onto a single kernel-to-kernel link, or to split the data flow of one block processing cluster and distribute it evenly across multiple kernel-to-kernel links, thereby achieving average resource allocation. Specifically, assuming a kernel hub connects to four block processing clusters, to ensure a balance between internal computing power and external communication bandwidth, two parallel kernel-to-kernel links can be configured between the kernel hub and the input / output kernels. In this configuration, the number of block processing clusters is twice the number of kernel-to-kernel links. At this point, the core hub can aggregate data streams from the four block processing clusters according to data transmission needs and distribute them to two core-to-core links for parallel transmission. Alternatively, each pair of block processing clusters can share one core-to-core link, or the core hub can dynamically and evenly distribute the data from all block processing clusters to all available core-to-core links based on the load. By setting the multiplier relationship, the core hub can flexibly manage internal and external data traffic and improve the efficiency of data transmission.

[0074] In one possible implementation, refer to Figure 4 , Figure 4 This is a schematic diagram of the transmission of multiple types of communication signals provided in the embodiments of this disclosure. The input / output core also includes an external switching port for connecting to an external network switching device, and a cross-core transmission port for connecting to other input / output cores. The signal types transmitted within the core hub include: cross-cluster communication signals, and / or cross-core communication signals, and / or external switching communication signals. The transmission path of the cross-cluster communication signal is that any processing cluster is transmitted to another processing cluster within the same core hub through an inter-cluster bus. The transmission path of the cross-core communication signal is that any processing cluster is transmitted to the cross-core transmission port through a link port of the core hub. The transmission path of the external switching communication signal is that any processing cluster is transmitted to the external switching port through a link port of the core hub.

[0075] The external switching port is an interface on the input / output chip used for data communication with network switching devices outside the processor. For example, it can be connected to an external Ethernet switch through a physical layer interface to exchange data with an external network at high speed.

[0076] The cross-chip transfer port is an interface on the input / output chip used for high-speed interconnection with other input / output chips. It supports extended communication between different computing chips or different processor systems in a multi-chip integrated processor architecture, enabling larger-scale parallel processing and data sharing. The cross-chip transfer port can employ D2D (Die-to-Die) interconnect technology to achieve high-bandwidth, low-latency inter-chip communication.

[0077] Cross-cluster communication signals refer to signals used for data exchange between different processing clusters within the same computing core. These signals are used for data sharing, task collaboration, and result aggregation between different processing clusters, enabling parallel processing and data flow scheduling within the computing core. For example... Figure 4 As shown, the new signal transmission path for cross-cluster communication is that any processing cluster transmits signals to another processing cluster connected within the same core hub via the inter-cluster bus.

[0078] Cross-core communication signals refer to signals transmitted from the block processing cluster within a computing core, through the core hub, and then to the cross-core transmission port of the input / output core, used for communication with other input / output cores. Cross-core communication signals are used for data exchange between the computing core and other external computing cores or processor systems. For example... Figure 4 As shown, the cross-core communication signal transmission path is that any processing cluster transmits to the cross-core transmission port through the link port of the core hub.

[0079] External exchange communication signals refer to signals transmitted from the block processing clusters within the computing core, through the core hub, and then to the external exchange ports of the input / output cores, used for communication with external network switching devices. External exchange signals are the signals used by the computing core to exchange data with the external network, allowing the processor to participate as a network node in large-scale distributed computing or data center applications. For example... Figure 4 As shown, the transmission path of the external exchange communication signal is that any processing cluster transmits to the external exchange port through the link port of the core hub.

[0080] Understandably, when a block processing cluster needs to communicate with another block processing cluster within the same kernel hub, the resulting cross-cluster communication signal can be transmitted to the kernel hub via the cluster connection link. Within the kernel hub, it is directly routed to the target block processing cluster's cluster connection link via the inter-cluster bus. Therefore, internal routing avoids cross-cluster communication consuming kernel-to-kernel link resources, thereby reducing the load on inter-kernel communication. When a block processing cluster needs to communicate with other external input / output kernels, the resulting cross-kernel communication signal is transmitted to the kernel hub via the cluster connection link. Then, through the kernel hub's link port, it is transmitted to the input / output kernel via the kernel-to-kernel link, and finally sent out by the input / output kernel's cross-kernel transmission port. When a block processing cluster needs to communicate with an external network switching device, the resulting external switching communication signal is transmitted to the kernel hub via the cluster connection link. Then, through the kernel hub's link port, it is transmitted to the input / output kernel via the kernel-to-kernel link, and finally sent out by the input / output kernel's external switching port. In this way, the scheme disclosed herein effectively distinguishes and optimizes paths for different communication types (cross-cluster, cross-core, and external switching). The core node hub, as the core scheduling unit, can guide the signal to the most suitable transmission path based on the signal type, thereby avoiding resource contention between different types of data streams and ensuring high throughput and low latency communication. It is worth noting that when sending cross-cluster communication signals, cross-core communication signals, and external switching communication signals, the cluster connection links of the processing cluster can be fully utilized, or only a portion of the cluster connection links can be used. For example, when the same processing cluster needs to send cross-cluster communication signals and cross-core communication signals simultaneously, the cross-cluster communication signal can be transmitted to another processing cluster through the corresponding cluster connection port and inter-cluster bus via the first cluster connection link, while the cross-core communication signal can be transmitted to the input / output core via the corresponding cluster connection port and link port via the second cluster connection link. In addition, the core node hub can contain a priority arbitrator. When multiple parallel transmission channels compete for bandwidth resources, the arbitrator can schedule the transmission order of cross-cluster communication signals, cross-core communication signals and external exchange communication signals according to a preset priority order, effectively improving communication efficiency and system responsiveness.

[0081] The following is a concrete example to illustrate this. Assume a multi-core integrated processor architecture includes one input / output core and two compute cores. Each compute core contains a core hub and four block processing clusters. The input / output core is configured with an external switching port for connecting to an Ethernet switch in the data center, and a cross-core transport port for connecting to the input / output core of another processor system. When block processing cluster A in a compute core needs to send data to block processing cluster B in the same compute core, for example, for data synchronization or result aggregation, the cross-cluster communication signal generated by block processing cluster A enters the core hub through the cluster connection link. The core hub recognizes this as cross-cluster communication and routes the signal to the inter-cluster bus within the core hub, forwarding it directly to the cluster connection link of block processing cluster B. The entire process is completed within the compute core, without consuming the bandwidth of the core-to-core link connecting the input / output cores. When block processing cluster C needs to send data to an input / output core connected to another processor system, for example, for data exchange in distributed computing, the cross-core communication signal generated by block processing cluster C enters the core hub through its cluster connection link. The core node recognizes this as inter-core communication and transmits the signal through its link port via a core-to-core link to the input / output core of the local processor system. The input / output core then sends the data to the input / output core of the target processor system via the inter-core transmission port. When block processing cluster D needs to upload computation results to an external data center or receive data from an external network, such as for model training or data analysis, the external switching communication signal generated by block processing cluster D enters the core node through its cluster connection link. The core node recognizes this as external switching communication and transmits the signal through its link port via a core-to-core link to the input / output core of the local processor system. The input / output core then sends the data to an external Ethernet switching device via its external switching port. Therefore, within a multi-core integrated processor architecture, multiple different types of communication signals can be transmitted simultaneously. In this process, the priority arbiter inside the core node hub can be configured, for example, to prioritize cross-core communication signals to ensure the real-time requirements of distributed computing, then process external exchange communication signals, and finally process cross-cluster communication signals, thereby intelligently managing data flow and avoiding delays in critical tasks when bandwidth resources are limited.

[0082] Secondly, based on the aforementioned multi-chip integrated processor architecture, this disclosure proposes a data transmission method applicable to multi-chip integrated processor architectures. This data transmission method can be applied to, for example... Figure 1 or Figure 2 or Figure 3 or Figure 4 The multi-chip integrated processor architecture shown is referenced. Figure 5 , Figure 5This is an optional flowchart illustrating a data transmission method provided in an embodiment of the present disclosure. The data transmission method includes, but is not limited to, the following steps 501 or 502.

[0083] Step 501: In response to the block processing cluster sending a cross-core communication signal, the cross-core communication signal is split into multiple cross-core communication sub-signals and sent in parallel to multiple link ports through multiple cluster connection links, so that multiple core-to-core links are driven to transmit simultaneously.

[0084] Step 502: In response to the block processing cluster sending an inter-cluster communication signal, the inter-cluster communication signal is split into multiple inter-cluster communication sub-signals and sent in parallel to multiple inter-cluster buses through multiple cluster connection links to transmit to the cluster connection links of other block processing clusters.

[0085] It is understandable that, in step 501, responding to the block processing cluster sending a cross-chip communication signal means that when any block processing cluster in a multi-chip integrated processor architecture needs to exchange data with input / output chips, a cross-chip communication request is generated, and a corresponding cross-chip communication signal is produced. Then, the cross-chip communication signal is decomposed into several smaller, independently transmittable cross-chip communication sub-signals. This decomposition can be achieved in various ways; for example, the original data packet can be segmented to form multiple sub-data packets; or, a continuous data stream can be divided into multiple parallel data sub-streams. The purpose of this decomposition is to distribute data that originally needed to be transmitted through a single path (and could not be transmitted in one go) across multiple parallel paths, thereby achieving one-time transmission to improve transmission efficiency and utilization. The split cross-core communication sub-signals are no longer transmitted through a single cluster connection link, but are instead transmitted in parallel using multiple cluster connection links within the computing core. These cluster connection links are each connected to different link ports within the core hub, allowing each sub-signal to enter the core hub through its corresponding link port. This fully utilizes the parallel communication capabilities within the processor, improving data transmission throughput and enabling multiple core-to-core links to be driven for transmission simultaneously. In other words, multiple link ports can simultaneously receive cross-core communication sub-signals, thereby simultaneously activating multiple core-to-core links connected to these link ports. This means that core-to-core link resources that might have been used sequentially or partially can now be used synchronously and fully, thereby maximizing the communication bandwidth between the computing core and the input / output cores.

[0086] Understandably, by splitting the cross-granular communication signals generated by the block processing cluster and transmitting them in parallel to multiple link ports of the granular hub using multiple cluster connection links, multiple granular-to-granular links are simultaneously driven for transmission, thus efficiently solving the bandwidth bottleneck problem in cross-granular communication. Specifically, when the block processing cluster needs to send a cross-granular communication signal, the signal is decomposed into multiple cross-granular communication sub-signals. These sub-signals are then transmitted in parallel to multiple link ports within the granular hub by calculating multiple pre-defined cluster connection links within the granular cluster. Since each link port of the core hub is connected to a core-to-core link, and the total bandwidth of all cluster connection links connected to a single block processing cluster is equal to the total bandwidth of multiple core-to-core links, when multiple sub-signals arrive at multiple link ports in parallel, multiple core-to-core links can be activated and driven synchronously. This enables efficient parallel data transmission between computing cores and input / output cores, making full use of the inherent parallel link resources in the multi-core integrated processor architecture. It transforms the congestion and latency problems that might have been caused by single-link bandwidth limitations into an efficient transmission mode of multi-link collaborative work, thereby effectively improving the overall throughput and response speed of cross-core communication.

[0087] For step 502, when a block processing cluster needs to exchange data with another block processing cluster within the same computing kernel, a cross-cluster communication signal can be generated. At this time, the communication management unit or related hardware will respond to this communication request. Similarly, in order to fully utilize the bandwidth of multiple parallel links, improve the throughput and efficiency of data transmission, and avoid a single link becoming a bottleneck, the original cross-cluster communication signal will be split into multiple cross-cluster communication sub-signals. For example, the original cross-cluster communication signal can be divided into fixed-size or variable-size data packets or frames, with each data packet being a cross-cluster communication sub-signal; or, the bits or bytes of the original signal can be allocated to different sub-signals in a certain order (such as cyclic interleaving).

[0088] Understandably, these cross-cluster communication sub-signals are transmitted in parallel to multiple inter-cluster buses via multiple cluster connection links. Cluster connection links are physical channels connecting block processing clusters and the kernel hub, while inter-cluster buses are logical or physical channels within the kernel hub connecting different block processing clusters. Parallel transmission utilizes the interconnection structure within the kernel hub to ensure efficient data routing between different block processing clusters. For example, the communication interface of a block processing cluster or the entry logic of the kernel hub can include a scheduler that dynamically allocates the split cross-cluster communication sub-signals to multiple available cluster connection links based on the current load of the cluster connection links and directs them to the corresponding inter-cluster buses; alternatively, each split cross-cluster communication sub-signal is mapped to a specific cluster connection link and inter-cluster bus, and then these sub-signals are transmitted to the cluster connection links of other block processing clusters to complete cross-cluster communication. Within the core hub, the routing logic can guide the sub-signal from the cluster connection link of the source block processing cluster to the corresponding inter-cluster bus based on the target block processing cluster address information contained in the cross-cluster communication sub-signal, and then forward it from the inter-cluster bus to the cluster connection link connected to the target block processing cluster.

[0089] Understandably, to overcome the bandwidth limitations of single-link transmission in traditional schemes, the original cross-cluster communication signal can be intelligently split into multiple smaller cross-cluster communication sub-signals. These sub-signals are no longer transmitted through a single path, but are instead distributed in parallel to multiple cluster connection links connected to the processing cluster. These cluster connection links are interconnected within the core node, forming multiple inter-cluster buses. Therefore, the parallel-transmitted sub-signals enter the core node through these cluster connection links and are routed to the corresponding inter-cluster buses. The routing mechanism within the core node forwards these sub-signals from the inter-cluster buses to the cluster connection links connected to the target processing cluster based on the target address of the sub-signals, and then the target processing cluster receives and reassembles them. In this way, the parallel transmission capabilities of multiple cluster connection links and inter-cluster buses within the computing core can be fully utilized, complementing the parallel transmission mechanism for cross-core communication signals. This embodiment of the present disclosure extends parallel transmission to cross-cluster communication within the core. Since the architecture has limited the total bandwidth of all cluster connection links connected to a single block processing cluster to be equal to the total bandwidth of multiple core-to-core links, by splitting and transmitting cross-cluster communication signals in parallel, the communication bandwidth within the core can be fully utilized, effectively avoiding internal link congestion and data accumulation caused by high-throughput data streams. It can also ensure that data exchange between the processing clusters within the computing core can be completed efficiently and in a timely manner, thereby improving the overall performance and response speed of the entire multi-core integrated processor architecture.

[0090] It is worth noting that cross-cluster communication signals and cross-core communication signals can be transmitted simultaneously within the core-core hub, or cross-core communication signals and external exchange communication signals can be transmitted simultaneously, or cross-cluster communication signals and external exchange communication signals can be transmitted simultaneously.

[0091] It is understood that although the steps in the above flowcharts are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated in this embodiment, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the above flowcharts may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.

[0092] Thirdly, embodiments of this disclosure also provide a processor comprising the multi-chip integrated processing architecture described in the first aspect above.

[0093] Fourthly, embodiments of this disclosure also provide an electronic device, the electronic device including a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for implementing connection communication between the processor and the memory, wherein the program is executed by the processor to implement the data transmission method described in the second aspect above.

[0094] This disclosure also provides an electronic device, including at least one processor and a memory communicatively connected to the at least one processor; wherein the processor and the memory are communicatively connected via a data bus, the memory stores a program, and the program is executed by the at least one processor to cause the at least one processor to implement the method as described in any of the above embodiments of this disclosure when executing instructions.

[0095] The following is combined with Figure 6 The hardware structure of the electronic device is described in detail. The electronic device includes: a processor 610, a memory 620, an input / output interface 630, a communication interface 640, and a bus 650.

[0096] The processor 610 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this disclosure.

[0097] The memory 620 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 620 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 620 and is called and executed by the processor 610 using the data transmission method of the embodiments of this disclosure.

[0098] The input / output interface 630 is used to realize information input and output;

[0099] The communication interface 640 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, Wi-Fi, Bluetooth, etc.).

[0100] Bus 650 transmits information between various components of the device (e.g., processor 610, memory 620, input / output interface 630, and communication interface 640);

[0101] The processor 610, memory 620, input / output interface 630 and communication interface 640 are connected to each other within the device via bus 650.

[0102] This disclosure also provides a computer-readable storage medium for storing a computer program for executing the data transmission methods of the foregoing embodiments.

[0103] This disclosure also provides a computer program product comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the data transmission method described above.

[0104] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in this disclosure and the foregoing drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this disclosure described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “including,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatuses.

[0105] It should be understood that in this disclosure, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0106] It should be understood that in the description of the embodiments disclosed herein, "multiple" means two or more, "greater than", "less than", "exceeding" etc. are understood to exclude the number itself, and "above", "below", "within" etc. are understood to include the number itself.

[0107] In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0108] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0109] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0110] It should also be understood that the various implementation methods provided in this disclosure can be combined arbitrarily to achieve different technical effects.

[0111] The above is a detailed description of the embodiments of this disclosure. However, this disclosure is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this disclosure. All such equivalent modifications or substitutions are included within the scope defined by the claims of this disclosure.

Claims

1. A multi-core chiplet integrated processor architecture, comprising: The system includes input / output kernels and compute kernels. Each compute kernel includes a kernel hub and multiple block processing clusters. Each kernel hub is connected to the input / output kernel via multiple parallel kernel-to-kernel links. The kernel hub includes multiple inter-cluster buses and multiple link ports corresponding one-to-one with each kernel-to-kernel link. Cluster connection links of multiple block processing clusters are interconnected within the kernel hub via the inter-cluster buses. Each cluster connection link within the kernel hub is connected to both the inter-cluster buses and the link ports. The total bandwidth of all cluster connection links connected to a single block processing cluster is equal to the total bandwidth of the multiple kernel-to-kernel links. Each block processing cluster includes a cluster hub, an intra-cluster shared memory, and multiple compute units. The intra-cluster shared memory and the multiple compute units are respectively connected to the cluster hub. The cluster hub is connected to the kernel hub via multiple cluster connection links. The cluster hub is used to distribute data from the multiple compute units to the multiple cluster connection links connected to it for parallel transmission. In this context, a single block processing cluster simultaneously drives multiple chip-to-chip links between the computing chip and the input / output chip to transmit data through the chip hub.

2. The multi-core tile integrated processor architecture of claim 1, wherein, The core hub structure includes multiple cluster ports, which are connected to multiple block processing clusters via multiple cluster connection links; the cluster ports connected to different block processing clusters are interconnected, and the link ports are respectively connected to each of the cluster ports.

3. The multi-core grid integrated processor architecture of claim 1, wherein, Each of the block processing clusters is connected to a cluster connection link, which includes a first cluster connection link and a second cluster connection link. The core-to-core link between the core hub and the input / output core includes a first core-to-core link and a second core-to-core link. The plurality of link ports include a first link port and a second link port. A first branch of the first cluster connection link is connected to a first inter-cluster bus, and a second branch is connected to the first link port. A first branch of the second cluster connection link is connected to a second inter-cluster bus, and a second branch is connected to the second link port. The first link port is connected to the first core-to-core link, and the second link port is connected to the second core-to-core link.

4. The multi-core grain integrated processor architecture of claim 1, wherein, The total bandwidth of all cluster connection links connected to a single block processing cluster is a multiple of the bandwidth of a single core-to-core link.

5. The multi-core grain integrated processor architecture of claim 1, wherein, The cluster hub is also used to aggregate the data generated by the multiple computing units, and distribute the aggregated data to the cluster connection links for parallel transmission at a rate equal to the total bandwidth of all the cluster connection links connected to the cluster hub.

6. The multi-core grain integrated processor architecture of claim 1, wherein, Data generated by a single computing unit is distributed through the cluster hub to multiple cluster connection links connected to the cluster hub for parallel transmission.

7. The multi-core grain integrated processor architecture of claim 1, wherein, The bandwidth of the transmission link between the intra-cluster shared memory and the cluster hub is equal to the total bandwidth of all cluster connection links connected to a single cluster hub.

8. The multi-core grain integrated processor architecture of claim 1, wherein, The number of block processing clusters connected to a single core hub is proportional to the number of core-to-core links between a single core hub and the input / output cores.

9. The multi-core grain integrated processor architecture of claim 1, wherein, The input / output core also includes an external switching port for connection to external network switching equipment, and a cross-core transmission port for connection to other input / output cores; wherein the signal types transmitted within the core hub include: Cross-cluster communication signal, wherein the transmission path of the cross-cluster communication signal is that any of the block processing clusters is transmitted to another block processing cluster within the same core hub via the inter-cluster bus; And / or, cross-core communication signals, wherein the transmission path of the cross-core communication signals is that any of the block processing clusters is transmitted to the cross-core transmission port through the link port of the core hub; And / or, external exchange communication signals, wherein the transmission path of the external exchange communication signals is that any of the block processing clusters is transmitted to the external exchange port through the link port of the core hub.

10. A data transmission method, characterized by, Applied to the multi-chip integrated processor architecture as described in any one of claims 1 to 9, the data transmission method includes: In response to the block processing cluster sending a cross-core communication signal, the cross-core communication signal is split into multiple cross-core communication sub-signals, and sent in parallel to multiple core-to-core links through all the cluster connection links connected by the core hub, so that multiple core-to-core links are simultaneously driven to transmit the cross-core communication signal.

11. The data transmission method of claim 10, wherein, The data transmission method further includes: In response to the block processing cluster sending an inter-cluster communication signal, the inter-cluster communication signal is split into multiple inter-cluster communication sub-signals and sent in parallel to multiple inter-cluster buses through all the cluster connection links connected by the core hub, so as to transmit the inter-cluster communication signal to the cluster connection links of other block processing clusters.

12. A processor, comprising: The processor includes the multi-chip integrated processor architecture as described in any one of claims 1 to 9.

13. An electronic device, comprising: The electronic device includes a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for enabling communication between the processor and the memory, wherein the program is executed by the processor to implement the data transmission method as described in any one of claims 10 to 11.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, which can be executed by one or more processors to implement the data transmission method as described in any one of claims 10 to 11.