Computing chiplet and electronic device

By computing the 3D stacking connection between the memory chip and the storage chip, the number of I/O interfaces is increased and the signal transmission path is shortened, which solves the problems of limited memory access bandwidth and high power consumption in traditional computing chips, and achieves a high-efficiency improvement in memory access performance.

WO2025214040A1PCT designated stage Publication Date: 2025-10-16MOORE THREADS TECH CO LTD

Patent Information

Application Number
PCT/CN2025/081803
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-07
Filing Date
2025-03-11
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

In traditional computing chips, it is difficult to increase the I/O interface density of the computing engine and storage array, the signal transmission bandwidth is limited, and the long signal transmission path leads to increased power consumption, which affects memory access performance.

Method used

By adopting a 3D stacking connection method between computing chips and storage chips, the number of I/O interfaces is increased and the signal transmission path is shortened, and efficient memory access command transmission is achieved through a routing system.

Benefits of technology

Significantly improve memory access bandwidth and reduce power consumption, improving signal transmission efficiency between the computing engine and the storage array.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025081803_16102025_PF_FP_ABST
    Figure CN2025081803_16102025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of chips. Provided are a computing chiplet and an electronic device. The computing chiplet is arranged on a first plane and is connected to a storage chiplet on a second plane, wherein storage arrays of the storage chiplet are connected to storage controllers of the computing chiplet, the number of transmission channels between each pair of connected storage controllers and storage arrays is greater than a first threshold value, and the connection direction of the storage controllers and the storage arrays intersects the first plane. The increase in the number of transmission channels can improve the memory access bandwidth, and arranging the computing chiplet and the storage chiplet on different planes can shorten a signal transmission path between the computing chiplet and the storage chiplet, thereby reducing power consumption. Using an optimized routing system to complete the communication between computing engines and the storage controllers of the computing chiplet reduces the number of specification types of routing channels, thereby reducing the design complexity of the routing system. The routing system can be arranged on the computing chiplet in a distributed manner, reducing the area overhead of the computing chiplet.
Need to check novelty before this filing date? Find Prior Art

Description

Computing core and electronic device

[0001] This application claims priority to the Chinese Patent Application No. 202410412656.8, filed on April 7, 2024, entitled “Computing Core and Electronic Device”, the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] The present disclosure relates to the field of chips, and in particular, to a computing core and an electronic device. BACKGROUND

[0003] With the increasing scale of Artificial Intelligence (AI) applications and High performance computing (HPC) applications in recent years, higher and higher requirements are put forward for the performance of the computing chips (such as Graphics Processing Unit (GPU) and the like) for implementing the AI applications and the HPC applications.

[0004] The performance of the computing chip can be evaluated from the aspects of memory access and operation. In the computing chip, a plurality of computing engines and a plurality of storage arrays are usually provided. In actual application, the computing engine first generates a read command, and the read command is executed to read the data to be processed from the storage array and return to the computing engine. The computing engine performs operation according to the read data, generates a write command, and the write command is executed to write the operation result into the storage array. That is, the signals transmitted between the computing engine and the storage array are mainly memory access commands and data. The memory access performance of the computing chip is determined by the bandwidth and power consumption of the signal transmission between the computing engine and the storage array.

[0005] In the conventional computing chip, the I / O interface density of the computing engine and the storage array is difficult to improve due to the packaging mode of the computing engine and the storage array and the size of the hardware input / output (I / O) interface, which directly limits the bandwidth of the signal transmission. In addition, in order to improve the memory access performance, each computing engine needs to be able to access any storage array. However, the distance between the computing engine and the storage array on the chip is usually far, resulting in a long transmission path and increased power consumption of the signal transmission. Therefore, how to improve the memory access bandwidth of the computing engine to the storage array and reduce the power consumption of the signal transmission between the computing engine and the storage array has become a research hotspot in the field. SUMMARY

[0006] Therefore, the disclosure provides a computing core and an electronic device. The computing core of the disclosure is arranged on a first plane and connected to a storage core on a second plane, forming a 3D stacked connection mode, thereby increasing the number of input / output interfaces between the computing core and the storage core, shortening the signal transmission path between the computing core and the storage core, and achieving a significant increase in memory bandwidth and a significant reduction in power consumption.

[0007] According to an aspect of the disclosure, a computing core is provided. The computing core is arranged on a first plane and connected to a storage core on a second plane. The storage core includes a plurality of storage arrays. The computing core includes a plurality of computing engines, a plurality of storage controllers, and a routing system connecting the computing engines and the storage controllers. Each storage controller is connected to one storage array. The number of transmission channels between each pair of connected storage controllers and storage arrays is greater than a first threshold, and the connection direction intersects the first plane. The computing engine is configured to generate a memory access command and transmit the memory access command to the routing system. The memory access command includes a memory access type and a memory access address. The routing system is configured to, in response to receiving the memory access command and the memory access address being the address of the storage array connected to the computing core, transmit the memory access command to the storage controller connected to the storage array. The storage controller is configured to, in response to receiving the memory access command, access the connected storage array according to the memory access type.

[0008] In a possible implementation, the computing core is connected to the storage core by hybrid bonding or micro-bump or through silicon via.

[0009] In a possible implementation, the connection direction of the storage controller and the storage array is perpendicular to the first plane.

[0010] In a possible implementation, the computing core is arranged in an electronic device, the electronic device includes a plurality of computing cores, the computing core further includes at least one first interconnection receiving controller and at least one first interconnection sending controller connected to the routing system, each first interconnection receiving controller is connected to a second interconnection sending controller included in another computing core, each first interconnection sending controller is connected to a second interconnection receiving controller included in another computing core, the routing system is further configured to, in response to receiving the memory access command and the memory access address being the address of the memory array connected to the other computing core, transmit the memory access command to the first interconnection sending controller of the computing core to which the other computing core is connected; the first interconnection sending controller is configured to, in response to receiving the memory access command from the routing system, transmit the memory access command to the second interconnection receiving controller of the other computing core; and the first interconnection receiving controller is configured to, in response to receiving the memory access command from the second interconnection sending controller in the other computing core, transmit the memory access command to the routing system.

[0011] In a possible implementation, the routing system includes an n-row and m-column router array and a plurality of routing channels connecting adjacent routers, a router in the ith row and the jth column is connected to A computing engines, B memory controllers, C first interconnection sending controllers, and D first interconnection receiving controllers, n and m are positive integers, A, B, C, and D are integers greater than or equal to 0, 1≤i≤n, and 1≤j≤m; the computing engine is configured to transmit the memory access command to the router connected to the computing engine in the routing system; and the router connected to the computing engine in the routing system is configured to transmit the memory access command to the memory controller connected to the memory array.

[0012] In a possible implementation, the routing system includes a first router, a second router, and a third router, the first router is connected to the computing engine and the memory controller connected to the memory array, and is configured to directly transmit the memory access command to the memory controller connected to the memory array; the second router is connected to the computing engine, and the third router is connected to the memory controller connected to the memory array, the second router is configured to transmit the memory access command to the third router through the routing channel, and the third router is configured to transmit the memory access command to the memory controller connected to the memory array.

[0013] In a possible implementation, in response to two routers connected by a single routing channel being in the same row, the routing direction of the routing channel is horizontal direction, in response to two routers connected by a single routing channel being in the same column, the routing direction of the routing channel is vertical direction, in response to signal transmission between different routers needing to pass through a routing channel in horizontal direction and a routing channel in vertical direction, the signal first passes through the routing channel in horizontal direction and then passes through the routing channel in vertical direction, or, the signal first passes through the routing channel in vertical direction and then passes through the routing channel in horizontal direction, the signal including the memory access command and data to be written into the storage array, or including data read from the storage array.

[0014] In a possible implementation, each router is connected with A computing engines, B storage controllers, C first interconnection sending controllers, and D first interconnection receiving controllers, and each routing channel connecting a jth column router and a (j+1)th column router is the same, and each routing channel connecting an ith row router and an (i+1)th row router is the same; when m=n, each routing channel in the routing system is the same.

[0015] In a possible implementation, the horizontal routing channel includes a first transmission direction from left to right and a second transmission direction from right to left, the vertical routing channel includes a third transmission direction from top to bottom and a fourth transmission direction from bottom to top, in the case of passing through the horizontal routing channel and then passing through the vertical routing channel, the maximum number of signals simultaneously transmitted in the first transmission direction of the routing channel connecting the router in the i th row and the j th column and the router in the i th row and the j + 1 th column is equal to a first number; wherein the first number is the minimum of a second number and a third number, the second number refers to the total number of storage controllers and first interconnection receiving controllers connected by the j + 1 th to m th column routers, and the third number refers to the total number of computing engines and first interconnection sending controllers connected by the 1 th to j th routers in the i th row; the maximum number of signals simultaneously transmitted in the second transmission direction of the routing channel connecting the router in the i th row and the j th column and the router in the i th row and the j + 1 th column is equal to a fourth number, wherein the fourth number is the minimum of a fifth number and a sixth number, the fifth number refers to the total number of storage controllers and first interconnection receiving controllers connected by the 1 th to j th column routers, and the sixth number refers to the total number of computing engines and first interconnection sending controllers connected by the j + 1 th to m th routers in the i th row; the maximum number of signals simultaneously transmitted in the third transmission direction of the routing channel connecting the router in the i th row and the j th column and the router in the i + 1 th row and the j th column is equal to a seventh number, wherein the seventh number is the minimum of an eighth number and a ninth number, the eighth number refers to the total number of storage controllers and first interconnection receiving controllers connected by the i + 1 th to n th routers in the j th column, and the ninth number refers to the total number of computing engines and first interconnection sending controllers connected by the 1 th to i th row routers; the maximum number of signals simultaneously transmitted in the fourth transmission direction of the routing channel connecting the router in the i th row and the j th column and the router in the i + 1 th row and the j th column is equal to a tenth number, wherein the tenth number is the minimum of an eleventh number and a twelfth number, the eleventh number refers to the total number of storage controllers and first interconnection receiving controllers connected by the 1 th to i th routers in the j th column, and the twelfth number refers to the total number of computing engines and first interconnection sending controllers connected by the i + 1 th to n th row routers.

[0016] In a possible implementation, the horizontal routing channel includes a first transmission direction from left to right and a second transmission direction from right to left, the vertical routing channel includes a third transmission direction from top to bottom and a fourth transmission direction from bottom to top, in the case of passing through the vertical routing channel first and then passing through the horizontal routing channel, the maximum number of signals transmitted simultaneously in the first transmission direction of the routing channel connecting the router in the i-th row and the j-th column and the router in the i-th row and the j+1-th column is equal to the thirteenth number, wherein the thirteenth number is the minimum of the fourteenth number and the fifteenth number, the fourteenth number refers to the total number of storage controllers and first interconnection receiving controllers connected by the j+1-th to m-th router in the i-th row, and the fifteenth number refers to the total number of computing engines and first interconnection sending controllers connected by the 1st to j-th router; the maximum number of signals transmitted simultaneously in the second transmission direction of the routing channel connecting the router in the i-th row and the j-th column and the router in the i-th row and the j+1-th column is equal to the sixteenth number, wherein the sixteenth number is the minimum of the seventeenth number and the eighteenth number, the seventeenth number refers to the total number of storage controllers and first interconnection receiving controllers connected by the 1st to j-th router in the i-th row, and the eighteenth number refers to the total number of computing engines and first interconnection sending controllers connected by the j+1-th to m-th router; the maximum number of signals transmitted simultaneously in the third transmission direction of the routing channel connecting the router in the i-th row and the j-th column and the router in the i+1-th row and the j-th column is equal to the nineteenth number, wherein the nineteenth number is the minimum of the twentieth number and the twenty-first number, the twentieth number refers to the total number of storage controllers and first interconnection receiving controllers connected by the i+1-th to n-th router, and the twenty-first number refers to the total number of computing engines and first interconnection sending controllers connected by the 1st to i-th router in the j-th column; the maximum number of signals transmitted simultaneously in the fourth transmission direction of the routing channel connecting the router in the i-th row and the j-th column and the router in the i+1-th row and the j-th column is equal to the twenty-second number, wherein the twenty-second number is the minimum of the twenty-third number and the twenty-fourth number, the twenty-third number refers to the total number of storage controllers and first interconnection receiving controllers connected by the 1st to i-th router, and the twenty-fourth number refers to the total number of computing engines and first interconnection sending controllers connected by the i+1-th to n-th router in the j-th column.

[0017] In a possible implementation, the router at the i-th row and the j-th column includes a plurality of sending interfaces and a plurality of receiving interfaces, wherein the maximum number of signals simultaneously transmitted in the first transmission direction of the routing channel connecting the router at the i-th row and the j-th column and the router at the i-th row and the j+1-th column is a first value, and the maximum number of signals simultaneously transmitted in the second transmission direction is a second value; the maximum number of signals simultaneously transmitted in the third transmission direction of the routing channel connecting the router at the i-th row and the j-th column and the router at the i+1-th row and the j-th column is a third value, and the maximum number of signals simultaneously transmitted in the fourth transmission direction is a fourth value; the maximum number of signals simultaneously transmitted in the first transmission direction of the routing channel connecting the router at the i-th row and the j-th column and the router at the i-th row and the j-1-th column is a fifth value, and the maximum number of signals simultaneously transmitted in the second transmission direction is a sixth value; the maximum number of signals simultaneously transmitted in the third transmission direction of the routing channel connecting the router at the i-th row and the j-th column and the router at the i-1-th row and the j-th column is a seventh value, and the maximum number of signals simultaneously transmitted in the fourth transmission direction is an eighth value; the plurality of sending interfaces include a first sending interface of the first value, a third sending interface of the third value, a sixth sending interface of the sixth value, an eighth sending interface of the eighth value, and a ninth sending interface; the plurality of receiving interfaces include a second receiving interface of the second value, a fourth receiving interface of the fourth value, a fifth receiving interface of the fifth value, a seventh receiving interface of the seventh value, and a tenth receiving interface; each first sending interface and each second receiving interface are connected to the routing channel connecting the router at the i-th row and the j-th column and the router at the i-th row and the j+1-th column; each third sending interface and each fourth receiving interface are connected to the routing channel connecting the router at the i-th row and the j-th column and the router at the i+1-th row and the j-th column; each fifth receiving interface and each sixth sending interface are connected to the routing channel connecting the router at the i-th row and the j-th column and the router at the i-th row and the j-1-th column; each seventh receiving interface and each eighth sending interface are connected to the routing channel connecting the router at the i-th row and the j-th column and the router at the i-1-th row and the j-th column; the ninth sending interface is connected to a storage controller connected to the router at the i-th row and the j-th column and a first interconnection receiving controller; and the tenth receiving interface is connected to a computing engine connected to the router at the i-th row and the j-th column and a first interconnection sending controller.

[0018] In a possible implementation, in the case of passing through the routing channel in the horizontal direction first and then passing through the routing channel in the vertical direction, each second receiving interface is connected to each third sending interface, each sixth sending interface, each eighth sending interface, and a ninth sending interface; each fourth receiving interface is connected to each eighth sending interface and a ninth sending interface; each fifth receiving interface is connected to each first sending interface, each third sending interface, each sixth sending interface, and a ninth sending interface; each seventh receiving interface is connected to each third sending interface and a ninth sending interface; and a tenth receiving interface is connected to each sending interface.

[0019] In a possible implementation, in the case of routing through a vertical routing channel and then through a horizontal routing channel, each second receiving interface is connected to each sixth sending interface and ninth sending interface; each fourth receiving interface is connected to each first sending interface, each sixth sending interface, each eighth sending interface, and ninth sending interface; each fifth receiving interface is connected to each first sending interface and ninth sending interface; each seventh receiving interface is connected to each first sending interface, each third sending interface, each sixth sending interface, and ninth sending interface; and the tenth receiving interface is connected to each sending interface.

[0020] According to another aspect of the present disclosure, an electronic device is provided, including at least one storage chip and at least one computing chip as described above.

[0021] According to the computing chip of the present disclosure, the storage chip is arranged on the first plane and connected to the second plane, the storage chip includes a plurality of storage arrays, the computing chip includes a plurality of computing engines, a plurality of storage controllers, and a routing system connecting the computing engines and the storage controllers, so that the computing chip and the storage chip form a 3D stacked structure, the distance between the computing engines and the storage arrays is shortened; each storage controller is connected to one storage array, and since the computing engines and the storage arrays are arranged on different planes, the number of input / output interfaces allowed to be arranged on the computing engines and the storage arrays is greatly increased, the number of transmission channels between each pair of connected storage controller and storage array is greater than a first threshold, so that the transmission bandwidth between the computing engines and the storage arrays is greatly increased, and the parasitic capacitance is reduced; the connection direction of each pair of connected storage controller and storage array intersects the first plane, so that the transmission path between the computing engines and the storage arrays is greatly shortened, and the signal transmission power consumption can be reduced. The computing engine is configured to generate a memory access command and transmit the memory access command to the routing system, the memory access command including a memory access type and a memory access address, the routing system is configured to, in response to receiving the memory access command and the memory access address being the address of the storage array connected to the computing chip, transmit the memory access command to the storage controller connected to the storage array, and the storage controller is configured to, in response to receiving the memory access command, access the connected storage array according to the memory access type, so that the computing chip of the present disclosure has the function of accessing the storage chip. In summary, the computing chip of the present disclosure can increase the number of input / output interfaces between the computing chip and the storage chip, shorten the signal transmission path between the computing chip and the storage chip, greatly improve the memory access bandwidth, and significantly reduce the power consumption.

[0022] Other features and aspects of the present disclosure will become apparent from the following detailed description of example embodiments, taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0023] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate examples of the present disclosure, and together with the description, serve to explain the principles of the present disclosure.

[0024] FIG. 1 shows a schematic diagram of an on-chip interconnect subsystem of a full interconnect architecture of the prior art.

[0025] FIG. 2 shows a schematic diagram of a setting manner of an on-chip interconnect subsystem of a full interconnect architecture of the prior art on a high-performance chip.

[0026] FIG. 3 shows an exemplary application scenario of a computing corelet according to an embodiment of the present disclosure.

[0027] FIG. 4 shows an exemplary application scenario of a computing corelet according to an embodiment of the present disclosure.

[0028] FIG. 5 shows an exemplary application scenario of a computing corelet according to an embodiment of the present disclosure.

[0029] FIG. 6a shows a schematic diagram of a structure of a computing corelet according to an embodiment of the present disclosure.

[0030] FIG. 6b shows a schematic diagram of a structure of a computing corelet according to an embodiment of the present disclosure.

[0031] FIG. 7 shows a schematic diagram of a structure of a computing corelet according to an embodiment of the present disclosure.

[0032] FIG. 8 shows an example of a layout of a router and a routing channel on a computing corelet according to an embodiment of the present disclosure.

[0033] FIG. 9 shows a schematic diagram of a transmission path of a signal in a routing system according to an embodiment of the present disclosure.

[0034] FIG. 10 shows a schematic diagram of a structure of a routing system according to an embodiment of the present disclosure.

[0035] FIG. 11 shows a schematic diagram of a structure of a routing system according to an embodiment of the present disclosure.

[0036] FIG. 12 shows a schematic diagram of a structure of a routing system according to an embodiment of the present disclosure.

[0037] FIG. 13 shows an exemplary structure diagram of a router according to an embodiment of the present disclosure.

[0038] FIG. 14 shows an exemplary structure diagram of a router according to an embodiment of the present disclosure.

[0039] FIG. 15 shows an example of adjusting a routing width of a routing channel according to an embodiment of the present disclosure.

[0040] FIG. 16 shows an example of adjusting a routing width of a routing channel according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0041] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numbers in different drawings represent the same or similar elements. Although various aspects of embodiments are illustrated in the drawings, the drawings are not necessarily drawn to scale unless specifically noted.

[0042] The term "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any implementation described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other implementations.

[0043] In addition, for the purpose of convenience and brevity, detailed descriptions of well-known functions and structures incorporated in the disclosure will be omitted. It will be appreciated that those skilled in the art will be able to devise various modes of implementing the disclosed disclosure, without the exercise of inventive faculty, and without departing from the scope of the disclosure as defined by the appended claims.

[0044] With the increasing scale of artificial intelligence (AI) applications and high performance computing (HPC) applications in recent years, higher and higher requirements are put forward for the performance of computing chips (such as graphics processing units (GPUs) and the like) for implementing AI applications and HPC applications.

[0045] The performance of a computing chip can be evaluated from the aspects of memory access and operation. A computing chip usually has multiple computing engines and multiple memory arrays, and the memory arrays usually use double data rate (DDR) / low power double data rate (LPDDR) / graphics double data rate (GDDR) synchronous dynamic random access memory and high bandwidth memory (HBM). In actual applications, a computing engine first generates a read command, and the read command is executed to read data to be processed from a memory array and return to the computing engine. The computing engine performs operation according to the read data, generates a write command, and the write command is executed to write the operation result into the memory array. That is, the signals transmitted between the computing engine and the memory array are mainly memory access commands and data. The memory access performance of the computing chip is determined by the bandwidth and power consumption of signal transmission between the computing engine and the memory array.

[0046] In a traditional computing chip, the I / O interface density of the computing engine and the storage array is difficult to improve due to the packaging mode of the computing engine and the storage array and the size of the hardware input / output interface (I / O), which directly limits the bandwidth of signal transmission. In addition, in order to improve the memory access performance, each computing engine needs to be able to access any storage array. However, the distance between the computing engine and the storage array on the chip is usually far, resulting in a long transmission path and increased signal transmission loss. In order to ensure the signal integrity of transmission, the power consumption of signal transmission is inevitable. Therefore, how to improve the memory access bandwidth of the computing engine to the storage array and reduce the power consumption of signal transmission between the computing engine and the storage array has become a research hotspot in the field.

[0047] In addition, an on-chip interconnection subsystem is also provided in the computing chip to connect the computing engine and the storage array and realize signal transmission between them. In a high-performance computing chip, the computing engine is more complex due to the increase in interconnection complexity. Therefore, the performance and efficiency of the on-chip interconnection subsystem are also key factors affecting the performance of the computing chip.

[0048] In the prior art, the commonly used interconnection architecture of the on-chip interconnection subsystem includes a shared bus interconnection architecture, a centralized arbitration full interconnection architecture, a star interconnection architecture, a hierarchical tree interconnection architecture, a ring bus interconnection architecture, a 2D mesh interconnection architecture, and a ring curved surface interconnection architecture. Among them, the centralized arbitration full interconnection architecture (hereinafter referred to as the full interconnection architecture) and the 2D mesh interconnection architecture are two main architectures used in high-performance computing chips.

[0049] FIG. 1 shows a schematic diagram of an on-chip interconnection subsystem of a full interconnection architecture in the prior art.

[0050] As shown in FIG. 1, the on-chip interconnection subsystem of the full interconnection architecture can realize complete interconnection of all computing engines and all storage arrays. The on-chip interconnection subsystem of the full interconnection architecture can be logically composed of three parts, which are a memory access distribution unit, a memory access arbitration unit, and a line connecting the memory access distribution unit and the memory access arbitration unit. The total number of the memory access distribution units is equal to the total number of the computing engines, and the total number of the memory access arbitration units is equal to the total number of the storage controllers. Each memory access distribution unit is connected to one computing engine and all memory access arbitration units; each memory access arbitration unit is connected to all memory access distribution units and one storage controller. Each storage controller is connected to one storage array.

[0051] Each memory access command issued by each computing engine first enters the memory access distribution unit of the on-chip interconnection subsystem. The memory access distribution unit transmits the memory access command (and memory data when the memory access command indicates a write operation) to the memory access arbitration unit corresponding to the address information included in the memory access command according to the address information (an address on a certain memory array) included in the memory access command. Each memory access arbitration unit can receive all memory access commands issued by all computing engines, and arbitrates a memory access command (and memory data when the memory access command indicates a write operation) to be output to the memory controller corresponding to the address information included in the memory access command. If the memory access command indicates a write operation, the memory data and the memory access command are transmitted to the memory array connected to the memory controller, and the memory array writes the received memory data to the address of the received memory access command. If the memory access command indicates a read operation, the memory access command is transmitted to the memory array connected to the memory controller, and the memory array reads data from the address information included in the received memory access command and returns the data to the computing engine.

[0052] When the on-chip interconnection subsystem using the full interconnection architecture is used, the high-performance chip has greater bus throughput capacity, but has obvious disadvantages in other aspects. FIG. 2 shows a schematic diagram of the setting mode of the on-chip interconnection subsystem of the prior art full interconnection architecture on a high-performance chip.

[0053] As shown in FIG. 2, in the high-performance chip, the computing engines and the memory controllers are generally arranged in a physical layout in which the computing engines are arranged at the center of the chip and the memory controllers are arranged around the periphery of the computing engines. In this case, a suitable position needs to be selected on the chip, such as the center of the chip, to arrange the on-chip interconnection subsystem of the full interconnection architecture, and the on-chip interconnection subsystem of the full interconnection architecture is connected to each computing engine and each memory controller, so that the maximum length of the connection line from the on-chip interconnection subsystem to each computing engine and each memory controller is minimized. The more the computing engines, the larger the logic and the wiring scale of the on-chip interconnection subsystem of the full interconnection architecture. As shown in FIG. 1, 9 computing engines and 9 memory arrays require 9x9=81 buses. Each bus has a memory access command channel, a write data channel and a read data channel. Due to the large number of buses, the wiring congestion problem is prominent, which leads to the need to reserve a relatively large area for the on-chip interconnection subsystem to route the wires.

[0054] In addition, in order to enable the computing engines and the memory controllers to be connected to the on-chip interconnection subsystem, a large number of wiring channels need to be reserved. In order to maintain a high operating clock rate of the bus, a large number of hardware logic modules (such as latches) need to be inserted in the wiring channels according to the wiring distance. In physical implementation, the area utilization rate of these long-distance wiring channels is relatively low, which also causes a waste of valuable chip area and becomes a limiting factor for further improving the chip performance scale and optimizing the chip efficiency.

[0055] In contrast, the on-chip interconnection subsystem of the prior art 2D mesh interconnection architecture sets up mesh routing nodes and routing channels between adjacent routing nodes, so that the computing engines and the storage controllers are connected to the routing nodes, and the interconnection between any one computing engine and any one storage controller can be completed through at least one router (optionally, also including routing channels). However, the prior art 2D mesh interconnection architecture is difficult to design a unified routing channel to provide the highest bandwidth throughput while meeting the adverse bandwidth conditions of simultaneous burst memory access of each computing engine.

[0056] Therefore, the present disclosure provides a computing core and an electronic device. The computing core of the present disclosure is arranged on a first plane and is connected to a storage core on a second plane to form a 3D stacked connection mode, thereby increasing the number of input / output interfaces between the computing core and the storage core, shortening the signal transmission path between the computing core and the storage core, and achieving a significant increase in memory bandwidth and a significant reduction in power consumption.

[0057] In the present disclosure, the routing mode supported by the routing system on the computing core is set, so that part of the routing channels is unified, the design complexity and manufacturing cost of the routing channels are reduced, the highest bandwidth throughput is provided, and the adverse bandwidth conditions of simultaneous burst memory access of each computing engine are met.

[0058] FIG. 3 shows an exemplary application scenario of a computing core according to an embodiment of the present disclosure.

[0059] As shown in FIG. 3, the computing core can be arranged on a first plane, and the storage core can be arranged on a second plane. The first plane and the second plane do not coincide.

[0060] The computing core is connected to the storage core through a plurality of transmission channels. The computing core can generate a memory access command, which can include a memory access type and a memory access address, and the memory access address can be an address on the storage core. The memory access type includes read access and write access. When the memory access type is write access, the computing core also generates memory access data (including data to be written into the storage array described below). The memory access command (and the memory access data) can be transmitted to the storage core through the transmission channel. The storage core can respond to the memory access command. If the memory access type is read access, the storage core reads the data stored at the memory access address in a read response and returns the read data to the computing core. If the memory access type is write access, the storage core writes the memory access data into the memory access address in a write response.

[0061] FIG. 4 shows an exemplary application scenario of a computing core according to an embodiment of the present disclosure.

[0062] As shown in FIG. 4, the scenario can include a plurality of computing dies and a plurality of memory dies, each computing die is connected to one memory die through a plurality of transmission channels, and adjacent computing dies are connected to each other. The planes where the two connected computing dies and the memory die are located do not coincide and are intersected by the transmission channels connecting the two.

[0063] Each computing die can generate a memory access command, which can include a memory access type and a memory access address, which can be an address on a certain memory die. The memory access type includes read access and write access. When the memory access type is write access, the computing die also generates memory access data (including data to be written into the memory array described below). The memory access command (and the memory access data) can be transmitted to the memory die to which the memory access address belongs through the transmission channel. The memory die can respond to the memory access command. If the memory access type is read access, the memory die reads out the data stored at the memory access address in a read response and returns the read data to the computing die. If the memory access type is write access, the memory die writes the memory access data into the memory access address in a write response.

[0064] FIG. 5 shows an exemplary application scenario of a computing die according to an embodiment of the present disclosure.

[0065] As shown in FIG. 5, the scenario can include a plurality of computing dies and a plurality of memory dies disposed on a silicon interposer or a package substrate, each computing die is connected to one memory die through a plurality of transmission channels (not shown), and adjacent computing dies are connected to each other. The planes where the two connected computing dies and the memory die are located do not coincide and are intersected by the transmission channels (not shown) connecting the two. The purpose of each computing die and each memory die is the same as described in the related description of FIG. 4, and will not be repeated here.

[0066] Those skilled in the art should understand that in the application scenarios shown in FIG. 4 and FIG. 5, the plurality of memory dies can be disposed on a plurality of planes and stacked with each other, and the adjacent memory dies can be connected through hybrid bonding or micro-bumps or through silicon vias; the plurality of memory dies can also be disposed on the same plane, and the present disclosure does not limit this.

[0067] An exemplary method for a computing die to implement the above functions is described below. FIG. 6a and FIG. 6b show a schematic diagram of the structure of a computing die according to an embodiment of the present disclosure.

[0068] As shown in FIG. 6a and FIG. 6b, in a possible implementation, the computing die is disposed on a first plane and connected to a memory die on a second plane, the memory die includes a plurality of memory arrays,

[0069] The computing die includes a plurality of computing engines, a plurality of storage controllers, and a routing system connecting the computing engines and the storage controllers, each storage controller is connected to a storage array, the number of transmission channels between each pair of connected storage controller and storage array is greater than a first threshold, and the connection direction intersects a first plane;

[0070] The computing engine is configured to generate a memory access command and transmit the memory access command to the routing system, the memory access command including a memory access type and a memory access address;

[0071] The routing system is configured to, in response to receiving the memory access command and the memory access address being the address of the storage array connected to the computing die to which the memory access command belongs, transmit the memory access command to the storage controller connected to the storage array;

[0072] The storage controller is configured to, in response to receiving the memory access command, access the connected storage array according to the memory access type.

[0073] For example, the computing die of the embodiment of the present disclosure is arranged on a first plane and can be connected to a storage die on a second plane. The storage die and the computing die are connected together in a 3D stacked manner. The I / O interface originally arranged on the side (perpendicular to the first plane) of the computing engine and the storage array can be arranged on the front (parallel to the first plane) with a larger area. The larger area allows more I / O interfaces to be arranged, thereby achieving a substantial increase in memory bandwidth, avoiding high-speed signal transceiver circuits and overheads, and reducing parasitic capacitance.

[0074] The computing die can include a plurality of computing engines, a plurality of storage controllers, and a routing system connecting the computing engines and the storage controllers. For example, referring to FIG. 6b, the computing die X1 includes the routing system A1 and the computing engine B10, the computing engine B11, the storage controller C10, and the storage controller C11 connected to the routing system A1. The computing engine can be various types of processors, the storage controller can be a controller for storage control, and the routing system is used for communication between the computing engine and the storage controller. The embodiment of the present disclosure does not limit the specific implementation and structure of the computing engine and the storage controller. For examples of exemplary implementation and structure of the routing system, refer to the further description below.

[0075] The memory core can include a plurality of memory arrays, and each memory controller is connected to one memory array. For example, referring to FIG. 6b, the memory core Y1 includes a memory array D10 and a memory array D11. The memory controller C10 is connected to the memory array D10, and the memory controller C11 is connected to the memory array D11. The memory array can be a Double Data Rate (DDR) / Low Power Double Data Rate (LPDDR) / Graphics Double Data Rate (GDDR) synchronous dynamic random access memory and a High Bandwidth memory (HBM). Embodiments of the present disclosure do not limit the specific implementation and structure of the memory array.

[0076] As shown in FIG. 6a, the number of transmission channels between each pair of connected memory controller and memory array is greater than a first threshold, and the connection direction intersects the first plane. The number of transmission channels can be proportional to the number of I / O interfaces. The first threshold can be set according to the requirements of the application scenario, and embodiments of the present disclosure do not limit the specific value of the first threshold.

[0077] Since the connection direction of each pair of connected memory controller and memory array intersects the first plane where the memory controller is located, the bus length between the computing engine and the memory array can be significantly shortened, the transmission loss from the computing engine to the memory array is also reduced, and the memory access power consumption can be reduced while ensuring the accuracy of signal transmission.

[0078] The computing engine on each computing core can generate a memory access command, which can include a memory access type and a memory access address. The memory access address can be the address of any memory array. The memory access type includes read access and write access. When the memory access type is write access, the computing engine also generates memory access data (including data to be written into the memory array described below). The computing engine can transmit the memory access command (and the memory access data) to the routing system. For example, the memory access command (and the memory access data) generated by the computing engine B10 and the computing engine B11 can be transmitted to the routing system A1.

[0079] The routing system can transmit the memory access command (and the memory access data) to the memory controller connected to the memory array in response to receiving the memory access command (and the memory access data) and the memory access address being the address of the memory array connected to the computing core. As shown in FIG. 6b, when the computing engine B10 generates a memory access command and transmits it to the routing system A1, if the memory access address included in the memory access command is the address of the memory array D11 connected to the computing core X1, the routing system A1 can transmit the memory access command to the memory controller C11 connected to the memory array D11.

[0080] The storage controller can access the connected storage array according to the access type in response to receiving the access command. When the access type is read access, the storage controller performs read access on the storage array, and the data read from the storage array is returned to the computing engine that generates the access command. When the access type is write access, the storage controller performs write access on the storage array, and the access data is written into the storage array. As shown in FIG. 6b, the storage controller C11 receives the access command. If the access type is read access, the storage controller C11 performs read access on the storage array D11, reads the data stored at the access address on the storage array D11, and returns the read data to the computing engine B10. If the access type is write access, the storage controller C11 performs write access on the storage array D11, and writes the access data into the storage array D11 at the access address.

[0081] Optionally, the storage controller can also be connected to the routing system through a cache to further reduce signal transmission delay. The present disclosure does not limit whether the storage controller is directly connected to the routing system.

[0082] According to the computing core of the embodiment of the present disclosure, the computing core is arranged on the first plane and connected to the storage core on the second plane, the storage core includes a plurality of storage arrays, the computing core includes a plurality of computing engines, a plurality of storage controllers, and a routing system connecting the computing engines and the storage controllers, so that the computing core and the storage core form a 3D stacked structure, the distance between the computing engine and the storage array is shortened; each storage controller is connected to one storage array, and since the two are arranged on different planes, the number of input / output interfaces allowed to be arranged on the computing engine and the storage array is greatly increased, the number of transmission channels between each pair of connected storage controller and storage array is greater than a first threshold, the transmission bandwidth between the computing engine and the storage array is greatly increased, and the parasitic capacitance is reduced; the connection direction of each pair of connected storage controller and storage array intersects the first plane, so that the transmission path between the computing engine and the storage array is greatly shortened, and the signal transmission power consumption can be reduced. The computing engine is configured to generate an access command and transmit the access command to the routing system, the access command including an access type and an access address, the routing system is configured to, in response to receiving the access command and the access address being the address of the storage array connected to the computing core, transmit the access command to the storage controller connected to the storage array, and the storage controller is configured to, in response to receiving the access command, access the connected storage array according to the access type, so that the computing core of the embodiment of the present disclosure has the function of accessing the storage core. In summary, the computing core of the embodiment of the present disclosure can increase the number of input / output interfaces between the computing core and the storage core, shorten the signal transmission path between the computing core and the storage core, greatly increase the access bandwidth, and significantly reduce the power consumption.

[0083] In a possible implementation, the computing core is connected to the storage core by mixed bonding or micro-bump or through-silicon via. In this case, the signal transmission path length between the computing core and the storage core can be further shortened to the micron level. Thus, the memory access power consumption is greatly reduced and the memory access efficiency is improved.

[0084] In a possible implementation, the connection direction of the storage controller and the storage array is perpendicular to the first plane. In this case, no matter what connection mode is used between the computing core and the storage core, the signal transmission path between the computing core and the storage core can be made shortest, and the memory access performance is improved.

[0085] In a possible implementation, the computing core is arranged in an electronic device, and the electronic device includes a plurality of computing cores. The computing core further includes at least one first interconnection receiving controller and at least one first interconnection sending controller connected to a routing system. Each first interconnection receiving controller is connected to a second interconnection sending controller included in another computing core. Each first interconnection sending controller is connected to a second interconnection receiving controller included in another computing core.

[0086] The routing system is further configured to, in response to receiving the memory access command and the memory access address being the address of the storage array connected by the other computing core, transmit the memory access command to the first interconnection sending controller of the computing core connected to the other computing core.

[0087] The first interconnection sending controller is configured to, in response to receiving the memory access command from the routing system, transmit the memory access command to the second interconnection receiving controller of the other computing core.

[0088] The first interconnection receiving controller is configured to, in response to receiving the memory access command from the second interconnection sending controller of the other computing core, transmit the memory access command to the routing system.

[0089] For example, the computing core and the storage core can be arranged in an electronic device. When applied to the application scenarios shown in FIG. 4 and FIG. 5, that is, when the electronic device includes a plurality of computing cores, the computing core can further access the storage core not directly connected to the computing core but connected to the other computing core. To achieve this, the computing core further includes at least one first interconnection receiving controller and at least one first interconnection sending controller connected to a routing system. Each first interconnection receiving controller is connected to a second interconnection sending controller included in another computing core. Each first interconnection sending controller is connected to a second interconnection receiving controller included in another computing core. At this time, the computing cores have communication capability.

[0090] Those skilled in the art should understand that "first" and "second" are only used to indicate that the interconnection receiving controller and the interconnection sending controller in mutual communication are not disposed on the same computing core, and the disclosure does not limit the number of interconnection receiving controllers and interconnection sending controllers on a single computing core. For example, from the perspective of computing core X1, the first interconnection sending controller included therein can be the first interconnection sending controller; but from the perspective of computing core X2, the first interconnection sending controller included in computing core X1 can be the second interconnection sending controller.

[0091] Those skilled in the art should understand that when the electronic device only includes one computing core, the first interconnection receiving controller and the first interconnection sending controller do not need to be disposed on the computing core. Embodiments of the disclosure are not limited to whether the first interconnection receiving controller and the first interconnection sending controller are disposed on the computing core.

[0092] FIG. 7 shows a schematic diagram of the structure of a computing core according to an embodiment of the disclosure.

[0093] As shown in FIG. 7, computing core X1 includes routing system A1 and computing engine B10, computing engine B11, storage controller C10, storage controller C11, first interconnection sending controller E10, and first interconnection receiving controller E11 connected to routing system A1. Storage core Y1 includes storage array D10 and storage array D11. Storage controller C10 is connected to storage array D10, and storage controller C11 is connected to storage array D11.

[0094] Computing core X2 includes routing system A2 and computing engine B20, computing engine B21, storage controller C20, storage controller C21, second interconnection sending controller E20, and second interconnection receiving controller E21 connected to routing system A2. Storage core Y2 includes storage array D20 and storage array D21. Storage controller C20 is connected to storage array D20, and storage controller C21 is connected to storage array D21.

[0095] First interconnection sending controller E10 is connected to second interconnection receiving controller E21, and second interconnection sending controller E20 is connected to first interconnection receiving controller E11.

[0096] The computing engines on each computing core can generate a memory access command, which can include a memory access type and a memory access address, which can be the address of any storage array. The memory access type includes read access and write access. When the memory access type is write access, the computing engine also generates memory access data. The computing engine can transmit the memory access command (and memory access data) to the routing system connected thereto. For example, the memory access command (and memory access data) generated by computing engine B10 and computing engine B11 can be transmitted to routing system A1, and the memory access command (and memory access data) generated by computing engine B20 and computing engine B21 can be transmitted to routing system A2.

[0097] The routing system can transmit the memory access command (and the memory access data) to the first interconnect transmit controller in the computing core to which the routing system belongs and connected to the other computing core in response to receiving the memory access command (and the memory access data) and the memory access address being the address of the memory array connected to the other computing core.

[0098] As shown in FIG. 7, when the computing engine B10 generates a memory access command and transmits the memory access command to the routing system Al, if the memory access address included in the memory access command is the address of the memory array D21 connected to the computing core X2 (i.e., the other computing core), the routing system Al can transmit the memory access command to the first interconnect transmit controller E10 (i.e., the first interconnect transmit controller in the computing core to which the routing system belongs and connected to the other computing core).

[0099] The first interconnect transmit controller is configured to transmit the memory access command to the second interconnect receive controller of the other computing core in response to receiving the memory access command from the routing system. The second interconnect receive controller of the other computing core can transmit the memory access command to the routing system on the other computing core.

[0100] As shown in FIG. 7, the first interconnect transmit controller E10 can transmit the memory access command to the second interconnect receive controller E21 on the computing core X2 (i.e., the other computing core) according to the memory access address in response to receiving the memory access command from the routing system Al. The second interconnect receive controller E21 can transmit the memory access command to the routing system A2 according to the memory access address. The response of the routing system A2 to the memory access command can be the same as the response of the routing system Al to the memory access command as described in the related description of FIG. 6b, and thus will not be described herein.

[0101] Similarly, when the computing engine on the computing core X2 (i.e., the other computing core) generates a memory access command and the memory access address is the address of the memory array connected to the computing core X1, the computing core X2 (i.e., the other computing core) can also transmit the memory access command to the first interconnect receive controller E11 on the computing core X1 through the second interconnect transmit controller E20 included therein.

[0102] The first interconnect receiving controller is configured to transmit the memory access command to the routing system in response to receiving the memory access command from the second interconnect transmitting controller in another computing core. For example, when the first interconnect receiving controller E11 receives the memory access command from the second interconnect transmitting controller E20 in the computing core X2 (i.e., another computing core), the memory access command can be transmitted to the routing system A1. The routing system A1 can respond to the memory access command in the same manner as described in the description of the routing system A1 responding to the memory access command in FIG. 6b, which will not be repeated here.

[0103] Those skilled in the art should understand that when there are enough computing cores and memory cores, two computing cores can also be indirectly connected through one or more other computing cores. In this case, the memory access command (and the memory access data) generated by a computing core can not be directly transmitted to the computing core connected to the memory core to which the memory access address included in the memory access command belongs. In this case, the computing core generating the memory access command (and the memory access data) can determine the transmission path of the memory access command (and the memory access data) among the computing cores according to the memory access address, and transmit the memory access command (and the memory access data) to the next computing core on the transmission path; the next computing core continues to transmit the memory access command (and the memory access data) until the memory access command (and the memory access data) is transmitted to the last computing core on the transmission path, and then the routing system on the computing core receives and processes the memory access command (and the memory access data).

[0104] In this case, except for the first computing core and the last computing core on the transmission path, the main purpose of the remaining computing cores on the transmission path is transit. In the implementation of the transit purpose, the computing core can receive the memory access command (and the memory access data) through the first interconnect receiving controller thereon, and then transmit the memory access command (and the memory access data) to the second interconnect receiving controller of the next computing core on the transmission path through the first interconnect transmitting controller thereon.

[0105] In this way, the computing core has stronger memory access capability.

[0106] Those skilled in the art should understand that the first interconnect transmitting controller and the first interconnect receiving controller can be used to realize the communication between the computing cores on different chips in addition to realizing the communication between different computing cores on the same chip.

[0107] The exemplary structure of the routing system of the embodiment of the present disclosure will be described below.

[0108] In a possible implementation, the routing system comprises an n-row m-column router array and a plurality of routing channels connecting adjacent routers, a router in the i-th row and the j-th column connects A computing engines, B storage controllers, C first interconnection sending controllers, and D first interconnection receiving controllers, n and m are positive integers, A, B, C, and D are integers greater than or equal to 0, 1≤i≤n, and 1≤j≤m;

[0109] The computing engine is configured to transmit the memory access command to the router connected to the computing engine in the routing system.

[0110] The router connected to the computing engine in the routing system is configured to transmit the memory access command to the storage controller connected to the storage array.

[0111] For example, in the embodiments of the present disclosure, the routing system can be implemented by an n-row m-column router array and a plurality of routing channels connecting adjacent routers. The router refers to a device or a module with routing function. A single router (a router in the i-th row and the j-th column) can connect A computing engines, B storage controllers, C first interconnection sending controllers, and D first interconnection receiving controllers. The computing engine and the first interconnection sending controller can be a module initiating memory access, and the storage controller and the first interconnection receiving controller can be a module responding to memory access. Each router is connected to both the module initiating memory access and the module receiving memory access.

[0112] The total sending bandwidth of a single router can be equal to the total receiving bandwidth. If the interface bandwidth of each computing engine, each storage controller, each first interconnection sending controller, and each first interconnection receiving controller is equal, the total number of the computing engines and the first interconnection sending controllers connected to a single router is equal to the total number of the storage controllers and the first interconnection receiving controllers connected to the single router. That is, A+C can be equal to B+D.

[0113] Those skilled in the art should understand that the embodiments of the present disclosure do not limit whether the interface bandwidth of each computing engine, each storage controller, each first interconnection sending controller, and each first interconnection receiving controller is equal, and do not limit whether A+C and B+D are equal, as long as the total sending bandwidth of a single router is equal to the total receiving bandwidth.

[0114] The router can be combined with the storage controller connected thereto into one module, or can be independent of each other, and the embodiments of the present disclosure do not limit this.

[0115] Those skilled in the art should understand that when the electronic device only comprises one computing core, C=D=0, that is, each router only connects the computing engine and the storage controller.

[0116] The connection mode of the router and each computing engine, each storage controller, each first interconnection sending controller and each first interconnection receiving controller can be set, so that each computing engine, each storage controller, each first interconnection sending controller and each first interconnection receiving controller is connected to at least one router.

[0117] In this case, the computing engine generates a memory access command and transmits it to the routing system, which can generate a memory access command and transmit it to the router connected to the computing engine in the routing system. The routing system transmits the memory access command to the storage controller connected to the storage array, which can be the router connected to the computing engine transmitting the memory access command to the storage controller connected to the storage array. The details of the router transmitting the memory access command can be referred to the relevant description below. FIG. 8 shows an example of the layout of the router, the routing channel on the computing core according to an embodiment of the present disclosure. As shown in FIG. 8, each computing engine and storage controller can be uniformly distributed on the entire computing core and interconnected by uniformly distributed routers and routing channels, and the first interconnection sending controller and the first interconnection receiving controller are distributed on the edge of the computing core. Those skilled in the art should understand that the physical layout of the router in the routing system is not necessarily an absolutely evenly distributed array, and adjacent routers and routing channels can be combined according to the requirements of the application scenario, and the specific structure of the routing system is not limited in the embodiments of the present disclosure.

[0118] In one possible implementation, the routing system includes a first router, a second router and a third router,

[0119] The first router is connected to the computing engine and the storage controller connected to the storage array, and is configured to directly transmit the memory access command to the storage controller connected to the storage array;

[0120] The second router is connected to the computing engine, and the third router is connected to the storage controller connected to the storage array. The second router is configured to transmit the memory access command to the third router through the routing channel, and the third router is configured to transmit the memory access command to the storage controller connected to the storage array.

[0121] For example, there are two possible cases for the router connected with the compute engine to transmit the memory access command to the storage controller connected with the storage array. In one case, the compute engine and the storage controller connected with the storage array (e.g., the storage array D11) are both connected with the first router in the routing system, and the first router is both the source router and the destination router, and the memory access command is transmitted only through the first router without passing through the routing channel and other routers. That is, the first router directly transmits the memory access command to the storage controller (e.g., the storage controller C11) connected with the storage array (e.g., the storage array D11).

[0122] In another case, the compute engine is connected with the second router in the routing system, and the storage controller connected with the storage array (e.g., the storage array D11) is connected with the third router in the routing system, and the second router is the source router and the third router is the destination router, and the memory access command cannot be transmitted to the storage controller C11 only through the second router, and needs to pass through the routing channel. That is, the second router transmits the memory access command to the third router through the routing channel, and the third router transmits the memory access command to the storage controller (e.g., the storage controller C11) connected with the storage array (e.g., the storage array D11).

[0123] In this case, if the second router and the third router are adjacent, the second router can transmit the memory access command to the third router only through the routing channel. If the second router and the third router are not adjacent, the second router can also transmit the memory access command to the third router through other routers and routing channels.

[0124] Since the routers and the routing channels form a mesh structure, there are many optional routing channels when a signal (including a memory access command and / or data) is transmitted from one router to another router. To avoid congestion of the routing channels, the transmission manner of the signal in the routing system can be set in advance. Examples of the transmission manner of the signal in the routing system according to the embodiments of the present disclosure are given below.

[0125] In one possible implementation, in response to two routers connected with a single routing channel being in the same row, the routing direction of the routing channel is a horizontal direction, and in response to two routers connected with a single routing channel being in the same column, the routing direction of the routing channel is a vertical direction,

[0126] In response to the signal transmission between different routers needing to pass through horizontal routing channels and vertical routing channels, the signal, including the memory access command and data to be written into the storage array, or including data read from the storage array, passes through the horizontal routing channels first and then the vertical routing channels, or passes through the vertical routing channels first and then the horizontal routing channels.

[0127] For example, since the routing system includes an n-row-m-column router array, each row includes m routers and each column includes n routers. The routing channels connect adjacent routers, that is, any two adjacent routers in each row are connected by a routing channel and any two adjacent routers in each column are connected by a routing channel. When the two routers connected by a single routing channel are in the same row, the routing direction of the routing channel can be horizontal, and when the two routers connected by a single routing channel are in the same column, the routing direction of the routing channel can be vertical.

[0128] In this case, the routing channels that the signal transmission between different routers needs to pass through can be only horizontal routing channels, only vertical routing channels, or both horizontal routing channels and vertical routing channels. The signal includes the memory access command and data to be written into the storage array, or includes data read from the storage array.

[0129] When the signal transmission between different routers needs to pass through horizontal routing channels and vertical routing channels, the routing channels of the two routing directions can be set to pass through the horizontal routing channels first and then the vertical routing channels, or set to pass through the vertical routing channels first and then the horizontal routing channels. FIG. 9 shows a schematic diagram of a transmission path of a signal in a routing system according to an embodiment of the present disclosure.

[0130] As shown in FIG. 9, it is assumed that m = 4 and n = 3. Taking the case of passing through the horizontal routing channels first and then the vertical routing channels as an example, when transmitting the signal from the source router 02 to the destination router 21, the transmission path can be router 02→router 12→router 22→router 21; when transmitting the signal from the source router 10 to the destination router 31, the transmission path can be router 10→router 20→router 30→router 31.

[0131] When passing through the vertical routing channels first and then the horizontal routing channels, the determination manner of the transmission path is similar to that of passing through the horizontal routing channels first and then the vertical routing channels, which will not be described here again.

[0132] Each routing channel can transmit signals in two directions, for example, the horizontal routing channel can transmit signals in a first transmission direction from left to right and a second transmission direction from right to left, and the vertical routing channel can transmit signals in a third transmission direction from top to bottom and a fourth transmission direction from bottom to top. In the case of different setting modes of the routing channel routing sequence in two routing directions, the routing width of the routing channel in each transmission direction can be different. The routing width of the routing channel in each transmission direction can be equal to the maximum number of signals transmitted simultaneously in the transmission direction.

[0133] The following describes an exemplary setting mode of the routing width of the routing channel in the embodiment of the present disclosure.

[0134] In a possible implementation, the horizontal routing channel includes a first transmission direction from left to right and a second transmission direction from right to left, and the vertical routing channel includes a third transmission direction from top to bottom and a fourth transmission direction from bottom to top,

[0135] In the case of first passing through the horizontal routing channel and then passing through the vertical routing channel,

[0136] The maximum number of signals transmitted simultaneously in the first transmission direction of the routing channel connecting the router in the i-th row and the j-th column and the router in the i-th row and the j+1-th column is equal to the first number, and the first number is the minimum value of the second number and the third number, the second number is the total number of storage controllers and first interconnection receiving controllers connected by the j+1-th to m-th column routers, and the third number is the total number of computing engines and first interconnection sending controllers connected by the 1st to j-th routers in the i-th row;

[0137] The maximum number of signals transmitted simultaneously in the second transmission direction of the routing channel connecting the router in the i-th row and the j-th column and the router in the i-th row and the j+1-th column is equal to the fourth number, and the fourth number is the minimum value of the fifth number and the sixth number, the fifth number is the total number of storage controllers and first interconnection receiving controllers connected by the 1st to j-th column routers, and the sixth number is the total number of computing engines and first interconnection sending controllers connected by the j+1-th to m-th routers in the i-th row;

[0138] The maximum number of signals transmitted simultaneously in the third transmission direction of the routing channel connecting the router in the i-th row and the j-th column and the router in the i+1-th row and the j-th column is equal to the seventh number, and the seventh number is the minimum value of the eighth number and the ninth number, the eighth number is the total number of storage controllers and first interconnection receiving controllers connected by the i+1-th to n-th routers in the j-th column, and the ninth number is the total number of computing engines and first interconnection sending controllers connected by the 1st to i-th row routers;

[0139] The maximum number of signals transmitted simultaneously by the routing channel connecting the router in the ith row and the jth column and the router in the (i+1)th row and the jth column in the fourth transmission direction is equal to the tenth number, where the tenth number is the minimum of the eleventh number and the twelfth number, the eleventh number is the total number of storage controllers and first interconnection receiving controllers connected by the routers in the jth column and the first i routers, and the twelfth number is the total number of computing engines and first interconnection sending controllers connected by the routers in the i+1th to nth rows.

[0140] As shown in FIG. 9, taking i=2 and j=2 as an example, the router in the ith row and the jth column can be router 11, the router in the ith row and the j+1th column can be router 21, and the router in the (i+1)th row and the jth column can be router 10.

[0141] When the routing channel connecting the router in the ith row and the jth column and the router in the ith row and the j+1th column transmits a signal in the first transmission direction, the destination router of the signal can be any one of the j+1th to mth routers, and the source router of the signal can be any one of the first j routers in the ith row.

[0142] Referring to FIG. 9, when the routing channel between router 11 and router 21 transmits a signal in the first transmission direction from left to right, since it is already in the horizontal direction, the destination router of the signal can be any one of the routers to the right of router 11, i.e., any one of routers 20, 21, 22, 30, 31, and 32. The source router of the signal can only be any one of the routers to the left of router 21 and in the same row as router 21, i.e., one of router 01 and router 11.

[0143] Without considering the routing direction, assuming that the routing width is large enough, the maximum number of signals received simultaneously by each router is equal to the total number of storage controllers and first interconnection receiving controllers connected by the router. The maximum number of signals transmitted simultaneously by each router is equal to the total number of computing engines and first interconnection sending controllers connected by the router. In this case, the maximum number of signals transmitted simultaneously by router 01 and router 11 is equal to the total number of computing engines and first interconnection sending controllers connected by router 01 and router 11, and the maximum number of signals received simultaneously by routers 20, 21, 22, 30, 31, and 32 is equal to the total number of storage controllers and first interconnection receiving controllers connected by routers 20, 21, 22, 30, 31, and 32.

[0144] In considering the routing direction, the signals transmitted by the routing channel connecting the router 11 and the router 21 in the first transmission direction are transmitted by the router 01 and the router 11 and received by the routers 20, 21, 22, 30, 31, 32, so the maximum number of signals simultaneously transmitted by the routing channel connecting the router 11 and the router 21 in the first transmission direction is actually equal to the smaller one of the maximum number of signals simultaneously transmitted by the router 01 and the router 11 and the maximum number of signals simultaneously received by the routers 20, 21, 22, 30, 31, 32. The reason is that if the maximum number of signals simultaneously transmitted by the router 01 and the router 11 is smaller than the maximum number of signals simultaneously received by the routers 20, 21, 22, 30, 31, 32, even if the maximum number of signals simultaneously transmitted by the routing channel connecting the router 11 and the router 21 in the first transmission direction is set to a value greater than the maximum number of signals simultaneously transmitted by the router 01 and the router 11, the maximum number of signals actually simultaneously transmitted cannot reach the set value, resulting in bandwidth waste. Similarly, if the maximum number of signals simultaneously transmitted by the router 01 and the router 11 is greater than the maximum number of signals simultaneously received by the routers 20, 21, 22, 30, 31, 32, even if the maximum number of signals simultaneously transmitted by the routing channel connecting the router 11 and the router 21 in the first transmission direction is set to a value greater than the maximum number of signals simultaneously received by the routers 20, 21, 22, 30, 31, 32, the maximum number of signals actually simultaneously received cannot reach the set value, and the signals will occupy the bandwidth for a long time, resulting in that the bandwidth cannot be released.

[0145] Therefore, it is more appropriate to set the maximum number of signals simultaneously transmitted by the routing channel connecting the router 11 and the router 21 in the first transmission direction to the smaller one of the maximum number of signals simultaneously transmitted by the router 01 and the router 11 and the maximum number of signals simultaneously received by the routers 20, 21, 22, 30, 31, 32. That is, the maximum number of signals simultaneously transmitted by the routing channel connecting the router in the i-th row and the j-th column and the router in the i-th row and the j+1-th column in the first transmission direction is equal to the first number, wherein the first number is the minimum of the second number and the third number, the second number is the total number of storage controllers and first interconnection receiving controllers connected by the routers in the j+1-th to m-th columns (that is, the maximum number of signals simultaneously received by the routers in the j+1-th to m-th columns), and the third number is the total number of computing engines and first interconnection transmitting controllers connected by the routers in the i-th row and the 1st to j-th columns (that is, the maximum number of signals simultaneously transmitted by the routers in the i-th row and the 1st to j-th columns). In this case, neither bandwidth waste nor bandwidth release failure occurs.

[0146] Similarly, when the routing channel connecting the router in the i-th row and the j-th column and the router in the i-th row and the j+1-th column transmits a signal in the second transmission direction, the destination router of the signal can be any one of the routers in the 1st to j-th columns, and the source router of the signal can be any one of the routers in the j+1-th to m-th columns in the i-th row.

[0147] Referring to FIG. 9, when the routing channel between the router 11 and the router 21 transmits a signal in the second transmission direction from right to left, since it is already in the horizontal direction, the destination router of the signal can be any one of the routers 00, 01, 02, 10, 11, 12 to the left of the router 21. The source router of the signal can only be any one of the routers 21 and 31 to the right of the router 11 and in the same row as the router 11.

[0148] In this case, the maximum number of signals simultaneously transmitted by the routers 21 and 31 is equal to the total number of the computation engines and the first interconnect transmission controllers connected to the routers 21 and 31, and the maximum number of signals simultaneously received by the routers 00, 01, 02, 10, 11, 12 is equal to the total number of the storage controllers and the first interconnect reception controllers connected to the routers 00, 01, 02, 10, 11, 12.

[0149] The signal transmitted in the second transmission direction by the routing channel connecting the router 11 and the router 21 is transmitted by the routers 21 and 31 and received by the routers 00, 01, 02, 10, 11, 12, and thus the maximum number of signals simultaneously transmitted in the second transmission direction by the routing channel connecting the router 11 and the router 21 is actually equal to the smaller one of the maximum number of signals simultaneously transmitted by the routers 21 and 31 and the maximum number of signals simultaneously received by the routers 00, 01, 02, 10, 11, 12. That is, the maximum number of signals simultaneously transmitted in the second transmission direction by the routing channel connecting the router in the i-th row and the j-th column and the router in the i-th row and the j+1-th column is equal to the fourth number, where the fourth number is the smaller one of the fifth number and the sixth number, the fifth number is the total number of the storage controllers and the first interconnect reception controllers connected to the routers in the 1st to j-th columns (i.e., the maximum number of signals simultaneously received by the routers in the 1st to j-th columns), and the sixth number is the total number of the computation engines and the first interconnect transmission controllers connected to the routers in the j+1-th to m-th columns in the i-th row (i.e., the maximum number of signals simultaneously transmitted by the routers in the j+1-th to m-th columns in the i-th row).

[0150] When the routing channel connecting the router in the i-th row and the j-th column and the router in the (i+1)-th row and the j-th column transmits a signal in the third transmission direction, the destination router of the signal can be any one of the routers in the j-th column from the (i+1)-th to the n-th row, and the source router of the signal can be any one of the routers in the first to the i-th row.

[0151] Referring to FIG. 9, when the routing channel between the router 11 and the router 10 transmits a signal in the third transmission direction from top to bottom, the destination router of the signal can only be the router 10 because the third transmission direction is already a vertical direction. The source router of the signal can be any one of the routers 01, 02, 11, 12, 21, 22, 31, 32.

[0152] In this case, the maximum number of signals simultaneously transmitted by the routers 01, 02, 11, 12, 21, 22, 31, 32 is equal to the total number of the computation engines and the first interconnect transmission controllers connected to the routers 01, 02, 11, 12, 21, 22, 31, 32, and the maximum number of signals simultaneously received by the router 10 is equal to the total number of the storage controllers and the first interconnect reception controllers connected to the router 10.

[0153] The signal transmitted in the third transmission direction by the routing channel between the router 11 and the router 10 is transmitted by the routers 01, 02, 11, 12, 21, 22, 31, 32 and received by the router 10, and thus the maximum number of signals simultaneously transmitted in the third transmission direction by the routing channel between the router 11 and the router 10 is actually equal to the smaller one of the maximum number of signals simultaneously transmitted by the routers 01, 02, 11, 12, 21, 22, 31, 32 and the maximum number of signals simultaneously received by the router 10. That is, the maximum number of signals simultaneously transmitted in the third transmission direction by the routing channel connecting the router in the i-th row and the j-th column and the router in the (i+1)-th row and the j-th column is equal to the seventh number, where the seventh number is the smaller one of the eighth number and the ninth number, the eighth number is the total number of the storage controllers and the first interconnect reception controllers connected to the routers in the j-th column from the (i+1)-th to the n-th row (i.e., the maximum number of signals simultaneously received by the routers in the j-th column from the (i+1)-th to the n-th row), and the ninth number is the total number of the computation engines and the first interconnect transmission controllers connected to the routers in the first to the i-th row (i.e., the maximum number of signals simultaneously transmitted by the routers in the first to the i-th row).

[0154] When the routing channel connecting the router in the i-th row and the j-th column and the router in the (i+1)-th row and the j-th column transmits a signal in the fourth transmission direction, the destination router of the signal can be any one of the routers in the j-th column from the first to the i-th row, and the source router of the signal can be any one of the routers in the (i+1)-th to the n-th row.

[0155] Referring to FIG. 9, when a signal transmitted in the fourth transmission direction from bottom to top in the routing channel between the router 11 and the router 10 is already in the vertical direction, the destination router of the signal can only be one of the router 11 and the router 12. The source router of the signal can be any one of the routers 00, 10, 20, 30.

[0156] In this case, the maximum number of signals simultaneously transmitted by the routers 00, 10, 20, 30 is equal to the total number of the computing engines and the first interconnection transmission controllers connected to the routers 00, 10, 20, 30, and the maximum number of signals simultaneously received by the routers 11 and 12 is equal to the total number of the storage controllers and the first interconnection reception controllers connected to the routers 11 and 12.

[0157] The signal transmitted in the fourth transmission direction in the routing channel between the router 11 and the router 10 is transmitted by the routers 00, 10, 20, 30 and received by the routers 11 and 12, and thus the maximum number of signals simultaneously transmitted in the fourth transmission direction in the routing channel between the router 11 and the router 10 is actually equal to the smaller one of the maximum number of signals simultaneously transmitted by the routers 00, 10, 20, 30 and the maximum number of signals simultaneously received by the routers 11 and 12. That is, the maximum number of signals simultaneously transmitted in the fourth transmission direction in the routing channel between the router in the i-th row and the j-th column and the router in the i+1-th row and the j-th column is equal to the tenth number, where the tenth number is the smaller one of the eleventh number and the twelfth number, the eleventh number is the total number of the storage controllers and the first interconnection reception controllers connected to the routers in the j-th column and from the 1st to the i-th row (i.e., the maximum number of signals simultaneously received by the routers in the j-th column and from the 1st to the i-th row), and the twelfth number is the total number of the computing engines and the first interconnection transmission controllers connected to the routers from the i+1-th to the n-th row (i.e., the maximum number of signals simultaneously transmitted by the routers from the i+1-th to the n-th row).

[0158] In this way, the routing width of each routing channel is optimally set in the case of first passing through the routing channel in the horizontal direction and then passing through the routing channel in the vertical direction, and neither bandwidth waste nor bandwidth release failure occurs.

[0159] In a possible implementation, the routing channel in the horizontal direction includes a first transmission direction from left to right and a second transmission direction from right to left, and the routing channel in the vertical direction includes a third transmission direction from top to bottom and a fourth transmission direction from bottom to top,

[0160] In the case of first passing through the routing channel in the vertical direction and then passing through the routing channel in the horizontal direction,

[0161] The maximum number of signals simultaneously transmitted in the first transmission direction through the routing channel connecting the router in the ith row and the router in the jth column and the router in the ith row and the router in the j+1th column is equal to the thirteenth number, wherein the thirteenth number is the minimum of the fourteenth number and the fifteenth number, the fourteenth number refers to the total number of storage controllers connected by the j+1th to mth routers in the ith row and the first interconnection receiving controller, and the fifteenth number refers to the total number of computing engines connected by the 1st to jth routers and the first interconnection sending controller;

[0162] The maximum number of signals simultaneously transmitted in the second transmission direction through the routing channel connecting the router in the ith row and the router in the jth column and the router in the ith row and the router in the j+1th column is equal to the sixteenth number, wherein the sixteenth number is the minimum of the seventeenth number and the eighteenth number, the seventeenth number refers to the total number of storage controllers connected by the 1st to jth routers in the ith row and the first interconnection receiving controller, and the eighteenth number refers to the total number of computing engines connected by the j+1th to mth routers and the first interconnection sending controller;

[0163] The maximum number of signals simultaneously transmitted in the third transmission direction through the routing channel connecting the router in the ith row and the router in the jth column and the router in the ith+1th row and the router in the jth column is equal to the nineteenth number, wherein the nineteenth number is the minimum of the twentieth number and the twenty-first number, the twentieth number refers to the total number of storage controllers connected by the ith+1th to n th routers and the first interconnection receiving controller, and the twenty-first number refers to the total number of computing engines connected by the 1st to ith routers in the jth column and the first interconnection sending controller;

[0164] The maximum number of signals simultaneously transmitted in the fourth transmission direction through the routing channel connecting the router in the ith row and the router in the jth column and the router in the ith+1th row and the router in the jth column is equal to the twenty-second number, wherein the twenty-second number is the minimum of the twenty-third number and the twenty-fourth number, the twenty-third number refers to the total number of storage controllers connected by the 1st to ith rows and the first interconnection receiving controller, and the twenty-fourth number refers to the total number of computing engines connected by the ith+1th to n th routers in the jth column and the first interconnection sending controller.

[0165] In the case of first passing through the routing channel in the vertical direction and then passing through the routing channel in the horizontal direction, the routing width setting mode of each routing channel is similar to the case of first passing through the routing channel in the horizontal direction and then passing through the routing channel in the vertical direction, as long as it does not cause bandwidth waste or bandwidth release, which will not be described here.

[0166] FIGS. 10-12 show schematic diagrams of the structure of a routing system according to embodiments of the present disclosure. FIGS. 10 and 11 take the case of first passing through the routing channel in the horizontal direction and then passing through the routing channel in the vertical direction as an example. FIG. 12 takes the case of first passing through the routing channel in the vertical direction and then passing through the routing channel in the horizontal direction as an example.

[0167] In one possible implementation, each router is connected with A computing engines, B storage controllers, C first interconnection sending controllers, and D first interconnection receiving controllers, and each routing channel connecting the jth column router and the j+1th column router is the same, and each routing channel connecting the ith row router and the ith+1 row router is the same.

[0168] When m=n, each routing channel in the routing system is the same.

[0169] For example, as shown in FIG. 10, in order to reduce the design complexity of each routing channel in the routing system, each router can be connected with A computing engines, B storage controllers, C first interconnection sending controllers, and D first interconnection receiving controllers. In this case, the total number of computing engines and first interconnection sending controllers connected by each router is the same and equal to the total number of storage controllers and first interconnection receiving controllers connected by each router. At this time, the maximum number of signals received by each router is equal to the maximum number of signals sent and equal to t.

[0170] As shown in FIG. 10, assuming t=1, for a routing system including an array of n rows and m columns of routers, in the case of routing channels in the horizontal direction first and routing channels in the vertical direction second, in the m-1 horizontal direction transmission channels in the same row, from the leftmost routing channel to the rightmost routing channel, the routing width in the first transmission direction from left to right is 1, 2, 3, …, m-2, m-1 in turn; from the rightmost routing channel to the leftmost routing channel, the routing width in the second transmission direction from right to left is 1, 2, …, m-3, m-2, m-1 in turn. In the n-1 vertical direction transmission channels in the same column, from the uppermost routing channel to the lowermost routing channel, the routing width in the third transmission direction from top to bottom is n-1, n-2, n-3, …, 2, 1 in turn; from the lowermost routing channel to the uppermost routing channel, the routing width in the third transmission direction from bottom to top is n-1, n-2, …, 3, 2, 1 in turn.

[0171] As can be seen in combination with FIG. 10, each routing channel connecting the jth column router and the j+1th column router is the same, and each routing channel connecting the ith row router and the ith+1 row router is the same. Therefore, the number of routing channel specifications required by the whole is m+n-2.

[0172] Similarly, in the case of routing channels in the vertical direction of the first approach and routing channels in the horizontal direction of the second approach, assuming t = 1, among the m-1 horizontal routing channels in the same row, from the leftmost routing channel to the rightmost routing channel, the routing width in the first transmission direction from left to right is m-1, m-2,..., 3, 2, 1 in turn; from the rightmost routing channel to the leftmost routing channel, the routing width in the second transmission direction from right to left is m-1, m-2,..., 3, 2, 1 in turn. Among the n-1 vertical routing channels in the same column, from the topmost routing channel to the bottommost routing channel, the routing width in the third transmission direction from top to bottom is 1, 2, 3,..., n-2, n-1 in turn; from the bottommost routing channel to the topmost routing channel, the routing width in the third transmission direction from bottom to top is 1, 2, 3,..., n-2, n-1 in turn.

[0173] At this time, each routing channel connecting the jth column of routers and the j+1th column of routers is the same, and each routing channel connecting the ith row of routers and the ith+1 row of routers is the same. Therefore, the number of routing channel specification types required by the whole is m+n-2.

[0174] When m = n, the ith routing channel of each row is equal to the (m-i)th routing channel of each column (for an example, see FIG. 11), and therefore the number of routing channel specification types required by the whole is m+n.

[0175] As shown in FIG. 11, assuming that in the routing system, m = n = 3, the total number of storage controllers and first interconnection receiving controllers connected by each router, and the total number of computing engines and interconnection sending controllers connected by each router are equal, and equal to 1.

[0176] For example, the routing system is designed in the way of first routing channel in horizontal direction and then routing channel in vertical direction, the specifications of the routing channels in the routing system can be: the routing channels connecting the routers 00 and 10, the routing channels connecting the routers 01 and 11, the routing channels connecting the routers 02 and 12 have the same specification, the maximum number of signals simultaneously transmitted in the first transmission direction is equal to 1, and the maximum number of signals simultaneously transmitted in the second transmission direction is equal to 2; the routing channels connecting the routers 10 and 20, the routing channels connecting the routers 11 and 21, the routing channels connecting the routers 12 and 22 have the same specification, the maximum number of signals simultaneously transmitted in the first transmission direction is equal to 2, and the maximum number of signals simultaneously transmitted in the second transmission direction is equal to 1; the routing channels connecting the routers 00 and 01, the routing channels connecting the routers 10 and 11, the routing channels connecting the routers 20 and 21 have the same specification, the maximum number of signals simultaneously transmitted in the third transmission direction is equal to 1, and the maximum number of signals simultaneously transmitted in the fourth transmission direction is equal to 2; the routing channels connecting the routers 01 and 02, the routing channels connecting the routers 11 and 12, the routing channels connecting the routers 21 and 22 have the same specification, the maximum number of signals simultaneously transmitted in the third transmission direction is equal to 2, and the maximum number of signals simultaneously transmitted in the fourth transmission direction is equal to 1.

[0177] In the case of first routing channel in vertical direction and then routing channel in horizontal direction, the specifications of each routing channel are shown in FIG. 12, which will not be repeated here.

[0178] In this way, the number of routing channel specifications is reduced, facilitating logic design and physical design, thereby reducing the design complexity of the routing system.

[0179] The structure design of the router is introduced below.

[0180] In a possible implementation, the router in the ith row and the jth column includes a plurality of sending interfaces and a plurality of receiving interfaces, where,

[0181] The maximum number of signals simultaneously transmitted in the first transmission direction of the routing channel connecting the router in the ith row and the jth column and the router in the ith row and the j+1th column is a first value, and the maximum number of signals simultaneously transmitted in the second transmission direction is a second value;

[0182] The maximum number of signals simultaneously transmitted in the third transmission direction of the routing channel connecting the router in the ith row and the jth column and the router in the i+1th row and the jth column is a third value, and the maximum number of signals simultaneously transmitted in the fourth transmission direction is a fourth value;

[0183] The maximum number of signals simultaneously transmitted in the first transmission direction by the routing channel connecting the router in the i-th row and the j-th column and the router in the i-th row and the j-1-th column is a fifth value, and the maximum number of signals simultaneously transmitted in the second transmission direction is a sixth value;

[0184] The maximum number of signals simultaneously transmitted in the third transmission direction by the routing channel connecting the router in the i-th row and the j-th column and the router in the i-1-th row and the j-th column is a seventh value, and the maximum number of signals simultaneously transmitted in the fourth transmission direction is an eighth value;

[0185] The plurality of sending interfaces includes a first sending interface of the first value, a third sending interface of the third value, a sixth sending interface of the sixth value, an eighth sending interface of the eighth value, and a ninth sending interface;

[0186] The plurality of receiving interfaces includes a second receiving interface of the second value, a fourth receiving interface of the fourth value, a fifth receiving interface of the fifth value, a seventh receiving interface of the seventh value, and a tenth receiving interface;

[0187] Each first sending interface and each second receiving interface are connected to the routing channel connecting the router in the i-th row and the j-th column and the router in the i-th row and the j+1-th column;

[0188] Each third sending interface and each fourth receiving interface are connected to the routing channel connecting the router in the i-th row and the j-th column and the router in the i+1-th row and the j-th column;

[0189] Each fifth receiving interface and each sixth sending interface are connected to the routing channel connecting the router in the i-th row and the j-th column and the router in the i-th row and the j-1-th column;

[0190] Each seventh receiving interface and each eighth sending interface are connected to the routing channel connecting the router in the i-th row and the j-th column and the router in the i-1-th row and the j-th column;

[0191] The ninth sending interface is connected to the storage controller connected to the router in the i-th row and the j-th column and the first interconnection receiving controller;

[0192] The tenth receiving interface is connected to the computing engine connected to the router in the i-th row and the j-th column and the first interconnection sending controller.

[0193] Taking FIG. 11 as an example, assuming that the router 11 is the router in the i-th row and the j-th column, the first value can be equal to 2, the second value is equal to 1, the third value can be equal to 1, the fourth value can be equal to 2, the fifth value can be equal to 1, the sixth value can be equal to 2, the seventh value can be equal to 2, and the eighth value can be equal to 1.

[0194] The router can connect the routing channels through the interfaces arranged thereon. In order to reduce the complexity of signal transmission, the interfaces on the router can be classified by type into sending interfaces and receiving interfaces. The sending interfaces are used for sending signals, and the receiving interfaces are used for receiving signals. Signals sent through different routing channels to the same router can be sent to different receiving interfaces on the router; signals sent by the same router to different routing channels can be sent from different sending interfaces on the router.

[0195] Still taking the router 11 as an example, since the first value is equal to 2, the router 11 needs to have the capability of sending two signals to the routing channel connecting the router 11 and the router 21 at the same time, and therefore the router 11 can include two first sending interfaces; similarly, since the third value is equal to 1, the sixth value is equal to 2, and the eighth value is equal to 1, the router 11 can include one third sending interface, two sixth sending interfaces, and one eighth sending interface.

[0196] Since the second value is equal to 1, the router 11 needs to have the capability of receiving signals from the routing channel connecting the router 11 and the router 21 at the same time, and therefore the router 11 can include one second receiving interface; similarly, since the fourth value is equal to 2, the fifth value is equal to 1, and the seventh value is equal to 2, the router 11 can include two fourth receiving interfaces, one fifth receiving interface, and two seventh receiving interfaces.

[0197] Since the number of the first sending interfaces is equal to the first value, and the number of the second receiving interfaces is equal to the second value, the first value and the second value are the routing width of the routing channel connecting the router in the i-th row and the j-th column and the router in the i-th row and the j+1-th column, and therefore the first sending interfaces and the second receiving interfaces can be connected to the routing channel connecting the router in the i-th row and the j-th column and the router in the i-th row and the j+1-th column. Similarly, the third sending interfaces and the fourth receiving interfaces can be connected to the routing channel connecting the router in the i-th row and the j-th column and the router in the i+1-th row and the j-th column; the fifth receiving interfaces and the sixth sending interfaces can be connected to the routing channel connecting the router in the i-th row and the j-th column and the router in the i-th row and the j-1-th column; and the seventh receiving interfaces and the eighth sending interfaces can be connected to the routing channel connecting the router in the i-th row and the j-th column and the router in the i-1-th row and the j-th column.

[0198] The ninth sending interface can be connected to the storage controller connected to the router in the i-th row and the j-th column and the first interconnection receiving controller; and the tenth receiving interface can be connected to the computing engine connected to the router in the i-th row and the j-th column and the first interconnection sending controller.

[0199] FIG. 13 shows an exemplary structural diagram of a router according to an embodiment of the present disclosure. FIG. 13 shows the structure of the router 11 in the routing system shown in FIG. 11.

[0200] As shown in FIG. 13, the router 11 can include a first sending interface R11, R12, a third sending interface R31, a sixth sending interface R61, R62, an eighth sending interface R81, a ninth sending interface R91, and a second receiving interface R21, a fourth receiving interface R41, R42, a fifth receiving interface R51, a seventh receiving interface R71, R72, and a tenth receiving interface R101.

[0201] In this case, the first sending interface R11, R12 and the second receiving interface R21 can be connected with a routing channel connecting the router 11 and the router 21. The third sending interface R31 and the fourth receiving interface R41, R42 can be connected with a routing channel connecting the router 11 and the router 10. The fifth receiving interface R51 and the sixth sending interface R61, R62 can be connected with a routing channel connecting the router 11 and the router 01. The seventh receiving interface R71, R72 and the eighth sending interface R81 can be connected with a routing channel connecting the router 11 and the router 12. The ninth sending interface R91 can be connected with a storage controller connected with the router 11 and a first interconnection receiving controller. The tenth receiving interface R101 can be connected with a computing engine connected with the router 11 and a first interconnection sending controller.

[0202] FIG. 14 shows an exemplary structural diagram of a router according to an embodiment of the present disclosure. FIG. 14 shows the structure of the router 11 in the routing system shown in FIG. 12.

[0203] In the case of first passing through a routing channel in the vertical direction and then passing through a routing channel in the horizontal direction, the setting mode of the number and type of interfaces of the router is similar to the case of first passing through a routing channel in the horizontal direction and then passing through a routing channel in the vertical direction. At this time, when the router 11 is the i-th row and j-th column router, the first value can be equal to 1, the second value equal to 2, the third value equal to 2, the fourth value equal to 1, the fifth value equal to 2, the sixth value equal to 1, the seventh value equal to 1, and the eighth value equal to 2. That is, as shown in FIG. 14, the router 11 can include 1 first sending interface K11, 2 second receiving interfaces K21, K22, 2 third sending interfaces K31, K32, 1 fourth receiving interface K41, 2 fifth receiving interfaces K51, K52, 1 sixth sending interface K61, 1 seventh receiving interface K71, 2 eighth sending interfaces K81, K82, 1 ninth sending interface K91, and 1 tenth receiving interface K101.

[0204] In this way, each router can realize lower signal transmission complexity through fewer interfaces.

[0205] The exemplary connection mode of the sending interface and the receiving interface in the router according to an embodiment of the present disclosure will be introduced below.

[0206] In a possible implementation, when the first routing channel is in the horizontal direction and the second routing channel is in the vertical direction,

[0207] Each second receiving interface is connected with each third sending interface, each sixth sending interface, each eighth sending interface, and a ninth sending interface.

[0208] Each fourth receiving interface is connected with each eighth sending interface and a ninth sending interface.

[0209] Each fifth receiving interface is connected with each first sending interface, each third sending interface, each sixth sending interface, and a ninth sending interface.

[0210] Each seventh receiving interface is connected with each third sending interface and a ninth sending interface.

[0211] A tenth receiving interface is connected with each sending interface.

[0212] In the case that the first routing channel is in the horizontal direction and the second routing channel is in the vertical direction, still taking FIG. 13 as an example, at this time, the first sending interface R11 and the first sending interface R12 can send signals to the router 20 / 21 / 22; the second receiving interface R21 can receive signals sent by the router 21. The third sending interface R31 can send signals to the router 10; the fourth receiving interface R41 and the fourth receiving interface R42 can receive signals sent by the router 00 / 10 / 20. The fifth receiving interface R51 can receive signals sent by the router 01, and the sixth sending interface R61 and the sixth sending interface R62 can send signals to the router 00 / 01 / 02. The seventh receiving interface R71 and the seventh receiving interface R72 can receive signals sent by the router 02 / 12 / 22, and the eighth sending interface R81 can send signals to the router 12. The ninth sending interface R91 can send signals to the storage controller connected with the router 11 and the first interconnection receiving controller, and the tenth receiving interface R101 can receive signals from the computing engine and the first interconnection sending controller connected with the router 11.

[0213] As can be seen from FIG. 13, in the case of the first routing channel passing through the horizontal direction and the second routing channel passing through the vertical direction, the signal sent by the router 21 and received by the second receiving interface R21 is only sent to the routers 10 / 11 / 12 / 00 / 01 / 02, and thus the second receiving interface R21 only needs to be connected to the third sending interface R31, the eighth sending interface R81, the ninth sending interface R91, and the sixth sending interface R61, R62. The signal sent by the routers 00 / 10 / 20 and received by the fourth receiving interface R41, R42 is only sent to the router 12 / 11, and thus the fourth receiving interface R41, R42 only needs to be connected to the eighth sending interface R81 and the ninth sending interface R91. The signal sent by the router 01 and received by the fifth receiving interface R51 is only sent to the routers 10 / 11 / 12 / 20 / 21 / 22, and thus the fifth receiving interface R51 only needs to be connected to the first sending interface R11, R12, the third sending interface R31, the eighth sending interface R81, and the ninth sending interface R91. The signal sent by the router 02 / 12 / 22 and received by the seventh receiving interface R71, R72 is only sent to the routers 10 / 11, and thus the seventh receiving interface R71, R72 only needs to be connected to the third sending interface R31 and the ninth sending interface R91. The signal sent by the computing engine and the first interconnection sending controller connected to the router 11 itself and received by the tenth receiving interface R101 can be sent to the routers 00 / 01 / 02 / 10 / 11 / 12 / 20 / 21 / 22, and thus the tenth receiving interface R101 needs to be connected to all the sending interfaces.

[0214] In this way, in the case of the first routing channel passing through the horizontal direction and the second routing channel passing through the vertical direction, the router can meet the demand for distributing / arbitrating the signal, ensure that the router has the maximum bandwidth, and avoid the resource redundancy problem caused by the general structure of the router.

[0215] The structure design of the other routers in FIG. 11 is similar to that of the router 11, and thus the specific structure of the other routers in FIG. 11 will not be described herein.

[0216] In a possible implementation, in the case of the first routing channel passing through the vertical direction and the second routing channel passing through the horizontal direction,

[0217] Each second receiving interface is connected to each sixth sending interface and ninth sending interface;

[0218] Each fourth receiving interface is connected to each first sending interface, each sixth sending interface, each eighth sending interface, and ninth sending interface;

[0219] Each fifth receiving interface is connected to each first sending interface and ninth sending interface;

[0220] Each seventh receiving interface connects each first sending interface, and each third sending interface, each sixth sending interface, ninth sending interface;

[0221] The tenth receiving interface connects each sending interface.

[0222] For example, in the case of routing channels in the vertical direction first and in the horizontal direction second, still taking FIG. 14 as an example, the first sending interface K11 can send signals to the router 21; the second receiving interface K21, K22 can receive signals sent by the routers 20 / 21 / 22. The third sending interface K31, K32 can send signals to the routers 00 / 10 / 20; the fourth receiving interface K41 can receive signals sent by the router 10. The fifth receiving interface K51, K52 can receive signals sent by the routers 00 / 01 / 02, and the sixth sending interface K61 can send signals to the router 01. The seventh receiving interface K71 can receive signals sent by the router 12, and the eighth sending interface K81, K82 can send signals to the routers 02 / 12 / 22. The ninth sending interface K91 can send signals to the storage controller and the first interconnection receiving controller connected to the router 11; and the tenth receiving interface K101 can receive signals from the computing engine and the first interconnection sending controller connected to the router 11.

[0223] As can be seen from FIG. 14, in the case of the first routing channel passing through the vertical direction and then the horizontal direction, the signals sent by the router 20 / 21 / 22 and received by the second receiving interface K21, K22 are only sent to the router 01 / 11, and thus the second receiving interface K21, K22 only needs to be connected to the sixth sending interface K61 and the ninth sending interface R91. The signals sent by the router 10 and received by the fourth receiving interface K41 are only sent to the router 01 / 11 / 21 / 02 / 12 / 22, and thus the fourth receiving interface K41 only needs to be connected to the first sending interface K11, the sixth sending interface K61, the eighth sending interface K81, K82 and the ninth sending interface K91. The signals sent by the router 00 / 01 / 02 and received by the fifth receiving interface K51, K52 are only sent to the router 11 / 21, and thus the fifth receiving interface K51, K52 only needs to be connected to the first sending interface K11 and the ninth sending interface K91. The signals sent by the router 12 and received by the seventh receiving interface K71 are only sent to the router 00 / 10 / 20 / 01 / 11 / 21, and thus the seventh receiving interface K71 only needs to be connected to the first sending interface K11, the third sending interface K31, K32, the sixth sending interface K61 and the ninth sending interface K91. The signals sent by the computing engine and the first interconnection sending controller connected to the router 11 itself and received by the tenth receiving interface K101 can be sent to the router 00 / 01 / 02 / 10 / 11 / 12 / 20 / 21 / 22, and thus the tenth receiving interface K101 needs to be connected to all the sending interfaces.

[0224] In this way, in the case of the first routing channel passing through the vertical direction and then the horizontal direction, the router can meet the demand for distributing / arbitrating signals, ensure that the router has the maximum bandwidth, and avoid the problem of resource redundancy caused by the general structure of the router.

[0225] The structures of the other routers in FIG. 12 are similar to that of the router 11, and thus the specific structures of the other routers in FIG. 12 will not be described here.

[0226] The present disclosure also provides an electronic device including at least one storage core and at least one computing core as described above. The electronic device can be a terminal device or a server, and the embodiments of the present disclosure do not limit the specific type of the electronic device.

[0227] The following describes an exemplary planning process for the number of computing cores and storage cores and the structure of the computing cores on the electronic device.

[0228] Step 1: Plan the number of computing cores and the computing power of each computing core according to the computing power demand provided by the user, and plan the number of storage cores and the memory bandwidth of each storage core according to the memory bandwidth demand provided by the user.

[0229] Step 2, for each compute die, determine the number of compute engines and first interconnect sending controllers according to the compute power of the compute die. For each memory die, determine the number of storage arrays on the memory die and the memory bandwidth of each storage array according to the memory bandwidth of the memory die. For each compute die, determine the memory dies connected to the compute die, determine the number of storage controllers on the compute die according to the number of storage arrays of the memory dies, and the interface bandwidth of each storage controller is equal to the memory bandwidth of the corresponding storage array. And determine the interface bandwidth of each compute engine, first interconnect sending controller and first interconnect receiving controller on the compute die, so that the interface bandwidth of each compute engine, each storage controller, each first interconnect sending controller and each first interconnect receiving controller is equal.

[0230] Step 3, determine the number of routers in the routing system, so that any one router satisfies the condition of connecting A compute engines, B storage controllers, C first interconnect sending controllers and D first interconnect receiving controllers, and A+C is equal to B+D.

[0231] Step 4, according to the physical layout of the compute engines, storage controllers and other devices on the compute die, determine the number of rows and columns of the router array included in the routing system.

[0232] Step 5, according to the number of rows and columns of the router array, determine the maximum number of signals transmitted simultaneously in both transmission directions of each routing channel when the maximum number of signals transmitted and received simultaneously by each router is 1, and an example can be referred to the related description of FIG. 11 and FIG. 12.

[0233] Step 6, according to the number of compute engines and first interconnect sending controllers (or the number of storage controllers and first interconnect receiving controllers) actually connected to each router, adjust the maximum number of signals transmitted simultaneously in both transmission directions of the routing channel (i.e. routing width) obtained in step 5. When the number of compute engines and first interconnect sending controllers (or the number of storage controllers and first interconnect receiving controllers) actually connected to a certain router is 1, the routing width of the routing channel involving the router does not need to be adjusted; if the number of compute engines and first interconnect sending controllers (or the number of storage controllers and first interconnect receiving controllers) actually connected to a certain router is not 1, then the routing width of the routing channel involving the router is adjusted.

[0234] FIG. 15 and FIG. 16 show an example of adjusting the routing width of the routing channel according to an embodiment of the present disclosure.

[0235] For example, on the basis of the example of Fig. 11, assume that the total number of computing engines and first interconnect send controllers of access router 01, 22 is 2, while the total number of computing engines and first interconnect send controllers of access router 00, 12 is 0. At this time, adjusting the routing widths of the routing channels involving routers 01, 22 and routers 00, 12 (assuming in the order of first the horizontal channels and then the vertical channels), can be as shown in Fig. 15, where,

[0236] the routing width of the routing channel connecting router 12 and router 22 is adjusted to 1 in the first transmission direction from left to right, and to 2 in the second transmission direction from right to left;

[0237] the routing width of the routing channel connecting router 01 and router 11 is adjusted to 2 in the first transmission direction from left to right;

[0238] the routing width of the routing channel connecting router 11 and router 21 is adjusted to 3 in the first transmission direction from left to right;

[0239] the routing width of the routing channel connecting router 00 and router 10 is adjusted to 0 in the first transmission direction from left to right;

[0240] the routing width of the routing channel connecting router 10 and router 20 is adjusted to 1 in the first transmission direction from left to right.

[0241] For another example, on the basis of the example of Fig. 11, assume that the total number of computing engines and first interconnect send controllers of access router 11 is 2, while the total number of computing engines and first interconnect send controllers of access router 00 is 0. At this time, adjusting the routing widths of the routing channels involving routers 11 and routers 00 (assuming in the order of first the horizontal channels and then the vertical channels), can be as shown in Fig. 16, where,

[0242] the routing width of the routing channel connecting router 01 and router 11 is adjusted to 3 in the second transmission direction from right to left;

[0243] the routing width of the routing channel connecting router 11 and router 21 is adjusted to 3 in the first transmission direction from left to right;

[0244] the routing width of the routing channel connecting router 00 and router 10 is adjusted to 0 in the first transmission direction from left to right;

[0245] the routing width of the routing channel connecting router 10 and router 20 is adjusted to 1 in the first transmission direction from left to right.

[0246] Step 7, determine the number of sending interfaces and receiving interfaces of the router according to the routing width of the routing channel connected by each router in two directions, the position of the router in the router array, and the sequence of the horizontal channel and the vertical channel, and determine which receiving interfaces are connected by each sending interface, see the related description of FIG. 13 and FIG. 14 for examples.

[0247] Step 8, determine the number of transmission channels between the computing core and the storage core connected thereto.

[0248] Step 9, set the position of the router accessing the first interconnection receiving controller and the first interconnection sending controller at the edge of the computing core, so as to be connected with the second interconnection receiving controller and the second interconnection sending controller of other computing cores.

[0249] It can be understood that if the total number of computing engines and first interconnection receiving controllers connected by most of the routers in the routing system is the same in actual application, it is more convenient to plan the routing channel according to steps 5 and 6; if the total number of computing engines and first interconnection receiving controllers connected by most of the routers is the same, steps 5 and 6 can also be replaced by step 10 to reduce the number of adjustments.

[0250] Step 10, determine the maximum number of signals simultaneously sent and received by each router and the maximum number of signals simultaneously transmitted in two transmission directions of each routing channel (i.e., routing width) according to the number of computing engines and first interconnection sending controllers (or the number of storage controllers and first interconnection receiving controllers) actually connected by each router and the number of rows and columns of the router array.

[0251] The embodiments of the present disclosure propose a connection mode of 3D stacking of computing cores and storage cores according to the physical distribution characteristics of storage resources and computing engines. The beneficial effects are as follows:

[0252] I. Occupying limited silicon resources to meet the maximum bandwidth throughput of electronic devices and meet the bandwidth demand when accessing storage at peak in the worst case;

[0253] II. The design of the routing channel and the router ensures that the path of the memory access operation from the computing engine to the storage array is shorter, minimizing the memory access delay;

[0254] III. There is a uniform specification of routing channels, so that the number of specifications of interconnection channels and routing channels in the routing system is limited, facilitating unified modular design and improving design reusability;

[0255] IV. By using the first interconnection receiving controller and the first interconnection sending controller, the expansion of computing power and storage resources can be conveniently realized.

[0256] The electronic device adopting the embodiment of the present disclosure can obtain greater memory bandwidth, lower memory power consumption, smaller chip area, and improved overall computing power performance and memory performance with less resource cost.

[0257] Compared with the on-chip interconnection subsystem of the full interconnection architecture, the routing system of the present disclosure avoids the huge wiring resource overhead caused by the long-distance routing of all memory buses converging in a certain area for centralized arbitration and then being distributed to each place for storage operation.

[0258] Meanwhile, aiming at the problem that centralized arbitration makes all memory access delays very large, the routing system of the present disclosure adopts an array form. The memory access operation of the computing engine to the near-distance storage array can be completed with the shortest distance, and the average memory access delay is greatly reduced.

[0259] Compared with the full interconnection architecture of centralized arbitration, the average transmission distance of the routing system of the present disclosure is shortened, and the effect of saving power consumption is also very obvious.

[0260] The interconnection line scale of the full interconnection architecture of centralized arbitration is very large, causing congestion of physical wiring, low area utilization rate, waste of valuable silicon area, and increase of chip cost. The routing system of the embodiment of the present disclosure is distributed, which can distribute the routing channels and routers to different areas on the chip, avoid congestion of wiring, and improve the area utilization rate.

[0261] Compared with the traditional 2D mesh interconnection architecture, the routing system of the present disclosure has an independent dedicated line for providing the routing channel between the computing engine and the storage array according to the physical layout of the storage array and the computing engine, which can meet the peak bandwidth requirement in the worst case and avoid the routing system from becoming the bottleneck of memory access.

[0262] The electronic device of the embodiment of the present disclosure can greatly improve the memory efficiency and reduce the power consumption, which helps the computing power system to alleviate the problems of memory wall and power wall.

[0263] The computer program product of the second aspect can include a computer readable storage medium. The computer readable storage medium can include instructions. The instructions can include one or both of: instructions for causing a computer to enable a user equipment device to receive a configuration message from a base station, the configuration message comprising an indication of a set of one or more parameters for a first type of hybrid automatic repeat request process, the first type of hybrid automatic repeat request process being associated with a first type of data; and instructions for causing a computer to enable a user equipment device to receive a configuration message from a base station, the configuration message comprising an indication of a set of one or more parameters for a first type of hybrid automatic repeat request process, the first type of hybrid automatic repeat request process being associated with a first type of data.

[0264] Embodiments of the present disclosure have been described above, with the understanding that these embodiments are exemplary only, and are not restrictive, and are not limited to the disclosed embodiments. Many modifications and changes to this disclosure would be apparent to those of ordinary skill in the art. The scope of the technology disclosed is not to be limited by the specific illustrative embodiments presented above, but only by the claims that follow. The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting.

Claims

1. A computing core particle, characterized in that: The computing core is arranged on the first plane and connected to the storage core on the second plane. The storage core includes a plurality of storage arrays. The computing core comprises a plurality of computing engines, a plurality of storage controllers, and a routing system connecting the computing engines and the storage controllers, each storage controller being connected to a storage array, the number of transmission channels between each pair of connected storage controllers and storage arrays being greater than a first threshold, and the connection direction intersecting the first plane; The computing engine is used to generate a memory access command and transmit it to the routing system, wherein the memory access command includes a memory access type and a memory access address; The routing system is configured to transmit the memory access command to a memory controller connected to the memory array in response to receiving the memory access command and the memory access address being the address of the memory array connected to the computing core; The storage controller is configured to access the connected storage array according to the memory access type in response to receiving the memory access command.

2. The computing core particle according to claim 1, characterized in that The computing core is connected to the storage core via hybrid bonding, microbumps, or through-silicon vias.

3. The computing core particle according to claim 1, characterized in that A connection direction between the storage controller and the storage array is perpendicular to the first plane.

4. The computing core particle according to claim 1, wherein: The computing core is set in an electronic device, and the electronic device includes multiple computing cores. The computing core also includes at least one first interconnection receiving controller and at least one first interconnection sending controller connected to the routing system. Each first interconnection receiving controller is connected to a second interconnection sending controller included in another computing core, and each first interconnection sending controller is connected to a second interconnection receiving controller included in another computing core. The routing system is further configured to, in response to receiving the memory access command and the memory access address being an address of a storage array connected to another computing core particle, transmit the memory access command to a first interconnection sending controller in the computing core particle to which the memory access command belongs and connected to the other computing core particle; The first interconnection sending controller is used for transmitting the memory access command to the second interconnection receiving controller of the other computing core particle in response to receiving the memory access command from the routing system; The first interconnect receiving controller is configured to transmit the memory access command to the routing system in response to receiving the memory access command from the second interconnect sending controller in the other computing core.

5. The computing core particle according to claim 4, characterized in that The routing system includes a router array of n rows and m columns and a plurality of routing channels connecting adjacent routers, wherein the router in the i-th row and j-th column is connected to A computing engines, B storage controllers, C first interconnection sending controllers, and D first interconnection receiving controllers, where n and m are positive integers, A, B, C, and D are integers greater than or equal to 0, 1≤i≤n, and 1≤j≤m; The computing engine is used to transmit the memory access command to a router connected to the computing engine in the routing system; In the routing system, the router connected to the computing engine is used to transmit the memory access command to the storage controller connected to the storage array.

6. The computing core particle according to claim 5, characterized in that The routing system includes a first router, a second router and a third router, The first router is connected to the computing engine and the storage controller connected to the storage array, and is used to directly transmit the memory access command to the storage controller connected to the storage array; The second router is connected to the computing engine, and the third router is connected to the storage controller connected to the storage array. The second router is used to transmit the memory access command to the third router through the routing channel, and the third router is used to transmit the memory access command to the storage controller connected to the storage array.

7. The computing core particle according to claim 5, characterized in that In response to the two routers connected by a single routing channel being in the same row, the routing direction of the routing channel is horizontal; in response to the two routers connected by a single routing channel being in the same column, the routing direction of the routing channel is vertical. In response to the fact that signal transmission between different routers needs to pass through horizontal routing channels and vertical routing channels, the signal first passes through the horizontal routing channel and then through the vertical routing channel, or first passes through the vertical routing channel and then through the horizontal routing channel, the signal includes the memory access command and the data to be written to the storage array, or includes the data read from the storage array.

8. The computing core particle according to claim 7, characterized in that: Each router is connected to A computing engines, B storage controllers, C first interconnect sending controllers, and D first interconnect receiving controllers. At the same time, each routing channel connecting the j-th column router and the j+1-th column router is the same, and each routing channel connecting the i-th row router and the i+1-th row router is the same. When m=n, each routing channel in the routing system is the same.

9. The computing core particle according to claim 7, characterized in that: The horizontal routing channel includes a first transmission direction from left to right and a second transmission direction from right to left. The vertical routing channel includes a third transmission direction from top to bottom and a fourth transmission direction from bottom to top. When passing through the horizontal routing channel first and then the vertical routing channel, The maximum number of signals that can be simultaneously transmitted in the first transmission direction of the routing channel connecting the router in the i-th row and j-th column to the router in the i-th row and j+1-th column is equal to the first number; the first number is the minimum of the second number and the third number, the second number refers to the total number of storage controllers and first interconnected receiving controllers connected to the routers in the j+1-th to m-th columns, and the third number refers to the total number of computing engines and first interconnected sending controllers connected to the routers in the i-th row and j-th columns; The maximum number of signals that can be simultaneously transmitted in the second transmission direction of the routing channel connecting the router in the i-th row and j-th column to the router in the i-th row and j+1-th column is equal to the fourth number, where the fourth number is the minimum of the fifth number and the sixth number, the fifth number refers to the total number of storage controllers and first interconnect receiving controllers connected to the routers in the 1st to j-th columns, and the sixth number refers to the total number of computing engines and first interconnect sending controllers connected to the routers in the i-th row and j+1-m-th columns; The maximum number of signals simultaneously transmitted in the third transmission direction of the routing channel connecting the router in the i-th row and j-th column and the router in the i+1-th row and j-th column is equal to the seventh number, where the seventh number is the minimum of the eighth number and the ninth number, the eighth number refers to the total number of storage controllers and first interconnect receiving controllers connected to the routers in the i+1-th to n-th rows in the j-th column, and the ninth number refers to the total number of computing engines and first interconnect sending controllers connected to the routers in rows 1 to i; The maximum number of signals simultaneously transmitted in the fourth transmission direction of the routing channel connecting the router in the i-th row and j-th column and the router in the (i+1)-th row and j-th column is equal to the tenth number, where the tenth number is the minimum of the eleventh number and the twelfth number, the eleventh number refers to the total number of storage controllers and first interconnect receiving controllers connected to the routers from the 1st to the i-th column, and the twelfth number refers to the total number of computing engines and first interconnect sending controllers connected to the routers in the (i+1)-nth rows.

10. The computing core particle according to claim 7, characterized in that: The horizontal routing channel includes a first transmission direction from left to right and a second transmission direction from right to left. The vertical routing channel includes a third transmission direction from top to bottom and a fourth transmission direction from bottom to top. When the vertical routing channel is passed first and then the horizontal routing channel is passed, The maximum number of signals simultaneously transmitted in the first transmission direction of the routing channel connecting the router in the i-th row and j-th column to the router in the i-th row and j+1-th column is equal to the thirteenth number, where the thirteenth number is the minimum of the fourteenth and fifteenth numbers, the fourteenth number is equal to the total number of storage controllers and first interconnect receiving controllers connected to the routers in the i-th row from j+1 to m, and the fifteenth number is equal to the total number of computing engines and first interconnect sending controllers connected to the routers in the 1st to j-th columns; The maximum number of signals simultaneously transmitted in the second transmission direction of the routing channel connecting the router in the i-th row and j-th column and the router in the i-th row and j+1-th column is equal to the sixteenth number, where the sixteenth number is the minimum of the seventeenth number and the eighteenth number, the seventeenth number refers to the total number of storage controllers and first interconnect receiving controllers connected to the routers from the 1st to the jth in the i-th row, and the eighteenth number refers to the total number of computing engines and first interconnect sending controllers connected to the routers from the j+1st to the mth columns; The maximum number of signals simultaneously transmitted in the third transmission direction of the routing channel connecting the router in the i-th row and j-th column and the router in the i+1-th row and j-th column is equal to the nineteenth number, where the nineteenth number is the minimum of the twentieth number and the twenty-first number, the twentieth number refers to the total number of storage controllers and first interconnect receiving controllers connected to the routers in rows (i+1) to (n), and the twenty-first number refers to the total number of computing engines and first interconnect sending controllers connected to the routers in rows (i) to (i) of the j-th column; The maximum number of signals simultaneously transmitted in the fourth transmission direction of the routing channel connecting the router in the i-th row and j-th column and the router in the (i+1)-th row and j-th column is equal to the twenty-second number, where the twenty-second number is the minimum of the twenty-third number and the twenty-fourth number, the twenty-third number refers to the total number of storage controllers and first interconnected receiving controllers connected to the routers in rows 1 to i, and the twenty-fourth number refers to the total number of computing engines and first interconnected sending controllers connected to the routers in rows (i+1) to (n) in the j-th column.

11. The computing core particle according to claim 7, characterized in that: The router in row i and column j includes multiple sending interfaces and multiple receiving interfaces, where: The maximum number of signals that can be simultaneously transmitted in the first transmission direction of the routing channel connecting the router at row i and column j and the router at row i and column j+1 is a first value, and the maximum number of signals that can be simultaneously transmitted in the second transmission direction is a second value; The maximum number of signals simultaneously transmitted in the third transmission direction by the routing channel connecting the router in the i-th row and j-th column and the router in the (i+1)-th row and j-th column is a third value, and the maximum number of signals simultaneously transmitted in the fourth transmission direction is a fourth value; The maximum number of signals that can be simultaneously transmitted in the first transmission direction of the routing channel connecting the router at row i and column j and the router at row i and column j-1 is the fifth value, and the maximum number of signals that can be simultaneously transmitted in the second transmission direction is the sixth value; The maximum number of signals simultaneously transmitted in the third transmission direction by the routing channel connecting the router at row i and column j and the router at row i-1 and column j is the seventh value, and the maximum number of signals simultaneously transmitted in the fourth transmission direction is the eighth value; The multiple sending interfaces include a first sending interface for the first value, a third sending interface for the third value, a sixth sending interface for the sixth value, an eighth sending interface for the eighth value, and a ninth sending interface; The multiple receiving interfaces include a second receiving interface of the second value, a fourth receiving interface of the fourth value, a fifth receiving interface of the fifth value, a seventh receiving interface of the seventh value, and a tenth receiving interface; Each first sending interface and each second receiving interface is connected to the routing channel connecting the router in the i-th row and j-th column and the router in the i-th row and j+1-th column; Each third sending interface and each fourth receiving interface is connected to the routing channel connecting the router in the i-th row and j-th column and the router in the (i+1)-th row and j-th column; Each fifth receiving interface and each sixth sending interface are connected to the routing channel connecting the router in the i-th row and j-th column and the router in the i-th row and j-1-th column; Each seventh receiving interface and each eighth sending interface are connected to the routing channel connecting the router in the i-th row and j-th column and the router in the i-1-th row and j-th column; The ninth sending interface is connected to the storage controller connected to the router in the i-th row and the j-th column and the first interconnected receiving controller; The tenth receiving interface is connected to the computing engine connected to the router in the i-th row and the j-th column and the first interconnected sending controller.

12. The computing core particle according to claim 11, characterized in that: When passing through the horizontal routing channel first and then the vertical routing channel, Each second receiving interface is connected to each third sending interface, each sixth sending interface, each eighth sending interface, and the ninth sending interface; Each fourth receiving interface is connected to each eighth sending interface and ninth sending interface; Each fifth receiving interface is connected to each first sending interface, each third sending interface, each sixth sending interface, and the ninth sending interface; Each seventh receiving interface is connected to each third sending interface and the ninth sending interface; The tenth receiving interface is connected to each sending interface.

13. The computing core particle according to claim 11, characterized in that: When the vertical routing channel is passed first and then the horizontal routing channel is passed, Each second receiving interface is connected to each sixth sending interface and the ninth sending interface; Each fourth receiving interface is connected to each first sending interface, each sixth sending interface, each eighth sending interface, and the ninth sending interface; Each fifth receiving interface is connected to each first sending interface and the ninth sending interface; Each seventh receiving interface is connected to each first sending interface, each third sending interface, each sixth sending interface, and the ninth sending interface; The tenth receiving interface is connected to each sending interface.

14. An electronic device, characterized in that: The device comprises at least one storage core particle and at least one computing core particle according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Data processing device and method, chip, processor, equipment and storage medium

    CN111080510A

  • Permutated ring network interconnected computing architecture

    CN113544658A

  • Three-dimensional chip and computing system

    CN113656346A

  • Many-core processing device, data processing method and equipment, and medium

    CN114721993A

  • Reconfigurable 3D chip and integration method thereof

    CN116246963A

Cited By

  • Chip system, data transmission method and related equipment

    CN121919165A

  • Distributed memory access device and distributed memory access method suitable for three-dimensional stacked storage

    CN122262077A