3D stacked and interconnected processor, system and equipment
By decoupling computing logic and interconnection logic in the processor and adopting 3D stacking interconnection technology, the existing processors have solved the problem of insufficient winding resources and increased power consumption when increasing internal interconnection bandwidth, and achieved higher computing power and interconnection bandwidth.
Patent Information
- Application Number
- CN202411883650.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-27
AI Technical Summary
When increasing the internal interconnect bandwidth, existing processors face problems such as insufficient winding resources and increased power consumption due to increased data bit width, or increase clock frequency but limited by process limitations.
By decoupling the computing logic from the interconnection logic and splitting it on different chips, 3D stacking interconnection technology is adopted. The computing chip is responsible for data calculation, and the interconnection chip is responsible for routing and forwarding data, and 3D stacking interconnection is realized using TSV connection.
It realizes that while saving the area of computing chips and winding resources, it increases routing nodes, improves the computing power and interconnection bandwidth of the processor, and improves bandwidth utilization.
Smart Images

Figure CN120045510A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of processor architecture design, and particularly relates to a 3D stacked interconnected processor, system, and device. Background Art
[0002] With the continuous improvement of the computing performance of processors, new requirements have been put forward for the interconnection bandwidth within the processor. Currently, for the improvement of the interconnection bandwidth within the processor, mainly two methods are relied on: one is to increase the bit width of the transmitted data. However, this requires consuming a large amount of wiring resources. If the data bit width is too large, it will lead to insufficient wiring resources, and then cause problems such as difficult implementation, excessive area occupation, and increased power consumption; the other method is to increase the frequency of the interconnection clock within the processor, but this method is limited by the current processor process. Summary of the Invention
[0003] To solve at least one of the problems mentioned in the above background art, this application proposes a 3D stacked interconnected processor, system, and device.
[0004] According to the first aspect of this application, this application provides a 3D stacked interconnected processor, which includes an interconnection chip and at least one computing chip. Among them, the computing chip is used for data calculation, and the interconnection chip is used for routing and forwarding data within and / or between the computing chips.
[0005] Optionally, 3D stacked interconnection is performed between the computing chip and the interconnection chip. Further, 3D stacked interconnection can be achieved between the computing chip and the interconnection chip through TSV connection.
[0006] The 3D stacked interconnected processor proposed in this application can decouple the computing logic and the interconnection logic on a single processor and split them onto different chips. More computing units can be arranged in the saved area and wiring resources within the computing chip, thereby improving the computing power of the processor. And decoupling the interconnection logic onto a separate interconnection chip can increase the routing nodes, thereby enhancing the interconnection bandwidth.
[0007] Specifically, the function of data calculation is realized by arranging computing units within the computing chip, and routing nodes are arranged within the interconnection chip to realize the functions of routing and forwarding. Moreover, each computing unit within the computing chip is respectively connected to the routing nodes within the interconnection chip.
[0008] For the 3D stacked interconnected processor proposed in this application, since only computing units need to be arranged within the computing chip, without the need to arrange on-chip interconnection and on-chip storage, the area and wiring resources of the computing chip can be saved, and more computing units can be arranged within the computing chip that has saved area and wiring resources, which can improve the computing power of the processor.
[0009] In addition, since the redundant area on the interconnected chip can also be used to arrange the memory control logic and the arithmetic logic unit, the data transfer frequency between the computing chip and the memory chip can be reduced, and the bandwidth utilization rate of the processor can be improved.
[0010] An arithmetic logic unit is arranged in the interconnected chip for processing the data on the routing node.
[0011] By arranging an arithmetic logic unit on the interconnected chip in the present application, the data calculation function on the interconnected chip during data transmission can be realized.
[0012] The processor further includes at least one memory chip for data storage, and the interconnected chip is further used for routing and forwarding the data between the computing chip and the memory chip.
[0013] Optionally, 3D stacked interconnection is performed between the memory chip and the interconnected chip. Further, 3D stacked interconnection can be performed between the memory chip and the interconnected chip through TSV connection.
[0014] For the 3D stacked interconnected processor proposed in the present application, by decoupling the interconnection logic and the memory logic on a single processor and splitting them onto different chips, since only computing units need to be arranged in the computing chip without arranging on-chip interconnection and on-chip memory, the area and wiring resources of the computing chip can be saved, and by arranging more computing units in the computing chip that has saved area and wiring resources, the computing power of the processor can be improved; while decoupling the interconnection logic onto a separate interconnected chip can increase the routing nodes, thereby enhancing the interconnection bandwidth.
[0015] According to the second aspect of the present application, the present application further provides a processor system, which includes any one of the 3D stacked interconnected processors described in the first aspect of the present application, and the 3D stacked interconnected processors are interconnected through UCIe standard interconnection technology.
[0016] According to the third aspect of the present application, the present application further provides an electronic device, which includes the processor system described in the second aspect of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a schematic diagram of the 3D stacked interconnected processor proposed in the present application;
[0018] Figure 2 is a schematic diagram of the horizontal expansion of the processor in the 3D stacked interconnected processor proposed in the present application;
[0019] Figures 3 to 6It is a schematic diagram of different data flow directions under a 3D stacked interconnected processor proposed in this application. Detailed implementation mode
[0020] In order to make the objectives and features of this application more obvious and understandable, the technical solution will be described in detail below through embodiments in conjunction with the accompanying drawings.
[0021] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments.
[0022] Embodiment 1:
[0023] As shown in Figure 1 , a 3D stacked interconnected processor includes an interconnected chip and at least one computing chip. Among them, the computing chip is used for data calculation, and the interconnected chip is used for routing and forwarding data within and / or between computing chips. Optionally, 3D stacked interconnection is performed between the computing chip and the interconnected chip. Further, 3D stacked interconnection can be performed between the computing chip and the interconnected chip through TSV connection.
[0024] By way of example and not limitation, the processor includes at least 1 computing chip and 1 interconnected chip. It should be noted that different numbers of computing chips and interconnected chips can also be set according to the actual application scenario or application requirements, and the present invention does not limit this.
[0025] It should be noted that in the embodiments of this application, the computing chip can refer to a computing chiplet or a computing die. Similarly, the interconnected chip can refer to an interconnected chiplet or an interconnected die; the storage chip can refer to a storage chiplet or a storage die.
[0026] For the 3D stacked interconnected processor proposed in this application, by decoupling the computing logic and the interconnection logic on a single processor and splitting them onto different chips, more computing units can be arranged in the saved area and wiring resources within the computing chip, thereby improving the computing power of the processor. And decoupling the interconnection logic onto a separate interconnected chip can increase routing nodes, thereby improving the interconnection bandwidth.
[0027] Specifically, the function of data calculation is realized by arranging computing units within the computing chip, and routing nodes are arranged within the interconnected chip to realize the functions of routing and forwarding. Moreover, each computing unit within the computing chip is respectively connected to the routing nodes within the interconnected chip.
[0028] Optionally, the computing unit can be a traditional central processing unit, or a neural network processor for artificial intelligence computing, or an accelerator responsible for processing specialized tasks. The present invention does not limit this.
[0029] For the 3D stacked interconnected processor proposed in this application, since only computing units need to be arranged in the computing chip, without the need to arrange on-chip interconnection and on-chip storage, the area and wiring resources of the computing chip can be saved. By arranging more computing units in the computing chip with saved area and wiring resources, the computing power of the processor can be improved.
[0030] In addition, since the redundant area on the interconnection chip can also be used to arrange storage control logic and arithmetic logic units, the data transfer frequency between the computing chip and the storage chip can be reduced, and the bandwidth utilization rate of the processor can be improved.
[0031] An arithmetic logic unit is arranged in the interconnection chip to process the data on the routing node.
[0032] In this application, by arranging an arithmetic logic unit on the interconnection chip, the data calculation function on the interconnection chip during data transfer can be realized.
[0033] The processor further includes at least one storage chip for data storage, and the interconnection chip is also used for routing and forwarding data between the computing chip and the storage chip.
[0034] Optionally, 3D stacked interconnection is performed between the storage chip and the interconnection chip. Further, 3D stacked interconnection can be performed between the storage chip and the interconnection chip through TSV connection.
[0035] By way of example and not limitation, the processor includes at least 1 storage chip and 1 interconnection chip. It should be noted that different numbers of storage chips and interconnection chips can also be set according to the actual application scenario or application requirements. The present invention does not limit this.
[0036] Specifically, a storage unit is arranged in the storage chip to provide the function of data storage. A storage control logic is arranged in the interconnection chip. The storage control logic is connected to the routing node, and the storage unit in the storage chip is connected to the storage control logic in the interconnection chip.
[0037] The 3D stacked interconnected processor proposed in this application decouples the interconnection logic and storage logic on a single processor and splits them onto different chips. Since only computing units need to be arranged in the computing chip without on-chip interconnection and on-chip storage, the area and wiring resources of the computing chip can be saved. By arranging more computing units in the computing chip with saved area and wiring resources, the computing power of the processor can be improved. Decoupling the interconnection logic onto a separate interconnection chip can increase routing nodes, thus enhancing the interconnection bandwidth.
[0038] In an embodiment of this application, a processor system is further provided. The processor system includes any one of the above 3D stacked interconnected processors.
[0039] As Figure 2 shown, in the specific implementation process, the 3D stacked interconnected processor proposed in this application can use standard interconnection technologies such as UCIe to horizontally expand the processor, thereby forming a larger-scale processor system.
[0040] In an embodiment of this application, an electronic device is further provided. The electronic device includes the above processor system.
[0041] Embodiment 2:
[0042] Currently, for the improvement of the memory bandwidth in a processor, it mainly depends on improving the bandwidth of the memory itself. However, this method not only affects the integrity of system signals but also cannot keep up with the continuous improvement of the processor's computing performance.
[0043] This embodiment provides a 3D stacked interconnected processor. At least one storage unit is arranged in the storage chip of the processor, and the number of storage units in the storage chip can be arranged according to the actual required storage bandwidth.
[0044] Specifically, the storage chip provides the function of data storage by arranging storage units in the storage chip, and the storage units in the storage chip are connected to the storage control logic in the interconnection chip through TSVs.
[0045] Optionally, the storage unit can be SRAM or DRAM.
[0046] By decoupling the computing logic, interconnection logic, and storage logic on a single processor in this application, since only storage units need to be arranged in the storage chip and the number of storage units to be arranged can be selected according to the actual required storage bandwidth, the storage bandwidth and storage capacity of the processor can be effectively improved.
[0047] Embodiment 3:
[0048] In Figures 3 to 6Among them, PE represents the computing unit in the computing chip, MEM represents the storage unit in the storage chip, ALU represents the arithmetic logic unit in the interconnection chip, XP represents the routing node in the interconnection chip, and MEM CTRL represents the storage control logic in the interconnection chip.
[0049] In a preferred embodiment, as Figure 3 shown, for the 3D stacked interconnected processor proposed in this application, the data flow between different PEs inside the computing chip is as follows:
[0050] First, the request sent by the source PE in the computing chip is sent to the XP in the interconnected chip connected thereto through the TSV. Then, this XP will route the received request to the XP connected to the target PE according to the routing information. Next, the request is sent from the XP in the interconnected chip to the target PE in the computing chip through the TSV.
[0051] In a preferred embodiment, as Figure 4 shown, for the 3D stacked interconnected processor proposed in this application, the data flow from the PE in the computing chip to the MEM in the storage chip is as follows:
[0052] First, the request sent by the source PE in the computing chip is sent to the XP in the interconnected chip connected thereto through the TSV. Then, this XP will route the received request to the XP connected to the control logic of the target MEM according to the routing information, and after being processed by the MEM CTRL, it is sent to the target MEM in the storage chip through the TSV.
[0053] In a preferred embodiment, as Figure 5 shown, for the 3D stacked interconnected processor proposed in this application, the data flow of the data sent by the source PE in the computing chip and sent to the target PE in the computing chip after being processed by the ALU in the interconnected chip is as follows:
[0054] First, the request sent by the source PE on the computing chip is sent to the XP in the interconnected chip connected thereto through the TSV. Then, this XP will send the received request to the ALU for calculation according to the routing information, and route the calculated request to the XP connected to the target PE. Next, the request is sent from the XP in the interconnected chip to the target PE in the computing chip through the TSV.
[0055] In a preferred embodiment, as Figure 6 shown, for the 3D stacked interconnected processor proposed in this application, the data flow of the data sent by the source PE in the computing chip and sent to the MEM in the storage chip after being processed by the ALU in the interconnected chip is as follows:
[0056] First, the requests sent by the source PEs within the chip are transmitted through the TSVs to the XPs within the interconnected chip connected thereto. Then, the XPs will send the received requests to the ALUs for calculation according to the routing information, and route the requests after calculation to the XPs connected to the control logic of the target MEM. Finally, after being processed by the MEM CTRL, they are sent to the target MEM within the storage chip through the TSVs.
[0057] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
[0058] It should be noted that the serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments. And the term "comprising" in this text, "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, apparatus, article or method including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, apparatus, article or method. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, apparatus, article or method including that element.
Claims
1. A 3D stacked interconnect processor, characterized in that: The processor includes an interconnect chip and at least one computing chip; wherein the computing chip is used for data computing, and the interconnect chip is used for routing and forwarding data within the computing chip and / or between computing chips.
2. The 3D stacked interconnect processor according to claim 1, characterized in that: The computing chip is provided with computing units to realize the function of data computing, the interconnection chip is provided with routing nodes to realize the functions of routing and forwarding, and each computing unit in the computing chip is respectively connected to the routing nodes in the interconnection chip.
3. The 3D stacked interconnect processor according to claim 1, characterized in that: An arithmetic logic unit is arranged in the interconnect chip to process data on the routing node.
4. The 3D stacked interconnect processor according to claim 2, characterized in that: The processor also includes at least one memory chip, which is used for data storage. The interconnect chip is also used for routing and forwarding data between the computing chip and the memory chip.
5. The 3D stacked interconnect processor according to claim 3, characterized in that: The storage chip has storage units arranged in it to provide a data storage function, the interconnect chip has storage control logic arranged in it, the storage control logic is connected to the routing node, and the storage units in the storage chip are connected to the storage control logic in the interconnect chip.
6. The 3D stacked interconnect processor according to claim 5, characterized in that: The computing chip and the interconnection chip are connected by TSV to perform 3D stacking interconnection; the storage chip and the interconnection chip are connected by TSV to perform 3D stacking interconnection.
7. The 3D stacked interconnect processor according to claim 2, characterized in that: At least one storage unit needs to be arranged in the storage chip, and the number of storage units in the storage chip is determined according to the actual required storage bandwidth.
8. A processor system, characterized in that: The processor system comprises a 3D stacked interconnected processor as described in any one of claims 1 to 7.
9. The processor system according to claim 8, characterized in that The processor system includes at least two 3D stacked interconnected processors, and the 3D stacked interconnected processors are interconnected by using UCIe standard interconnection technology.
10. An electronic device, characterized in that: The electronic device comprises the processor system as claimed in claim 9 above.
Citation Information
Cited By
Processor verification system and method, electronic equipment and readable storage medium
CN121480403A