Chip design method and chip system

By integrating multiple DRAM controllers into the processor core and utilizing three-dimensional stacked storage technology, the problem of low data flow management efficiency in large-scale parallel data processing is solved, efficient data transmission and processing is achieved, and system performance is significantly improved.

CN120144534AActive Publication Date: 2025-06-13CORE ARK (SHANGHAI) INTEGRATED CIRCUIT CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510629574.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-13
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

In the large-scale parallel data processing scenario, it is difficult for the prior art to efficiently manage data flows, resulting in bus congestion and affect storage response speed and energy efficiency ratio.

Method used

By integrating the local DRAM controller, global DRAM controller and shared DRAM controller in the processor core, and vertically stacking DRAM chips through three-dimensional stacking storage technology, memory-shared data interaction between processor cores is realized and data transmission path is optimized.

Benefits of technology

Shorten the data path, improve the data exchange speed, significantly improve data throughput and overall system performance, and enhance the ability to process real-time data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144534A_ABST
    Figure CN120144534A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of semiconductors, and particularly relates to a chip design method and a chip system. The chip design method comprises the following steps of: integrating a local DRAM (Dynamic Random Access Memory) controller, a global DRAM controller and a shared DRAM controller on a processor core, and integrating a plurality of processor cores on a logic chip. The DRAM chips are vertically stacked on the logic chip, all the global DRAM controllers are connected with the shared DRAM storage space, and the shared DRAM controllers are connected with the corresponding private DRAM storage spaces respectively. The global DRAM controller is directly connected to the shared DRAM storage space, intermediate links are reduced, the data path is shortened, the data exchange speed is increased, the shared DRAM controller is directly connected with the private DRAM storage space of each processor core, and memory shared data interaction between the processor cores is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of semiconductor technology, and particularly relates to a chip design method and a chip system. Background Art

[0002] Three-dimensional DRAM technology greatly improves the storage density and data access speed by vertically stacking memory cells. However, in large-scale parallel data processing scenarios, how to efficiently manage data streams and avoid bus congestion has become a key factor restricting system performance.

[0003] Traditional data routing algorithms are unable to cope when dealing with highly concurrent accesses and high bandwidth requirements. Especially in a distributed bus architecture, the flexibility and efficiency of data routing directly affect the response speed and energy efficiency ratio of the entire storage system. Summary of the Invention

[0004] Aiming at the technical problem that the existing data transmission path has delays or congestion, which affects the storage response speed and energy efficiency ratio, the present invention aims to provide a chip design method and a chip system.

[0005] To solve the foregoing technical problems, a first aspect of the present invention provides a chip design method, which includes:

[0006] Integrating a local DRAM controller, a global DRAM controller, and a shared DRAM controller respectively connected to the local DRAM controller on a processor core, integrating a plurality of the processor cores on a logic chip, and connecting the local DRAM controllers to each other in the logic chip;

[0007] Vertically stacking a DRAM chip having a shared DRAM storage space and a plurality of private DRAM storage spaces on the logic chip by three-dimensional stacking storage technology, connecting each of the global DRAM controllers to the shared DRAM storage space, and connecting each of the shared DRAM controllers to their respective corresponding private DRAM storage spaces.

[0008] Optionally, in the chip design method as described above, the local DRAM controllers are connected to each other through a NoC structure, so that the local DRAM controller of each processor core is responsible for managing its local memory access requests and cooperates with the local controllers of other processor cores through the NoC.

[0009] Optionally, in the chip design method as described above, a logic layer is further integrated on the logic chip, and each of the local DRAM controllers is connected to the logic layer through an independent bus to form multiple parallel data transmission channels, allowing multiple data transmissions to occur simultaneously.

[0010] Optionally, in the chip design method described above, the logic layer is a CPU or an NPU.

[0011] Optionally, in the chip design method described above, the chip design method further includes:

[0012] Design a central arbiter, and the central arbiter is used to allocate the bandwidth resources of the bus.

[0013] Optionally, in the chip design method described above, the chip design method further includes:

[0014] Integrate a memory data prefetch model in the local DRAM controller, and the memory data prefetch model is used to predict future data access, and use the global DRAM controller or the shared DRAM controller to load data from the shared DRAM storage space or the private DRAM storage space to the cache in advance.

[0015] Optionally, in the chip design method described above, when the global DRAM controller is connected to the shared DRAM storage space, the connection is implemented by copper interconnection of the metal layer.

[0016] Optionally, in the chip design method described above, when the shared DRAM controller is connected to the corresponding private DRAM storage space, the connection is implemented by copper interconnection of the metal layer.

[0017] To solve the foregoing technical problems, a second aspect of the present invention provides a chip system, and the chip system includes:

[0018] A logic chip, the logic chip integrates a plurality of processor cores, and each processor core integrates a local DRAM controller, a global DRAM controller, and a shared DRAM controller. The local DRAM controller is respectively connected to the global DRAM controller and the shared DRAM controller, and the local DRAM controllers of each processor core are connected to each other;

[0019] A DRAM chip, the DRAM chip is vertically stacked on the logic chip, the DRAM chip has a shared DRAM storage space and a plurality of private DRAM storage spaces, the shared DRAM storage space is respectively connected to the global DRAM controller, and each private DRAM storage space is connected to a corresponding shared DRAM controller.

[0020] Optionally, in the chip system as described above, the local DRAM controllers are connected through the NoC structure, so that the local DRAM controller of each processor core is responsible for managing its local memory access requests and cooperates with the local controllers of other processor cores through the NoC.

[0021] Optionally, in the chip system as described above, the logic chip is further integrated with a logic layer, and each local DRAM controller is connected to the logic layer through an independent bus to form multiple parallel data transmission channels, allowing multiple data transmissions to occur simultaneously.

[0022] Optionally, in the chip system as described above, the logic layer is a CPU or an NPU.

[0023] Optionally, in the chip system as described above, the chip system further includes:

[0024] A central arbiter, which is used to allocate the bandwidth resources of the bus.

[0025] Optionally, in the chip system as described above, a memory data prefetch model is integrated in the local DRAM controller, and the memory data prefetch model is used to predict future data accesses and load data from the shared DRAM storage space or the private DRAM storage space to the cache in advance by using the global DRAM controller or the shared DRAM controller.

[0026] Optionally, in the chip system as described above, when the global DRAM controller is connected to the shared DRAM storage space, the connection is implemented by copper interconnection of the metal layer.

[0027] Optionally, in the chip system as described above, when the shared DRAM controller is connected to the corresponding private DRAM storage space, the connection is implemented by copper interconnection of the metal layer.

[0028] The positive and progressive effects of the present invention are as follows:

[0029] 1. In the present invention, when there is a local DRAM controller in the processor core, two DRAM controllers are added. Specifically, by directly connecting the global DRAM controller to the shared DRAM storage space, the intermediate links are reduced, the data path is shortened, and the data exchange speed is improved. By directly connecting the shared DRAM controller to the private DRAM storage spaces of each processor core, memory shared data interaction between the processor cores is achieved.

[0030] 2. The local DRAM controller of each processor core in the present invention is connected to the logic layer through an independent high-speed bus, forming multiple parallel data transfer channels, allowing multiple data transfers to occur simultaneously, and significantly improving data throughput.

[0031] 3. The present invention adopts a distributed memory management strategy. The local DRAM controllers of each processor core are responsible for managing their local memory access requests, reducing the pressure of global arbitration, and working in coordination with the local DRAM controllers of other processor cores through NoC (Network-on-Chip) technology to ensure the efficient scheduling and consistency of global memory access.

[0032] 4. Through the design of a central arbiter, the present invention can dynamically allocate resources according to task priority, data locality, and bus idle status to ensure the efficient and orderly data access.

[0033] 5. The present invention predicts future data access through the memory data prefetch model integrated in the local DRAM controller, closely cooperates with the cache subsystem, optimizes data flow, reduces unnecessary data migration, and further improves the overall performance and response speed of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] With reference to the accompanying drawings, the disclosure of the present invention will become more apparent. It should be understood that these drawings are only for illustrative purposes and are not intended to limit the scope of protection of the present invention. In the figures:

[0035] Figure 1 is a schematic diagram of a partial connection relationship of the present invention;

[0036] Figure 2 is a partial connection diagram of the DRAM chip and the logic chip of the present invention;

[0037] Figure 3 is a partial structural schematic diagram of the DRAM chip vertically stacked on the logic chip of the present invention;

[0038] Figure 4 is a connection block diagram of the DRAM chip and the logic chip of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0039] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention.

[0040] It should be noted that, without conflict, the following embodiments and the features in the embodiments may be combined with each other.

[0041] In the description of the present invention, it should be noted that for orientation terms, such as the terms "outer side", "middle section", "inner", "outer", etc., indicating orientation and position relationships are based on the orientation or position relationships shown in the drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and should not be construed as limiting the specific protection scope of the present invention.

[0042] In addition, such terms as "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, the meanings of "several" and "a number of" are two or more, unless otherwise specifically defined.

[0043] An embodiment of the present invention provides a chip design method, and the chip design method includes:

[0044] S1, integrating a local DRAM controller, a global DRAM controller connected to the local DRAM controller, and a shared DRAM controller connected to the local DRAM controller on a processor core, integrating a plurality of processor cores on a logic chip, and in the logic chip, connecting the local DRAM controllers to each other.

[0045] S2, vertically stacking a DRAM chip having a shared DRAM storage space and a plurality of private DRAM storage spaces on the logic chip through three-dimensional stacked storage technology, connecting each global DRAM controller to the shared DRAM storage space, and connecting each shared DRAM controller to its corresponding private DRAM storage space respectively.

[0046] In the prior art, usually a local DRAM controller is provided on a processor core, and the local DRAM controller is used to manage its local memory access requests. On this basis, the present invention additionally integrates two DRAM controllers in the processor core, and these two DRAM controllers are directly connected to the DRAM stacked on the upper layer.

[0047] In a DRAM, the present invention divides it into a shared DRAM storage space and several private DRAM storage spaces. Among them, the shared DRAM storage space is the storage space shared by all processor cores and can be accessed by all. The private DRAM storage space is the unique storage space of each processor core, and each processor core cannot directly access the private DRAM storage space of other processor cores. Therefore, the two DRAM controllers of the present invention are a global DRAM controller and a shared DRAM controller respectively. The global DRAM controller is directly connected to the shared DRAM storage space, reducing intermediate links, shortening the data path, and improving the data exchange speed. The shared DRAM controller is connected to its own unique private DRAM storage space to achieve memory shared data interaction between processor cores through the shared DRAM controller.

[0048] Through the above method, the data transmission path is optimized, the latency is reduced, and the bandwidth utilization rate is improved, thereby significantly enhancing the overall performance of the storage system in big data processing, high-performance computing, and cloud computing environments.

[0049] In some embodiments, the local DRAM controllers are connected through a NoC structure, so that the local DRAM controller of each processor core is responsible for managing its local memory access requests and works with the local controllers of other processor cores through the NoC to ensure the efficient scheduling and consistency of global memory access.

[0050] In some embodiments, a logic layer is also integrated on the logic chip, and each local DRAM controller is connected to the logic layer through an independent bus to form multiple parallel data transmission channels, allowing multiple data transmissions to be carried out simultaneously, significantly improving the data throughput.

[0051] In some embodiments, the logic layer is a CPU or an NPU.

[0052] In some embodiments, the chip design method further includes: designing a central arbiter, and the central arbiter is used to allocate the bandwidth resources of the bus.

[0053] The central arbiter in this embodiment can directly adopt the existing technology. For example, the central arbiter processes the bus access requests according to preset rules and real-time load information. For another example, the central arbiter dynamically allocates resources according to at least one factor such as task priority, data locality, and bus idle state to ensure the efficient and orderly data access.

[0054] In some embodiments, the chip design method further includes: integrating a memory data prefetch model in a local DRAM controller, where the memory data prefetch model is used to predict future data accesses, and using a global DRAM controller or a shared DRAM controller to load data in advance from a shared DRAM storage space or a private DRAM storage space to a cache.

[0055] The memory data prefetch model in this embodiment can directly adopt existing technologies, such as a model that uses the locality principle to predict future data accesses. By loading data from the DRAM to the cache in advance and closely collaborating with the cache subsystem, the data flow is optimized, unnecessary data migrations are reduced, and the overall performance and response speed of the system are further improved.

[0056] In some embodiments, when the global DRAM controller is connected to the shared DRAM storage space, the connection is implemented using copper interconnection of a metal layer.

[0057] In some embodiments, when the shared DRAM controller is connected to the corresponding private DRAM storage space, the connection is implemented using copper interconnection of a metal layer.

[0058] An embodiment of the present invention further provides a chip system, which includes a logic chip and a DRAM chip, and the DRAM chip is vertically stacked on the logic chip through three-dimensional stacked storage technology.

[0059] The logic chip integrates a number of processor cores, and each processor core integrates a local DRAM controller, a global DRAM controller, and a shared DRAM controller. The local DRAM controller is respectively connected to the global DRAM controller and the shared DRAM controller, and the local DRAM controllers of each processor core are connected to each other.

[0060] The DRAM chip has a shared DRAM storage space and a number of private DRAM storage spaces. The shared DRAM storage space is respectively connected to the global DRAM controller, and each private DRAM storage space is connected to a corresponding shared DRAM controller.

[0061] As Figure 1 shown, four processor cores in the logic chip are shown, which are the first processor core 10, the second processor core 20, the third processor core 30, and the fourth processor core 40. Of course, the logic chip may also have more or fewer processor cores. Taking four processor cores as an example,

[0062] The first processor core 10 has a first local DRAM controller 11, a first global DRAM controller 12, and a first shared DRAM controller 13, and its first local DRAM controller 11 is respectively connected to the first global DRAM controller 12 and the first shared DRAM controller 13.

[0063] The second processor core 20 has a second local DRAM controller 21, a second global DRAM controller 22, and a second shared DRAM controller 23. The second local DRAM controller 21 is respectively connected to the second global DRAM controller 22 and the second shared DRAM controller 23.

[0064] The third processor core 30 has a third local DRAM controller 31, a third global DRAM controller 32, and a third shared DRAM controller 33. The third local DRAM controller 31 is respectively connected to the third global DRAM controller 32 and the third shared DRAM controller 33.

[0065] The fourth processor core 40 has a fourth local DRAM controller 41, a fourth global DRAM controller 42, and a fourth shared DRAM controller 43. The fourth local DRAM controller 41 is respectively connected to the fourth global DRAM controller 42 and the fourth shared DRAM controller 43.

[0066] The first local DRAM controller 11, the second local DRAM controller 21, the third local DRAM controller 31, and the fourth local DRAM controller 41 are signal-connected to achieve the collaborative work of each processor core through a distributed memory management strategy, ensuring the efficient scheduling and consistency of global memory access.

[0067] Each of the four processor cores is respectively assigned an independent private DRAM storage space. Specifically, the first processor core 10 corresponds to the first private DRAM storage space 51, and the first private DRAM storage space 51 is connected to the first shared DRAM controller 13. The second processor core 20 corresponds to the second private DRAM storage space 52, and the second private DRAM storage space 52 is connected to the second shared DRAM controller 23. The third processor core 30 corresponds to the third private DRAM storage space 53, and the third private DRAM storage space 53 is connected to the third shared DRAM controller 33. The fourth processor core 40 corresponds to the fourth private DRAM storage space 54, and the fourth private DRAM storage space 54 is connected to the fourth shared DRAM controller 43.

[0068] The four processor cores have a common shared DRAM storage space 55. Among them, the shared DRAM storage space 55 is respectively connected to the first global DRAM controller 12, the second global DRAM controller 22, the third global DRAM controller 32, and the fourth global DRAM controller 42.

[0069] Such as Figure 2 and Figure 3As shown, the logic chip 6 has 8 processor cores, and the local DRAM controllers corresponding to the respective processor cores form a local DRAM controller cluster. Each local DRAM controller is connected to each storage space in the DRAM chip 7 through a shared DRAM controller or a global DRAM controller.

[0070] As Figure 4 shown, the logic chip 6 has 2 processor cores, namely processor core 61 and processor core 62. The way that processor core 61 or processor core 62 is connected to each storage space in the DRAM chip 7 is hybrid bonding (or Hybrid Bonding).

[0071] Since the DRAM chip 7 of the present invention is vertically stacked on the logic chip 6 through three-dimensional stacked storage technology, the hybrid bonding technology can be better applied to the three-dimensional stacking scenario.

[0072] In some embodiments, the local DRAM controllers of the respective processor cores are connected through a NoC structure, so that the local DRAM controller of each processor core is responsible for managing its local memory access requests, and cooperates with the local controllers of other processor cores through the NoC to ensure the efficient scheduling and consistency of global memory access.

[0073] In some embodiments, the logic chip is also integrated with a logic layer, and each local DRAM controller is connected to the logic layer through an independent bus to form multiple parallel data transmission channels, allowing multiple data transmissions to be performed simultaneously, significantly improving data throughput.

[0074] The chip system further includes: a central arbiter, and the central arbiter is used to allocate the bandwidth resources of the bus.

[0075] In some embodiments, the logic layer is a CPU or an NPU.

[0076] In some embodiments, a memory data prefetch model is integrated in the local DRAM controller. The memory data prefetch model is used to predict future data access, and uses the global DRAM controller or the shared DRAM controller to load data from the shared DRAM storage space or the private DRAM storage space into the cache in advance.

[0077] In some embodiments, when the global DRAM controller is connected to the shared DRAM storage space, the connection is implemented by copper interconnection of the metal layer.

[0078] In some embodiments, when the shared DRAM controller is connected to the corresponding private DRAM storage space, the connection is implemented by copper interconnection of the metal layer.

[0079] The above embodiments of the present invention adopt distributed bus DRAM control technology, especially dispersing control logic in a three-dimensional stacked structure, optimizing the data path, and realizing efficient data read and write operations, which are particularly applicable to high-performance computing scenarios such as intelligent driving. In this way, not only the storage bandwidth and density are improved, but also the system's ability to process real-time data is enhanced. The present invention adopts an architecture that combines distributed control logic with a multi-bus design to further improve the efficiency of data read and write operations.

[0080] The present invention has been described in detail in combination with the embodiments with reference to the accompanying drawings. Those of ordinary skill in the art can make various variations of the present invention according to the above description. Therefore, certain details in the embodiments should not constitute a limitation to the present invention, and the present invention will take the scope defined by the appended claims as the protection scope.

Claims

1. A chip design method, characterized in that: The chip design method comprises: Integrate a local DRAM controller, a global DRAM controller and a shared DRAM controller respectively connected to the local DRAM controller on a processor core, integrate several processor cores on a logic chip, and connect the local DRAM controllers in the logic chip; A DRAM chip having a shared DRAM storage space and a plurality of private DRAM storage spaces is vertically stacked on the logic chip through three-dimensional stacking storage technology, each of the global DRAM controllers is connected to the shared DRAM storage space, and each of the shared DRAM controllers is respectively connected to the corresponding private DRAM storage space.

2. The chip design method according to claim 1, characterized in that: The local DRAM controllers are connected to each other via a NoC structure, so that the local DRAM controller of each processor core is responsible for managing its local memory access request and works in coordination with the local controllers of other processor cores via the NoC; And / or, a logic layer is also integrated on the logic chip, and each local DRAM controller is connected to the logic layer through an independent bus to form a plurality of parallel data transmission channels, allowing multiple data transmissions to be performed simultaneously.

3. The chip design method according to claim 2, characterized in that: The logic layer is a CPU or an NPU; And / or, the chip design method further includes: designing a central arbitrator, wherein the central arbitrator is used to allocate bandwidth resources of the bus.

4. The chip design method according to claim 1, characterized in that: The chip design method further includes: A memory data prefetch model is integrated in the local DRAM controller, and the memory data prefetch model is used to predict future data access, and the global DRAM controller or the shared DRAM controller is used to load data from the shared DRAM storage space or the private DRAM storage space to the cache in advance.

5. The chip design method according to claim 1, characterized in that: The global DRAM controller is connected to the shared DRAM storage space by using metal layer copper interconnection to achieve connection; The shared DRAM controller is connected to the corresponding private DRAM storage space by using metal layer copper interconnection.

6. A chip system, characterized in that: The chip system comprises: A logic chip, wherein the logic chip integrates a plurality of processor cores, wherein the processor core integrates a local DRAM controller, a global DRAM controller, and a shared DRAM controller, wherein the local DRAM controller is respectively connected to the global DRAM controller and the shared DRAM controller, and the local DRAM controllers of the processor cores are connected to each other; A DRAM chip, wherein the DRAM chip is vertically stacked on the logic chip, the DRAM chip has a shared DRAM storage space and a plurality of private DRAM storage spaces, the shared DRAM storage spaces are respectively connected to the global DRAM controllers, and each of the private DRAM storage spaces is connected to a corresponding one of the shared DRAM controllers.

7. The chip system according to claim 6, characterized in that: The local DRAM controllers are connected via a NoC structure; And / or, a logic layer is also integrated on the logic chip, and each of the local DRAM controllers is connected to the logic layer via an independent bus.

8. The chip system according to claim 7, characterized in that: The logic layer is a CPU or an NPU; And / or, the chip system further includes: a central arbitrator, and the central arbitrator is used to allocate bandwidth resources of the bus.

9. The chip system according to claim 6, characterized in that: The local DRAM controller integrates a memory data prefetch model, which is used to predict future data access and use the global DRAM controller or the shared DRAM controller to load data from the shared DRAM storage space or the private DRAM storage space to the cache in advance.

10. The chip system according to claim 6, characterized in that: The global DRAM controller is connected to the shared DRAM storage space by using metal layer copper interconnection to achieve connection; The shared DRAM controller is connected to the corresponding private DRAM storage space by using metal layer copper interconnection.

Citation Information

Patent Citations

  • Business processing device and method, and business processing control device

    CN103019809A

  • LLC chip and cache system

    CN113643739A

  • Three-dimensional stacked programmable logic architecture and processor design architecture

    CN115858439A

  • Data transmission method and device, electronic equipment and storage medium

    CN116151345A

  • Access controller

    CN116737617A