Three-dimensional stacked chip system, memory access method and electronic equipment
By using a three-dimensional stacked chip system design, the direct access path between the chip and the memory chip is calculated. Combined with a dual-path memory access architecture, the transmission latency and bandwidth problems in the existing technology are solved, and efficient intelligent memory management is achieved.
Patent Information
- Application Number
- CN202511476735.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-10-15
AI Technical Summary
Existing chip technologies still fall short in reducing chip design complexity and manufacturing costs, especially in addressing the transmission latency and bandwidth issues between computing chips and memory chips.
The chip system design employs a three-dimensional stacked architecture, which stacks computing chips and memory chips through through-silicon vias (TSVs) and connects them to the substrate via input/output chips, enabling direct access paths between computing chips and memory chips. This is combined with a dual-path memory access architecture for intelligent memory management.
It reduces transmission latency between computing chips and memory chips, increases bandwidth, and enables efficient management of intelligent memory access, adapting to the needs of different access mechanisms.
Smart Images

Figure CN120957429A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of integrated circuit technology, and more specifically, to a three-dimensional stacked chip system, a memory access method, and an electronic device. Background Technology
[0002] With the development of integrated circuit manufacturing processes, the production cost of chips is increasing. Under the same manufacturing process, the more functions a chip needs to perform, the higher its design complexity, leading to lower chip yield. Furthermore, to support the required functions, the chip area increases. Constrained by chip yield and chip area, the manufacturing cost of chips remains high.
[0003] To reduce chip manufacturing costs, chiplet technology emerged. Chiplet technology is a design approach that breaks down complex chips into multiple small, independent, and reusable modules. These modules can be processor cores, memory chips, sensors, or other types of integrated circuits, connected via high-speed interfaces or advanced packaging technologies to form a complete system-on-a-chip (SoC). This approach effectively addresses the challenges of chip design complexity and manufacturing costs.
[0004] However, existing chip technology still needs improvement.
[0005] The information disclosed in the background section is only intended to enhance the understanding of the background of this application, and therefore may contain information that is not part of the prior art known to those skilled in the art. Summary of the Invention
[0006] This application provides a three-dimensional stacked chip system, a memory access method, and an electronic device that can reduce the transmission latency between computing chips and memory chips.
[0007] According to a first aspect of the embodiments of this application, a three-dimensional stacked chip system is provided, comprising: substrate; One or more computing chips are disposed on the substrate; One or more memory chips, each of which is stacked on top of a corresponding computing chip via through-silicon vias, and the computing chip and the memory chip are packaged into a first chip; One or more input / output chips are disposed on a substrate on one side of the computing chip and connected to each computing chip via substrate wiring. The input / output chips are packaged as a second chip and have no direct communication path with one or more of the memory chips.
[0008] Optionally, the computing chip includes one or more of a central processing unit, a graphics processing unit, and a neural network processor; The memory chips include: high-bandwidth memory chips and double data rate memory chips.
[0009] Optionally, the chip system further includes solder balls located between the computing chip and the substrate, and between the input / output chip and the substrate; The solder balls are used to connect the computing chip and the substrate, as well as the input / output chip and the substrate, and the substrate wiring passes through the solder balls to connect the input / output chip and the computing chip.
[0010] Optionally, the computing chip also transmits a first semantic tag carried by the memory access request through a multiplexed through-silicon via channel. The first semantic tag is parsed and executed by the memory controller, which is located within the first chip and connected to the one or more memory chips. Specifically, when the destination address of the memory access request points to the first storage space, the first semantic tag is generated, and the first storage space is the location of the one or more memory chips.
[0011] Optionally, the input / output kernel includes: a routing control unit configured to perform the following operations in response to a second semantic tag carried by a memory access request of the computing kernel: in write merge buffer mode, aggregating multiple write requests and sending them to the memory controller; in prefetch buffer mode, prompting for prefetching data to the cache based on the access mode. The computing chip sends a memory access request carrying a second semantic tag to the routing control unit through the substrate semantic channel, and generates the second semantic tag when the destination address of the memory access request points to the second storage space. The second storage space can be different from the first storage space.
[0012] Optionally, the second semantic tag includes: an operation mode identifier and parameter field information; The routing control unit includes: The semantic decoding module is configured to receive memory access requests from the computing chip, parse the second semantic tag carried by the request, and output a mode identifier and parameter field. The mode identifier includes: write merging mode and prefetch mode. The address matching module, coupled to the semantic decoding module, is configured to determine whether the target address of the memory access request falls within the accessible storage space based on the target address of the memory access request and the configurable storage space range in the programmable register group, and output a hit signal. The feature analysis module, coupled to the semantic decoding module, is configured to generate suggestion signals based on packet size, request type, and historical access trajectory. The arbitration module, coupled to the semantic decoding module, the address matching module, and the feature analysis module, is configured to, in response to the validity of the pattern identifier, adopt the processing mode of the memory access request corresponding to the pattern identifier according to a first priority; in response to the first memory space hit signal output by the address matching module, adopt the prefetch mode according to a second priority; and in response to the output of the second memory space hit signal, adopt the write merging mode according to a third priority; and in response to the suggestion signal, adopt the suggestion mode of the suggestion signal according to a fourth priority.
[0013] Optionally, the operation of the write merge buffer mode includes: collecting unaligned write requests within a time window and merging them into burst write operations after aligning them according to DRAM page boundaries; The operation of the prefetch buffer mode includes: calculating the prefetch step size according to the access mode prompt, and continuously reading subsequent data blocks according to the step size and storing them into the cache.
[0014] Optionally, the chip system further includes: a programmable interconnect layer disposed between corresponding computing chips and memory chips, the programmable interconnect layer including: a dynamic memory pooling module configured to manage heterogeneous memory resources, the heterogeneous memory resources being composed of multiple memory chips; The dynamic memory pooling module is configured to generate a unified resource vector based on the physical topology location attributes, bandwidth, latency, durability, and energy consumption characteristics of the heterogeneous memory resources; allocate, release, and map the unified resource vector using cache lines as the basic unit; perform topology-aware allocation based on the relative positional relationship between the computing kernels and the memory kernels, prioritizing the allocation of frequently accessed data to the physical memory region with the lowest access latency; divide the physical memory region into time slices and perform time-sharing multiplexing between different tasks, with the time-sharing multiplexing process coordinated with hardware context switching operations; monitor energy consumption status and migrate cold data to low-power memory regions or place idle memory kernels into a low-power state according to predefined energy efficiency strategies; perform data striping and verification calculations within the memory pooling layer; in response to the detection of memory kernel failures, perform online data recovery using verification information and redundant data, and mark and isolate faulty memory kernels.
[0015] According to a second aspect of the embodiments of this application, a memory scheduling method is provided, applied to the three-dimensional stacked chip system described in any of the foregoing claims, the memory scheduling method comprising: In response to a memory access request sent by a computing chip, a parsing operation is performed; In response to the parsing result of the parsing operation, the following operations are performed: in write merge buffer mode, multiple write requests are aggregated and sent to the memory controller, which is located in the first chip; in prefetch buffer mode, data is prefetched to the cache based on the access mode prompt.
[0016] According to a third aspect of the embodiments of this application, an electronic device is provided, comprising: a three-dimensionally stacked chip system as described in any one of the claims.
[0017] The embodiments of this application, by adopting the above technical solutions, have the following technical effects: One or more computing chips, one or more memory chips, and one or more input / output chips are all disposed on a substrate. Each memory chip is stacked above the corresponding computing chip through a through-silicon via (TSV) and packaged as a first chip. The input / output chips are packaged as a second chip and have no direct communication path with the one or more memory chips. Thus, compared to existing solutions where the input / output chips and memory chips are packaged together, the access path between the computing chips and memory chips in this application is shortened, thereby reducing the transmission latency between the computing chips and memory chips. Attached Figure Description
[0018] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a schematic diagram of a packaging structure for a computing chip and an input / output chip; Figure 2 This is a schematic diagram of the structure of a three-dimensional stacked chip system according to an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a routing control unit in an embodiment of this application. Detailed Implementation
[0019] To make the technical solutions and advantages of the embodiments of this application clearer, the exemplary embodiments of this application will be described in further detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not an exhaustive list of all embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other.
[0020] As mentioned in the background section, chiplet technology effectively addresses the challenges of chip design complexity and manufacturing costs. In current solutions, chipslets are divided into compute chiplets and input / output chiplets based on their functions.
[0021] The computing chips include various computing units such as the central processing unit (CPU), graphics processing unit (GPU), and neural network processor (NPU).
[0022] Input / output chips include Peripheral Component Interconnect Express (PCIe), Universal Serial Bus (USB), Double Data Rate (DDR) memory, High Bandwidth Memory (HBM), etc.
[0023] The computing chip and the input / output chip are manufactured using different processes, and then packaged together after production. This can greatly reduce manufacturing costs and iteration time.
[0024] See Figure 1 The diagram shown illustrates a packaging structure for a computing chip and input / output chips, as follows: Figure 1 As shown, after manufacturing the computing chip 12 and the input / output chip 10 respectively, communication between the computing chip 12 and the input / output chip 10 is achieved through interconnection lines. Thus, the computing chip 12 can access the memory chip 14, such as DDR and HBM, through the input / output chip 10.
[0025] However, due to protocol conversion, cross-chip transmission introduces additional latency, and the bandwidth of inter-chip interconnects is significantly lower than the bandwidth within a single chip. For computing units sensitive to data latency, this bandwidth and latency impact cannot be ignored.
[0026] Furthermore, the memory chip 14 is packaged together with the input / output chip 10, which allows for a longer access path between the memory chip 14 and the computing chip 12.
[0027] To address the aforementioned technical problems, this application provides a chip design scheme based on three-dimensional (3D) stacking, where a memory chip is stacked on top of a latency-sensitive computing chip. By stacking the memory chips, the computing units within the computing chip can directly access the memory chip within their own packaging structure, thus shortening the access path between the computing chip and the memory chip.
[0028] Furthermore, since it does not involve protocol conversion of memory chips, it can effectively reduce the latency of computing units accessing memory in computing chips, thereby reducing the transmission latency between computing chips and memory chips.
[0029] Furthermore, this application achieves a bus width of over 1024 bits through through-silicon via (TSV) technology. Compared to the Universal Chipplet Interconnect Express (UCIE), this significantly improves bandwidth.
[0030] In other words, the embodiments of this application simultaneously reduce the transmission latency of memory access and increase bandwidth.
[0031] To enable those skilled in the art to better understand and implement this solution, the following detailed description of the specific solution, principles, advantages, and effects of this application is provided with reference to the accompanying drawings and specific embodiments.
[0032] See Figure 2 The diagram shown is a structural schematic of a three-dimensional stacked chip system according to an embodiment of this application. Figure 2 As shown, the three-dimensional stacked chip system 200 may include: substrate 202; One or more computing cores 204 are disposed on the substrate 202; One or more memory chips 206, each memory chip 206 is stacked on top of a corresponding computing chip 204 through a through silicon via 208, and the computing chip 204 and the memory chip 206 are packaged into a first chip 210; One or more input / output chips 212 are disposed on a substrate 202 on one side of a computing chip 204 and are connected to each computing chip 204 via substrate wiring. The input / output chips 212 are packaged into a second chip 220 and have no direct communication path with one or more memory chips.
[0033] In some embodiments, substrate 202 provides a process platform for the formation of chip system 200.
[0034] Specifically, substrate 202 can serve as a bridge between the chip and the circuit board. It is made of materials such as organic materials, ceramics, or silicon and has wiring layers on it to provide electrical connections, mechanical support, and heat dissipation channels.
[0035] For example, in an advanced packaging solution, substrate 202 is used to connect a chip (or die) to an external system.
[0036] The computing core 204 is responsible for processing and computation, similar to a CPU or GPU core in a traditional solution.
[0037] The Memory Chip 206 provides storage and caching for storing data, similar to on-chip memory or external memory expansion.
[0038] Input / output chip 212 is used to handle external interfaces and signal transmission.
[0039] In short, the computing core 204 is the "brain," the memory core 206 is the "memory," and the input / output core 212 is the "neural network interface." The three work together to realize the data processing process.
[0040] In some embodiments, the memory chip 206 is stacked above the corresponding compute chip 204 via through-silicon vias 208. This allows the compute chip 204 to directly access the memory chip 206, reducing transmission latency.
[0041] Furthermore, by encapsulating the computing chip 204 and the memory chip 206 into a first chip 210, the interaction between the computing chip 204 and the memory chip 206 is achieved within the same package, further reducing transmission latency.
[0042] In some embodiments, through-silicon via interconnects are vertically connected, with extremely short paths, resulting in significantly reduced signal delay.
[0043] In some embodiments, the input / output chip 212 is connected to the computing chip 204 via substrate wiring and is packaged independently as a second chip 220, and has no direct communication path with one or more memory chips 206, so that the computing chip 204 can directly access the memory chip 206 without going through the input / output chip 212, thereby reducing transmission latency.
[0044] In other words, this application binds the computing chip 204 and the memory chip 206 together by 3D stacking, which effectively reduces the transmission latency of the computing unit in the computing chip 204 accessing the memory.
[0045] It should be pointed out that, firstly, Figure 2 The illustrated chip system 200's packaging structure is merely an example, used to represent packaging the computing chip 204 and memory chip 206 together. It is an abstract expression and should not be construed as a limitation of this application; secondly... Figure 2 The number and location of the computing chip 204, memory chip 206, through-silicon via 208, and input / output chip 212 shown are illustrative examples used to represent the constituent components of the chip system 200; third, the focus of this application is on structural changes and does not mention how the computing chip, memory chip, through-silicon via, and input / output chip are formed, which can be found in existing examples.
[0046] For example, a through-silicon via (TSV) can be formed by: forming an interlayer dielectric layer on a computing die; forming a via within the interlayer dielectric layer; and forming a metal interconnect structure within the via.
[0047] In some embodiments, the computing core can be used as a processing core to control different processes.
[0048] For example, the computing chip may include one or more of a central processing unit, a graphics processing unit, and a neural network processor.
[0049] In widely used processing, stacked computing and memory chips can reduce latency and meet the needs of most application scenarios.
[0050] In some embodiments, memory chips may include high-bandwidth memory chips and double-data-rate memory chips. High-bandwidth memory prioritizes bandwidth and is suitable for computationally intensive applications, reducing signal latency in scenarios such as AI model training; double-data-rate memory prioritizes capacity and cost and is suitable for large-scale general-purpose computing.
[0051] For example, double data rate memory can include different types of memory such as DDR4 and DDR5.
[0052] In some embodiments, see next. Figure 2 To enable interconnection between different chips, the three-dimensional stacked chip system 200 may also include solder balls 230 located between computing chip 204 and substrate 202, and between input / output chip 212 and substrate 202.
[0053] Solder balls 230 are used to connect computing chip 204 and substrate 202, as well as input / output chip 212 and substrate 202, and substrate wiring is connected to input / output chip 212 and computing chip 204 through solder balls 230.
[0054] In other words, the interconnection between internal components of the three-dimensionally stacked chip system 200 is achieved through solder balls 230.
[0055] Furthermore, solder balls provide a stable conductive path, ensuring the transmission of high-speed signals, clock signals, and power supplies. Compared to traditional wire bonding, solder ball connections have shorter paths and better signal integrity.
[0056] It should be noted that the solder ball connection method is only one example. In other embodiments, connections between different structures can also be achieved through one or more of the following methods: through-silicon vias, hybrid bonding, and pin / slot connections (e.g., two-layer structures interconnected by a ball grid array).
[0057] In some embodiments, a memory chip is bonded on top of a latency-sensitive computing chip, and the computing chip directly accesses the memory chip through a through-silicon via (TSV).
[0058] As system complexity increases, it becomes necessary to integrate multiple types of memory and to intelligently manage memory access (such as write merging and prefetching) to further improve efficiency.
[0059] The current direct access path scheme cannot achieve this kind of intelligent management, especially since other computing chips use the method of accessing memory chips through input / output chips.
[0060] To further overcome the limitations of different access mechanisms, this application also defines a scheme that can support heterogeneous memory access and intelligent memory management, namely, by introducing a dual-path memory access architecture to achieve the above objectives.
[0061] The dual-path memory access includes: The first access path, i.e., the access request for stacked memory (corresponding to the first storage space), is sent directly to the memory controller by the computing chip through the through-silicon via channel, ensuring ultra-low latency access of the core computing to high-speed memory.
[0062] The second access path, i.e., for access requests to onboard memory (corresponding to the second storage space), involves the computing chip sending the request to the input / output chip. The routing control unit built into this input / output chip improves the efficiency and bandwidth utilization of onboard memory access by parsing the semantic tags in the request.
[0063] In short, in a solution where the computing chip and memory chip are packaged together, the memory chip is accessed directly; while in a solution where the memory chip and input / output chip are packaged together, the memory chip is accessed through the input / output chip.
[0064] Specifically, this application includes two different types of memory chips. To accommodate this change, input / output chips and the operating logic of computation chips can be configured to execute different access strategies in response to memory access requests from computation chips.
[0065] This differentiated approach provides various configuration options to meet the needs of different users and cover a wide range of usage scenarios.
[0066] For example, the computing chip is configured to generate different semantic tags for memory access requests to different destination addresses, and send them respectively through the substrate semantic channel or through-silicon via channel.
[0067] In some embodiments, the computing chip also transmits a first semantic tag carried by the memory access request through a multiplexed through-silicon via (TSV) channel. The first semantic tag is parsed and executed by a memory controller located within a first chip and connected to one or more memory chips.
[0068] Specifically, when the destination address of a memory access request points to the first storage space, a first semantic tag is generated, and the first storage space is the location of one or more memory chips.
[0069] In other words, when it is determined that the access is to the memory chip inside itself, the data processing operation is performed directly in the traditional way of memory controller-memory chip.
[0070] Multiplexing through-silicon via (TSV) channels refers to reusing communication paths based on TSVs for transmitting data between chips or between chips and memory controllers.
[0071] The first semantic tag is transmitted to the memory controller through a multiplexed through-silicon via (TSV) channel, where it is parsed and performs corresponding operations (such as triggering prefetching, write merging, or other optimizations).
[0072] Adaptively, the input / output kernel includes: a routing control unit configured to perform the following operations in response to a second semantic tag carried by a memory access request from a computing kernel: in write merge buffer mode, aggregating multiple write requests and sending them to the memory controller; in prefetch buffer mode, prefetching data to the cache based on the access mode prompt; wherein the computing kernel sends a memory access request carrying the second semantic tag to the routing control unit through the substrate semantic channel.
[0073] In some embodiments, a memory access request initiated by a computing chip can carry indication information of the memory to be accessed. That is, a second semantic tag indicates the memory chip that the computing chip wants to access. For example, the type and address of the memory chip, and other information.
[0074] In this scenario, the routing control unit can parse the second semantic tag and execute different access paths based on the final parsing result.
[0075] For example, in one case, in write merge buffer mode, multiple write requests are aggregated and then sent to the memory controller.
[0076] In computer systems, write merging refers to the practice of not immediately sending a single write request to memory each time data needs to be written. Instead, multiple small write requests are temporarily stored in a buffer and merged into one or more larger requests before being sent.
[0077] Compared to direct write mode, write-merge mode reduces the number of memory accesses by merging requests, thus reducing bandwidth pressure.
[0078] In some embodiments, a buffer is a temporary storage area used to collect and organize multiple write requests. The physical location of the buffer can be a memory chip, or inside a computing chip, etc.
[0079] The memory controller is a hardware component responsible for managing data transfer between memory chips and compute chips. It handles read and write requests, controlling memory timing, refresh, and data transfer. The memory controller can be located inside the first chip. In write-merge buffer mode, multiple write requests are aggregated, and the buffer sends the merged data as a whole to the memory controller. The memory controller then writes this data to the memory chip.
[0080] Specifically, the write merge buffer mode operation includes: collecting unaligned write requests within a time window and merging them into burst write operations after aligning them according to DRAM page boundaries.
[0081] The time window refers to a specific time range during which the write merge buffer collects write requests issued by the processor, instead of immediately sending each request to the memory controller.
[0082] Unaligned means that the target address of a write request does not fully conform to some boundary requirement of memory (e.g., cache line boundary or DRAM page boundary).
[0083] For example, a typical cache line size is 64 bytes. If the target address of a write request is not on a 64-byte aligned boundary, it is called unaligned.
[0084] In DRAM, misalignment may refer to a write request address not being aligned with the DRAM page boundary.
[0085] Collecting unaligned write requests means that the buffer temporarily stores these unaligned requests instead of sending them directly to the memory controller.
[0086] Furthermore, a DRAM page boundary is a contiguous storage cell, typically 4KB, 8KB, or larger in size. Page boundary alignment refers to ensuring that the starting address of a write request is aligned with the boundary of the DRAM page.
[0087] In the buffer, collected unaligned write requests are reorganized to align their addresses with DRAM page boundaries. This alignment may involve: address remapping, i.e., adjusting the address of the request so that data is written to the aligned DRAM page; and data reorganization, i.e., rearranging multiple unaligned requests to fill the same DRAM page.
[0088] Burst writes refer to the operation of continuously transferring multiple data blocks in a single memory transaction. This can make efficient use of bandwidth, allowing multiple data blocks to be transferred at once and reducing control overhead on the memory bus.
[0089] In another scenario, under prefetch buffer mode, data is prefetched into the cache based on access pattern hints.
[0090] Among them, the cache is a fast storage unit (such as L1, L2, L3 cache) located between the processing core and the memory core, used to store frequently accessed data.
[0091] Write-merge buffer mode refers to a write-merge mechanism enabled under a specific operating state. Under this mechanism, a dedicated prefetch mechanism is activated, which determines the memory addresses that may be needed based on access pattern hints. Data is prefetched from main memory and loaded into a cache (or prefetch buffer) according to these hints, allowing the computing chip to access it quickly and reducing memory access latency.
[0092] Specifically, the prefetch buffer mode operation includes: calculating the prefetch step size based on the access mode hints, and continuously reading subsequent data blocks into the buffer according to the step size.
[0093] For example, access pattern hints are used to guide the prefetching mechanism in predicting memory addresses that the processor may access in the future. These hints can come from: analyzing the computing chip's past memory access behavior to identify regular patterns, such as sequential access, strafing, and pointer tracing; or, the compiler or metadata can explicitly specify addresses that need to be prefetched. These hints help the prefetching mechanism predict future memory access addresses more accurately.
[0094] Prefetch stride refers to the address interval used by the prefetch mechanism when predicting the address of subsequent data. The stride reflects the regularity of memory access, for example: sequential access: the stride is 1 cache line (usually 64 bytes), that is, subsequent memory blocks are loaded consecutively; strafing access: the stride is a fixed size, such as loading a data block every 256 bytes (common in matrix or multidimensional array access); dynamic stride: the stride may be dynamically adjusted according to the runtime access mode.
[0095] Once the prefetch step size is determined, the prefetch mechanism predicts subsequent memory addresses according to this step size and loads multiple data blocks consecutively according to the step size. This takes advantage of the burst transfer characteristics of DRAM to efficiently read data from main memory and store the read data blocks in the cache or prefetch buffer so that the processor can directly hit the cache when it makes a request in the future, avoiding access to the slower DRAM.
[0096] The computing chip sends a memory access request carrying a first semantic tag to the routing control unit through the semantic channel on the substrate. This means that the computing chip generates a memory request carrying a semantic tag and sends it to the routing control unit through the semantic channel on the substrate. The routing control unit optimizes the route based on the tag and address.
[0097] This approach enhances the intelligence and efficiency of traditional memory access by introducing semantic information, while reducing latency and power consumption. It is particularly suitable for high-performance computing and AI scenarios, and it integrates closely with memory optimization patterns from historical conversations (such as write merging and prefetching) to form a complete memory management system.
[0098] In some embodiments, when the destination address of a memory access request points to a second storage space, a second semantic tag is generated, and the second storage space may be different from the first storage space.
[0099] In other words, the computational core generates a memory access request and selectively attaches a first semantic tag or a second semantic tag based on the destination address. That is, if the destination address points to the first storage space, a second semantic tag is generated; if the destination address points to the second storage space, a first semantic tag is generated.
[0100] It should be noted that the reason for allowing the second storage space to differ from the first storage space is to account for unforeseen circumstances, such as access request signals from other computing chips carrying information indicating the first storage space, which is the address of a memory chip specific to another computing chip. Thus, inter-chip interaction is achieved when the second storage space can be identical to the first storage space.
[0101] By generating different semantic tags based on the destination address, the system achieves adaptive optimization based on storage space, enhancing flexibility and efficiency, and adapting to the heterogeneous storage needs of multi-core systems.
[0102] In short, for a multi-chip heterogeneous memory structure, for memory access requests destined for the second storage space, the computing chip sends a memory access request carrying a second semantic tag to the routing control unit through the substrate semantic channel. For memory access requests destined for the first storage space, the computing chip sends a memory access request carrying a first semantic tag to the memory controller through a multiplexed through-silicon via (TSV) channel.
[0103] In some embodiments, the second semantic tag includes: an operation mode identifier and parameter field information. The operation mode identifier indicates the processing mode of the memory access request, and the parameter field information indicates the behavior of that processing mode.
[0104] Specifically, the operation mode identifier is a field in the semantic tag that indicates which operating mode the dynamic routing control unit should use to process the current request.
[0105] For example: 000 represents normal pass-through mode, 001 represents write-merge mode, 010 represents prefetch mode, and 011 represents cache-write-through mode. It should be noted that more modes can be added according to actual needs, such as error correction mode and delay masking mode, and the above characters are only examples.
[0106] Thus, when the field is detected to be valid, the corresponding pattern will be used as the highest priority input.
[0107] The parameter field is a set of control parameters in the semantic tag that are paired with the pattern identifier. These parameters are used to refine the behavior of the execution pattern and include at least the following: time window size, which determines the aggregation range of write merging; prefetch step size (4~256 bytes), which is used for the granularity of continuous data reading; buffer depth limit; and triggering conditions, such as access count threshold.
[0108] Accordingly, see Figure 3 The schematic diagram shown in this application illustrates the structure of a routing control unit in an embodiment of the present application. Figure 3 As shown, the routing control unit 300 may include: The semantic decoding module 302 is configured to receive memory access requests from the computing core 204, parse the second semantic tag carried by it, and output the pattern identifier and parameter field. The pattern identifier includes: write merging mode and prefetch mode. Address matching module 304, coupled to semantic decoding module 302, is configured to determine whether the target address of the memory access request falls into the accessible storage space based on the target address of the memory access request and the configurable storage space range in the programmable register group, and output a hit signal. The feature analysis module 306, coupled to the semantic decoding module 302, is configured to generate suggestion signals based on packet size, request type, and historical access trajectory. Arbitration module 308, coupled to semantic decoding module 302, address matching module 304, and feature analysis module 306 respectively, is configured to, in response to a valid pattern identifier, adopt the processing mode of the memory access request corresponding to the pattern identifier according to a first priority; in response to a first memory space hit signal output by the address matching module, adopt a prefetch mode according to a second priority; in response to a second memory space hit signal output by the address matching module, adopt a write merging mode according to a third priority; and in response to a suggestion signal, adopt the suggestion mode of the suggestion signal according to a fourth priority.
[0109] Specifically, the second semantic tag carried in the memory access request output by the computing core 204 contains a large amount of useful information. By performing a parsing operation, the data used by the address matching module 304 can be obtained.
[0110] The target address of a memory access request can be understood as the memory chip to be accessed.
[0111] The programmable register set is written to during startup by the initialization software / firmware or the chip system configuration logic. By design, programmable registers are part of the dynamic routing control unit and can be configured by external software via register mapping (MMIO).
[0112] The programmable register may include: the start and end addresses of a first memory space, the start and end addresses of a second memory space, and may also include address ranges of other memory or peripherals.
[0113] The first storage space can be different from the second storage space. For example, the first storage space can be a memory chip pointing to a memory chip in a first chip, while the second storage space can be a memory chip pointing to a memory chip in a second chip, or another memory chip.
[0114] Thus, the address matching module 304 can determine the address range of the memory to be read based on the target address and storage space range of the memory access request, and then output a hit signal.
[0115] It is understandable that a portion of the space in the first storage space and the second storage space can be the same, that is, there is overlap between the two.
[0116] In some embodiments, the request type may be obtained from memory access requests sent by the computing chip. Each memory access request contains its type information, which is typically generated and embedded in the request by memory access control logic (such as load, store, write, or read).
[0117] In the routing control unit 300, when the semantic decoding module 302 parses the second semantic tag, it extracts and identifies the request type.
[0118] The packet size is typically defined within a memory access request. Each memory request carries not only the address but also the size of the requested data block. The packet size may also be dynamically determined by the compute chip or memory controller based on the needs of memory operations.
[0119] Historical access records can be obtained from the access history buffer of the compute chip or memory controller. The dynamic routing control unit can query these historical records to analyze access patterns and make optimization decisions.
[0120] For example, if an address is accessed frequently, it may be considered a locality-based access pattern, thus generating a prefetch request or enabling a write merging strategy.
[0121] In this case, the feature analysis module 306 no longer simply responds to every request, but is able to proactively analyze traffic characteristics, predict future behavior, and make reasonable suggestions.
[0122] For example, the computing chip issues a write request to write 128 bytes of data to an unaligned address. The feature analysis module 306 analyzes the packet size and finds that 128 bytes is a relatively small write operation, which is worth merging.
[0123] Since this is a write request, write merging mode is suitable. Furthermore, analysis of historical access patterns indicates that no other write requests have recently been made to this address or its vicinity, preventing the formation of a merge queue. Therefore, the urgency of this merge is not the highest. Thus, after comprehensive evaluation, a recommendation signal is generated, which could be: "It is recommended to use write merging mode for this request."
[0124] When all the above information reaches the arbitration module 308, the arbitration module 308 can, based on the output results, perform the following operations: When a pattern identifier is valid, the processing mode for the memory access request corresponding to that pattern identifier is adopted according to the first priority.
[0125] Specifically, the first priority is the highest priority. When the pattern identifier is determined to be valid, the processing is carried out based on the processing method indicated by the validity of the pattern identifier.
[0126] In response to the first memory space hit signal output by the address matching module, the prefetch mode is adopted according to the second priority.
[0127] Specifically, the first memory space hit signal is used to indicate access to the memory chip in the first chip. Although the optimal solution of directly accessing the memory chip through a through-silicon via (TSV) has been mentioned in the above example, it only provides a more efficient default path, while retaining the possibility of accessing the memory chip through an input / output chip for some advanced optimization purpose.
[0128] For example, in one application scenario, the data being processed by compute chip A will soon be needed by compute chip B. This data is stored in the HBM of compute chip A. If accessed via a direct path, compute chip B cannot directly access the HBM of compute chip A.
[0129] However, by configuring compute chip A's HBM address space to be mapped as the "first memory space" that the input / output chip path can manage, compute chip B can send a request with a prefetch hint to the input / output chip.
[0130] In this way, the routing control unit of the input / output chip can parse the request, and then actively pre-extract the data from the HBM of computing chip A through the substrate, and put it into a shared buffer or send it directly to computing chip B.
[0131] In other words, in the preferred embodiment, the first storage space is mapped to high-speed memory (such as HBM) that is directly accessed via a through-silicon via (TSV) to achieve the lowest possible access latency.
[0132] However, this application also supports more flexible configurations. In certain application scenarios, access requests to its primary storage space (such as HBM) can be routed to input / output granules for intelligent management. This mode is of great significance for achieving data coordination across computing granules and system-level prefetching.
[0133] Among them, the second priority is less than the first priority.
[0134] In response to the output of the second memory space hit signal, the write merge mode is adopted according to the third priority.
[0135] Specifically, the second storage space points to DDR. In this case, a write-merge mode is used.
[0136] In response to the suggestion signal, the suggestion mode of the suggestion signal is adopted according to the fourth priority. That is, the suggestion mode of the suggestion signal has the lowest priority.
[0137] Thus, by employing the above-described logic control, a multi-input, hierarchical priority arbitration scheme is provided, achieving a robust, efficient, and highly intelligent decision-making system. By clearly defining the decision-making logic, conflicts and uncertainties are avoided.
[0138] Specifically, this application adopts a method that prioritizes instructions, followed by configuration, and optimizes for minimum enablement, ensuring that the arbitration module can make a unique and definite decision, thus avoiding conflicts caused by multiple signals being valid simultaneously. That is, the priority levels from first to fourth decrease sequentially.
[0139] In some embodiments, the above scheme can be further optimized.
[0140] For example, a programmable interconnect layer is provided on computing chips and memory chips with corresponding relationships. The programmable interconnect layer includes a dynamic memory pooling module configured to manage heterogeneous memory resources, which are composed of multiple memory chips.
[0141] Specifically, the dynamic memory pooling module is configured to generate a unified resource vector (i.e., implement resource abstraction) based on the physical topology location attributes, bandwidth, latency, durability, and energy consumption characteristics of heterogeneous memory resources; allocate, release, and map addresses to the unified resource vector using cache lines as the basic unit; perform topology-aware allocation based on the relative positional relationship between computing granules and memory granules, prioritizing the allocation of frequently accessed data to the physical memory region with the lowest access latency; divide the physical memory region into time slices and perform time-sharing multiplexing between different tasks, with the time-sharing multiplexing process coordinated with hardware context switching operations; monitor energy consumption status and migrate cold data to low-power memory regions or place idle memory granules into a low-power state according to predefined energy efficiency strategies; perform data striping and verification calculations within the memory pooling layer, and in response to the detection of memory granule failures, perform online data recovery using verification information and redundant data, and mark and isolate faulty memory granules.
[0142] This involves identifying the memory resources (HBM, DDR, etc.) in the system and evaluating the static attributes (such as physical location and type) and dynamic attributes (such as current bandwidth and latency) of each resource in detail. It abstracts all this information into a unified, software-readable format (Uniform Resource Vector), which is the foundation for intelligent management.
[0143] Next, using cache lines as the basic unit, allocation, deallocation, and address mapping are performed, resulting in very fine-grained management compared to traditional large blocks of memory. This allows for extremely precise allocation of data locations, avoiding waste and achieving ultimate optimization.
[0144] Then, based on relative positional relationships, topology-aware allocation is performed, prioritizing frequently accessed data to the physical memory region with the lowest access latency. Specifically, the module knows that accessing the HBM stack is much faster than accessing DDR. Therefore, it intelligently places hot data (i.e., frequently accessed data) on the HBM closest to the compute unit, while placing cold data on DDR, thereby maximizing overall performance.
[0145] Furthermore, by dividing the physical memory region into time slices and performing time-sharing multiplexing between different tasks, it works in conjunction with hardware context switching operations. When switching from one task to another (i.e., hardware context switching), the memory pooling module can quickly allocate HBM space to the new task, enabling multiple tasks to share the same high-speed memory, greatly improving resource utilization.
[0146] Furthermore, by continuously monitoring data access frequency (hot data / cold data), infrequently accessed cold data is automatically migrated to more energy-efficient memory (such as certain low-power DDR modes). It can even put completely idle memory chips into sleep mode, thereby dynamically reducing overall system power consumption.
[0147] To further enhance reliability, when a memory chip fails, the system can reconstruct the lost data in real time using checksums and data from other chips, enabling seamless repair for upper-layer applications and ensuring uninterrupted service. Furthermore, the faulty chip is marked as unusable, preventing data allocation to it and thus preventing the spread of errors.
[0148] In other words, by setting up a dynamic memory pooling module, the paradigm of active computing and passive memory can be transformed, and the memory system can be changed from a simple storage warehouse into an intelligent, proactive, and schedulable resource service grid.
[0149] Its function can be summarized in the following three core aspects: First, by using Cache-Line granular allocation and topology-aware allocation, memory access latency is significantly reduced, bandwidth utilization is greatly improved, and memory wall is broken.
[0150] Second, by using time-slice reuse, physical memory resources are divided in time and allocated to multiple users or tasks in a time-sharing manner, which can maximize resource utilization and economic benefits.
[0151] Third, through energy consumption sensing strategies, data is intelligently placed in memory with the best energy efficiency ratio, and memory chips are put into a low-power state when idle, significantly reducing system energy consumption.
[0152] In this embodiment, "context" refers to the complete environmental state that a task (process, thread) depends on when it executes. "Hardware context" specifically refers to the critical state in this environment that is directly managed and used by the computing chip hardware.
[0153] The hardware context is typically stored inside the computing core or in specific registers very close to the computing core. It mainly includes: program counter, the address of the currently executing instruction; stack pointer, the memory address of the current stack; general-purpose registers, which store temporary computation data; status registers, which record the result of the previous operation (such as whether it is 0, whether it overflowed, etc.); and memory management unit registers, such as the page table base address register, which defines how the "virtual memory space" seen by the current task is mapped to "physical memory".
[0154] By switching hardware contexts, memory mappings or policies can be switched to ensure that new tasks can correctly access their allocated memory time slices, thereby achieving efficient and transparent memory sharing.
[0155] Physical memory resources refer to the sum of all available physical memory hardware in the system, while a physical memory region refers to a portion of physical memory resources that is divided according to specific attributes or functions.
[0156] In some embodiments, the programmable interconnect layer may further include a switch matrix network configured to dynamically reconfigure the connection paths between computing chips and memory chips.
[0157] In the switch matrix network, the generation of channel scheduling instructions depends on the resource vectors and topology information provided by the dynamic memory pooling module, and the topology-aware allocation of the dynamic memory pooling module depends on the channel status information provided by the switch matrix network.
[0158] For example, the switch matrix network is configured to receive memory access requests from computing chips; based on a predetermined monitoring cycle, it collects real-time status data of each silicon channel, including channel utilization, queuing depth, and access latency; combining historical channel status data and real-time status data, it generates multiple logical candidate paths through a path calculation algorithm, where each logical candidate path corresponds to a virtualization mapping and scheduling scheme for a silicon channel; through a conflict arbitration mechanism, it determines a target logical path group from the multiple logical candidate paths based on the priority and bandwidth requirements of the memory access request; and generates channel scheduling instructions based on the target logical path group to dynamically allocate data streams to multiple silicon channels in a logical manner, thereby achieving multi-channel load balancing and cross-cycle bandwidth allocation.
[0159] This application also provides a memory scheduling method that can be applied to the three-dimensional stacked chip system in the aforementioned example.
[0160] The memory scheduling method may include: in response to a memory access request sent by a computing chip, performing a parsing operation (i.e., step 1); in response to the parsing result of the parsing operation, performing the following operations: in write merge buffer mode, aggregating multiple write requests and sending them to the memory controller, which is located in the first chip; in prefetch buffer mode, prompting to prefetch data to the cache based on the access mode (i.e., step 2).
[0161] In one example, the memory scheduling method may include: generating a first semantic tag in response to a memory access request generated by a computing chip whose target address points to a first storage space, and transmitting the first semantic tag to a memory controller via a multiplexed silicon channel; and generating a second semantic tag in response to a memory access request generated by a computing chip whose target address points to a second storage space, and transmitting the second semantic tag to a routing control unit via a substrate semantic channel.
[0162] Step 2 is for the execution scheme of the second storage space.
[0163] This application also provides an electronic device that may include the three-dimensional stacked chip system described in the foregoing example.
[0164] Electronic devices can include mobile devices such as mobile phones and tablets, and can also include computers.
[0165] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0166] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0167] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A three-dimensional stacked chip system, characterized in that, include: substrate; One or more computing chips are disposed on the substrate; One or more memory chips, each of which is stacked on top of a corresponding computing chip via through-silicon vias, and the computing chip and the memory chip are packaged into a first chip; One or more input / output chips are disposed on a substrate on one side of the computing chip and connected to each computing chip via substrate wiring. The input / output chips are packaged as a second chip and have no direct communication path with one or more of the memory chips.
2. The chip system according to claim 1, characterized in that, The computing chip includes one or more of a central processing unit, a graphics processing unit, and a neural network processor. The memory chips include: high-bandwidth memory chips and double data rate memory chips.
3. The chip system according to claim 1, characterized in that, It also includes solder balls located between the computing chip and the substrate, and between the input / output chip and the substrate; The solder balls are used to connect the computing chip and the substrate, as well as the input / output chip and the substrate, and the substrate wiring passes through the solder balls to connect the input / output chip and the computing chip.
4. The chip system according to claim 1, characterized in that, The computing chip also transmits a first semantic tag carried by a memory access request through a multiplexed through-silicon via channel. The first semantic tag is parsed and executed by the memory controller, which is located within the first chip and connected to the one or more memory chips. Specifically, when the destination address of the memory access request points to the first storage space, the first semantic tag is generated, and the first storage space is the location of the one or more memory chips.
5. The chip system according to claim 4, characterized in that, The input / output core includes a routing control unit configured to perform the following operations in response to a second semantic tag carried by a memory access request of the computing core: in write merge buffer mode, aggregating multiple write requests and sending them to the memory controller; in prefetch buffer mode, prompting for prefetching data to the cache based on the access mode. The computing chip sends a memory access request carrying a second semantic tag to the routing control unit through the substrate semantic channel, and generates the second semantic tag when the destination address of the memory access request points to the second storage space. The second storage space can be different from the first storage space.
6. The chip system according to claim 5, characterized in that, The second semantic tag includes: operation mode identifier and parameter field information; The routing control unit includes: The semantic decoding module is configured to receive memory access requests from the computing chip, parse the second semantic tag carried by the request, and output a mode identifier and parameter field. The mode identifier includes: write merging mode and prefetch mode. The address matching module, coupled to the semantic decoding module, is configured to determine whether the target address of the memory access request falls within the accessible storage space based on the target address of the memory access request and the configurable storage space range in the programmable register group, and output a hit signal. The feature analysis module, coupled to the semantic decoding module, is configured to generate suggestion signals based on packet size, request type, and historical access trajectory. The arbitration module, coupled to the semantic decoding module, the address matching module, and the feature analysis module, is configured to, in response to the validity of the pattern identifier, adopt the processing mode of the memory access request corresponding to the pattern identifier according to a first priority; in response to the first memory space hit signal output by the address matching module, adopt the prefetch mode according to a second priority; and in response to the output of the second memory space hit signal, adopt the write merging mode according to a third priority; and in response to the suggestion signal, adopt the suggestion mode of the suggestion signal according to a fourth priority.
7. The chip system according to claim 5, characterized in that, The operation of the write merge buffer mode includes: collecting unaligned write requests within a time window and merging them into burst write operations after aligning them according to DRAM page boundaries; The operation of the prefetch buffer mode includes: calculating the prefetch step size according to the access mode prompt, and continuously reading subsequent data blocks according to the step size and storing them into the cache.
8. The chip system according to claim 1, characterized in that, Also includes: A programmable interconnect layer is disposed between corresponding computing chips and memory chips, the programmable interconnect layer including: a dynamic memory pooling module configured to manage heterogeneous memory resources, the heterogeneous memory resources being composed of multiple memory chips; The dynamic memory pooling module is configured to generate a unified resource vector based on the physical topology location attributes, bandwidth, latency, durability, and energy consumption characteristics of the heterogeneous memory resources; allocate, release, and map the unified resource vector using cache lines as the basic unit; perform topology-aware allocation based on the relative positional relationship between the computing kernels and the memory kernels, prioritizing the allocation of frequently accessed data to the physical memory region with the lowest access latency; divide the physical memory region into time slices and perform time-sharing multiplexing between different tasks, with the time-sharing multiplexing process coordinated with hardware context switching operations; monitor energy consumption status and migrate cold data to low-power memory regions or place idle memory kernels into a low-power state according to predefined energy efficiency strategies; perform data striping and verification calculations within the memory pooling layer; in response to the detection of memory kernel failures, perform online data recovery using verification information and redundant data, and mark and isolate faulty memory kernels.
9. A memory scheduling method, characterized in that, The memory scheduling method, applied to the three-dimensional stacked chip system according to any one of claims 1 to 8, comprises: In response to a memory access request sent by a computing chip, a parsing operation is performed; In response to the parsing result of the parsing operation, the following operations are performed: in write merge buffer mode, multiple write requests are aggregated and sent to the memory controller, which is located in the first chip; in prefetch buffer mode, data is prefetched to the cache based on the access mode prompt.
10. An electronic device, characterized in that, include: The three-dimensional stacked chip system as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Chip based on Chiplet architecture and control method
CN115617739A
Core particle system memory controller layout optimization method
CN118332999A
Implementation method of multi-core-particle 3D chip
CN118504496A
Three-dimensional reconfigurable hardware acceleration core chip
CN119441130A
Chip, chip system, processor and computing device
CN119759813A
Cited By
Chip system, data transmission method and related equipment
CN121919165A