Memory Controller with Pseudo-Channel Support
The memory controller architecture addresses the challenge of efficient command scheduling in HBM pseudo-channels by using separate virtual channel pipelines, enhancing access efficiency and reducing power consumption.
Patent Information
- Application Number
- JP2024575326
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-24
- Filing Date
- 2023-06-09
- Publication Date
- 2025-07-03
- Estimated Expiration
- 2043-06-09
AI Technical Summary
The challenge of efficiently scheduling memory access at a high command rate in pseudo-channel mode of High-Bandwidth Memory (HBM) due to the need for independent command execution in each sub-channel, which complicates memory controller operations.
A memory controller architecture that utilizes separate virtual channel pipeline circuits for each pseudo-channel, incorporating a front-end interface, command queue, arbiter, and data fabric to selectively convert and route memory access requests, enabling independent command execution and prioritization based on access patterns.
Enhances memory access efficiency by allowing independent scheduling and prioritization of commands in each pseudo-channel, optimizing bandwidth utilization and reducing power consumption while maintaining compatibility with existing standards.
Smart Images

Figure 2025520672000001_ABST
Abstract
Description
Background Art
[0001] Modern dynamic random-access memory (DRAM) provides high memory bandwidth by increasing the speed of data transmission on a bus that connects the DRAM to one or more data processors such as a graphics processing unit (GPU) and a central processing unit (CPU). DRAM is typically inexpensive and high-density, thereby enabling the integration of a large amount of DRAM per device. Most DRAM chips sold today are compatible with various double data rate (DDR) DRAM standards promoted by the Joint Electron Devices Engineering Council (JEDEC). Typically, several DDR DRAM chips are combined on a single printed circuit board to form a memory module that is not only relatively fast but also provides scalability.
[0002] Recently, a new form of memory known as high-bandwidth memory (HBM) has emerged. HBM promises to increase memory speed and bandwidth by integrating vertically stacked memory dies that allow for more data signal lines and shorter signal traces, enabling faster operation. An important feature of newer HBM memory is known as pseudo-channel mode. Pseudo-channel mode divides the channel into two individual sub-channels that operate semi-independently. The pseudo-channels share a command bus but execute commands individually. Pseudo-channel mode allows for further flexibility but presents the challenge of the memory controller efficiently scheduling access at a high command rate in each pseudo-channel.
Brief Description of the Drawings
[0003]
Figure 1
Figure 2
Figure 3
Figure 4
[0004] In the following description, the use of the same reference numerals in different drawings indicates the same or identical items. Unless otherwise noted, the word "coupled" and its related verbs include both direct and indirect electrical connections by means known in the art, and unless otherwise noted, any description of direct connection also means alternative embodiments that use suitable forms of indirect electrical connection.
[0005] The data processor accesses a memory having a first virtual channel and a second virtual channel. The data processor includes at least one memory access agent that generates a memory access request, a memory controller that selectively uses a first virtual channel pipeline circuit and a second virtual channel pipeline circuit to provide a memory command to the memory in response to a normalized request, and a data fabric that selectively converts the memory access request into a normalized request for the first virtual channel and the second virtual channel.
[0006] The memory controller provides memory commands to a physical interface circuit for a memory having a first virtual channel and a second virtual channel. The memory controller includes a front-end interface circuit for converting a normalized address to the decoded address of a decoded memory access request, a command queue coupled to the front-end interface circuit for storing the decoded memory access request, and an arbiter for selecting, according to a predetermined criterion, from among the decoded memory access requests from the command queue and providing the selected memory access request to its output. The memory controller further includes a first virtual channel pipeline circuit and a second virtual channel pipeline circuit, each having the above components.
[0007] A method for a data processor to provide commands to a memory having a first virtual channel and a second virtual channel includes generating a memory access request and a memory access response. The memory access request is selectively routed between an upstream port and a downstream port of a data fabric. The memory access response is selectively routed between the downstream port and the upstream port of the data fabric. The memory access request in the data fabric is decoded according to one of a plurality of virtual channels including the first virtual channel and the second virtual channel. The first memory access request of the first virtual channel is decoded in a first decoding and command arbitration circuit, and the second memory access request of the second virtual channel is decoded in a second decoding and command arbitration circuit independent of the first decoding and command arbitration circuit.
[0008] FIG. 1 is a block diagram of a data processing system 100 according to some embodiments. The data processing system 100 includes a data processor in the form of a system-on-chip (SOC) 110 and a low-power high-bandwidth memory, a memory 180 in the form of High Bandwidth Memory version 3 (HBM3). In other embodiments, the memory 180 may be an HBM, a High Bandwidth Memory version 4 (HBM4) memory, or another type of memory implementing a pseudo-channel. Although many other components of an actual data processing system are typically present, they are not shown in FIG. 1 for ease of explanation and are not relevant to understanding the present disclosure.
[0009] The SOC 110 generally includes a system management unit (SMU) 111, a system management network (SMN) 112, a central processing unit (CPU) core complex 120 labeled "CCX", a graphics controller 130 labeled "GFX", a real-time client subsystem 140, a memory / client subsystem 150, a data fabric 160, a memory channel 170, and a peripheral component interface express (PCIe) subsystem 190. As will be understood by those skilled in the art, the SOC 110 may not have all of these elements present in all embodiments and may further have additional elements included therein.
[0010] The SMU111 is bidirectionally connected to the main components within the SOC110 via the SMN112. The SMN112 forms a control fabric for the SOC110. The SMU111 is a local controller that controls the operation of the resources on the SOC110 and synchronizes the communication between them. The SMU111 manages the power-up sequencing of various processors on the SOC110 and controls a plurality of off-chip devices via reset, enable, and other signals. The SMU111 includes one or more clock sources (not shown), such as a phase locked loop (PLL), to provide clock signals to each of the components of the SOC110. Also, the SMU111 manages the power for various processors and other functional blocks and can receive measured power consumption values from the CPU cores and the graphics controller 130 within the CPU core complex 120 to determine appropriate P-states.
[0011] The CPU core complex 120 includes a set of CPU cores, each of which is bidirectionally connected to the SMU111 via the SMN112. Each CPU core may be a single core that shares only the last level cache with other CPU cores, or may be combined with some, but not all, of the other cores within the cluster.
[0012] The graphics controller 130 is bidirectionally connected to the SMU 111 via the SMN 112. The graphics controller 130 is a high-performance graphics processing unit capable of executing graphics operations such as vertex processing, fragment processing, shading, texture blending, etc. in a highly integrated parallel manner. The graphics controller 130 requires periodic access to external memory to execute its operations. In the embodiment shown in FIG. 1, the graphics controller 130 shares a common memory subsystem with the CPU cores within the CPU core complex 120, which is an architecture known as an integrated memory architecture. Since the SOC 110 includes both a CPU and a GPU, it is also referred to as an accelerated processing unit (APU).
[0013] The real-time client subsystem 140 includes a set of real-time clients such as representative real-time clients 142 and 143, and a memory management hub 141 labeled "MM HUB". Each real-time client is bidirectionally connected to the SMU 111 via the SMN 112 and bidirectionally connected to the memory management hub 141. A real-time client can be any type of peripheral controller that requires periodic movement of data, such as an image signal processor (ISP), an audio codec, a display controller that renders and rasterizes objects generated by the graphics controller 130 for display on a monitor, etc.
[0014] The memory / client subsystem 150 includes a set of memory elements or peripheral controllers such as representative memory / client devices 152 and 153, and a system and input / output hub 151 labeled "SYSHUB / IOHUB". Each memory / client device is bidirectionally connected to the SMU 111 via the SMN 112 and is bidirectionally connected to the system and input / output hub 151. The memory / client devices are circuits that store data aperiodically or require access to data, such as non-volatile memory, static random-access memory (SRAM), external disk controllers such as Serial Advanced Technology Attachment (SATA) interface controllers, universal serial bus (USB) controllers, system management hubs, and the like.
[0015] The data fabric 160 is an interconnect that controls the flow of traffic in the SOC 110. The data fabric 160 is bidirectionally connected to the SMU 111 via the SMN 112 and is bidirectionally connected to the CPU core complex 120, the graphics controller 130, the memory management hub 141, and the system and input / output hub 151. The data fabric 160 includes a crossbar switch for routing memory-mapped access requests and responses between any of the various devices in the SOC 110. The data fabric also includes a system memory map defined by the basic input / output system (BIOS) for determining the destination of memory accesses based on the system configuration, as well as buffers for each virtual connection. Further, it has a pseudo-channel decoder circuit 161 for providing a memory access request to two downstream ports. The pseudo-channel decoder circuit 161 has an input for receiving a memory access request that includes a multi-bit address labeled "ADDRESS", a first output connected to a first downstream port, a second output connected to a second downstream port, and a control input for receiving bits of the ADDRESS labeled "PC".
[0016] Memory channel 170 is a circuit that controls the transfer of data to and from memory 180. Memory channel 170 includes last-level cache 171 labeled "LLC0", last-level cache 172 labeled "LLC1", memory controller 173, and physical interface circuit 174 labeled "PHY" connected to memory 180. Last-level cache 171 is bidirectionally connected to SMU 111 via SMN 112 and has an upstream port connected to the first downstream port of data fabric 160 and a downstream port. Last-level cache 172 is bidirectionally connected to SMU 111 via SMN 112 and has an upstream port connected to the second downstream port of data fabric 160 and a downstream port. Memory controller 173 has a first upstream port connected to the downstream port of last-level cache 171, a second upstream port connected to the downstream port of last-level cache 172, a first downstream port, and a second downstream port. Physical interface circuit 174 has a first upstream port bidirectionally connected to the first downstream port of memory controller 173, a second upstream port bidirectionally connected to the second downstream port of memory controller 173, and a downstream port bidirectionally connected to memory 181.The downstream ports of the physical interface circuit 174 comply with the HBM3 standard and include a set of differential pairs of command clock signals for each channel and / or pseudo-channel collectively labeled "CK", a set of differential pairs of write data strobe signals for each channel and / or pseudo-channel collectively labeled "WDQS", a set of differential pairs of read data strobe signals for each channel and / or pseudo-channel collectively labeled "RDQS", a set of row command signals for each channel and / or pseudo-channel collectively labeled "ROW COM", a set of column command signals for each channel and / or pseudo-channel collectively labeled "COL COM", a set of data signals for pseudo-channel 0 for each channel labeled "DQ[31:0]", and a set of data signals for pseudo-channel 1 for each channel labeled "DQ[63:32]".
[0017] The Peripheral Component Interconnect Express (PCIe) subsystem 190 includes a PCIe controller 191 and a PCIe physical interface circuit 192. The PCIe controller 191 has an upstream port bi-directionally connected to the SMU 111 via the SMN 112 and bi-directionally connected to the system and the input / output hub 151, and a downstream port. The PCIe physical interface circuit 192 has an upstream port bi-directionally connected to the PCIe controller 191 and a downstream port bi-directionally connected to a PCIe fabric (not shown in FIG. 1). The PCIe controller 191 can form a PCIe root complex of the PCIe system for connection to a PCIe network including a PCIe switch, router, and devices.
[0018] In operation, the SOC 110 integrates a complex combination of computing devices and storage devices, including the CPU core complex 120 and the graphics controller 130, on a single chip. Most of these controllers are well-known and will not be described further. The SOC 110 includes multiple internal buses for conducting data at high speed between these circuits. For example, the CPU core complex 120 accesses data via a high-speed 32-bit bus through the upstream port of the data fabric 160. The data fabric 160 multiplexes accesses between any of a plurality of memory access agents connected to its upstream port and a memory access responder connected to its downstream port. Due to the large number of memory access agents and memory access responders, the number of internal bus lines is also very similar. The crossbar switch within the data fabric 160 multiplexes these wide buses to form virtual connections between the memory access requester and the memory access responder.
[0019] In addition, various processing nodes maintain their own cache hierarchies. In a typical configuration, the CPU core complex 120 includes four CPU cores, each CPU core having its own dedicated level 1 (L1) cache and level 2 (L2) cache, and a level 3 (L3) cache shared among the four CPU cores within the cluster. In this example, the last-level caches 171 and 172 form a level 4 (L4) cache, and operate as the last-level caches within the cache hierarchy regardless of the internal organization of the cache hierarchy within the CPU core complex 120. In one example, the last-level caches 171 and 172 implement an inclusive cache where any cache line stored in any higher-level cache within the SOC 110 is also stored in the last-level caches 171 and 172. In another example, the last-level caches 171 and 172 are victim caches, containing cache lines, each of which contained data requested by the data processor at an earlier time, but ultimately became the last-used cache line and was evicted from all upper-level caches.
[0020] The data processing system 100 uses High Bandwidth Memory 3 (HBM3), a currently emerging type of memory that provides an opportunity to increase the memory bandwidth over other types of DDR SDRAM. HBM uses vertically stacked memory dies that communicate using short through-silicon vias (TSVs) technology that reduces the physical distance between the processor and the memory and enables a larger bus width and faster operating speed. In particular, HBM3 supports "pseudo-channels" that allow different portions of the memory die to be accessed independently using separate data paths with a common command and address path. Since the command and address signals are provided at a lower frequency than the data signals, the command and address information can be transmitted from the memory controller to the HBM3 memory stack between two pseudo-channels in an interleaved manner. In other embodiments, the data processing system 100 can use HBM, version 4 (HBM4) memory, or other types of memory that implement pseudo-channels.
[0021] As further described below, the SOC 110 uses an architecture in which the data fabric 160 is pseudo-channel aware. For example, in one embodiment, the physical addresses provided by a memory access agent such as the CPU core complex 120 or the graphics controller 130 are mapped to different pseudo-channels based on a single address bit. Due to this feature, the data fabric 160 can separate accesses based on the accessed pseudo-channel by inspecting a single bit of the physical address and routing (demultiplexing) them to specific downstream ports based on the accessed pseudo-channel. The memory controller 173 receives the normalized addresses at each of its upstream ports that do not include the PC bits.
[0022] This pseudo-channel awareness enables the memory controller 173 to independently select accesses to each pseudo-channel, and as a result, the accesses are reordered and prioritized based on specific access patterns in each pseudo-channel. Prioritization can be achieved by an arbiter for each pseudo-channel using the same arbitration rules, and thus, the arbitration and other circuitry associated with each pseudo-channel can be easily replicated without the need for redesign.
[0023] The physical interface circuit 174 then receives accesses on the upstream ports corresponding to different timing slots and provides them to the HBM3 die in a deserialized manner using the pseudo-channel interleaving of HBM3. As described below, deserialization can be performed either by the PHY 174 or the memory controller 173 as shown in FIG. 1.
[0024] FIG. 2 is a block diagram of a memory controller 200 known in the prior art. The memory controller 200 includes a memory channel controller 210 and a power controller 250. The memory channel controller 210 generally includes an interface 212, a queue 214, a command queue 220, an address generator 222, a content addressable memory (CAM) 224, a replay queue 230, a refresh logic block 232, a timing block 234, a page table 236, an arbiter 238, an error correction code (ECC) check block 242, an ECC generation block 244, and a data buffer (DB) 246.
[0025] The interface 212 has a first bidirectional connection to a data fabric via an external bus and has an output. In the memory controller 200, this external bus is compatible with the Advanced eXtensible Interface version 4 specified by ARM Holdings, PLC of Cambridge, UK, known as "AXI4", but in other embodiments it may be other types of interfaces. The interface 212 converts a memory access request from a first clock domain known as the FCLK (or MEMCLK) domain to an internal second clock domain of the memory controller 200 known as the UCLK domain. Similarly, the queue 214 provides memory access from the UCLK domain to the DFICLK domain associated with the DFI interface.
[0026] The address generator 222 decodes the address of a memory access request received from the data fabric via the AXI4 bus. The memory access request includes an access address within the physical address space represented by a normalized address. The address generator 222 converts the normalized address into a format that can be used to address the actual memory devices within the memory system and to efficiently schedule the associated accesses. This format includes a region identifier that associates the memory access request with a specific rank, row address, column address, bank address, and bank group. At startup, the system BIOS queries the memory devices within the memory system to determine their sizes and configurations and programs a set of configuration registers associated with the address generator 222. The address generator 222 uses the configuration stored in the configuration registers to convert the normalized address into the appropriate format. The command queue 220 is a queue of memory access requests received from memory access agents within the data processing system 100, such as the CPU cores within the CPU core complex 120 and the graphics controller 130. The command queue 220 stores the address field decoded by the address generator 222, as well as other address information that enables the arbiter 238 to efficiently select memory accesses that include access type and quality of service (QoS) identifiers. The CAM 224 includes information for enforcing ordering rules, such as write after write (WAW) and read after write (RAW) ordering rules.
[0027] The replay queue 230 is a temporary queue for storing memory accesses selected by the arbiter 238 that are waiting for responses such as address and command parity responses, write cyclic redundancy check (CRC) responses for DDR4 DRAM, or write and read CRC responses for GDDR5 DRAM. The replay queue 230 accesses the ECC check block 242 to determine whether the returned ECC is correct or indicates an error. The replay queue 230 enables the access to be replayed in the case of a parity or CRC error in one of these cycles.
[0028] The refresh logic block 232 includes state machines for various power-down, refresh, and termination resistance (ZQ) calibration cycles that are generated separately from normal read and write memory access requests received from the memory access agent. For example, when a memory rank is in precharge power-down, the refresh logic must be periodically activated to execute a refresh cycle. The refresh logic block 232 periodically generates an auto-refresh command to prevent data errors caused by charge leakage from the storage capacitors of the memory cells within the DRAM chip. In addition, the refresh logic block 232 periodically calibrates ZQ to prevent mismatches in the on-die termination resistance due to thermal variations within the system. Also, the refresh logic block 232 determines when to put the DRAM device into different power-down modes.
[0029] The arbiter 238 is bi-directionally connected to the command queue 220 and is at the center of the memory channel controller 210. The arbiter 238 improves efficiency through intelligent scheduling of accesses to improve the use of the memory bus. The arbiter 238 enforces appropriate timing relationships by using the timing block 234 to determine whether a particular access within the command queue 220 is eligible for issue based on the DRAM timing parameters. For example, each DRAM has a "t RCIt has a minimum specified time between activation commands to the same bank, known as "___". The timing block 234 maintains a set of counters that determine eligibility based on this and other timing parameters defined by the JEDEC specification, and is connected bidirectionally to the replay queue 230. The page table 236 maintains state information regarding active pages in each bank and rank of the memory channel for the arbiter 238, and is connected bidirectionally to the replay queue 230.
[0030] In response to a write memory access request received from the interface 212, the ECC generation block 244 calculates an ECC according to the write data. The DB 246 stores the write data and the ECC regarding the received memory access request. The data buffer outputs the combined write data / ECC to the queue 214 when the arbiter 238 selects the corresponding write access for dispatch to the memory channel.
[0031] The power controller 250 includes an interface 252 to the Advanced eXtensible Interface, Version 1 (AXI), an APB interface 254, and a power engine 260. The interface 252 has a first bidirectional connection to the SMN, including an input and an output for receiving event signals labeled "EVENT_n" shown separately in FIG. 5. The APB interface 554 has an input connected to the output of the interface 252 and an output for connecting to the PHY via the APB. The power engine 260 has an input connected to the output of the interface 252 and an output connected to the input of the queue 214. The power engine 260 includes a set of configuration registers 262, a microcontroller (μC) 264, a self refresh controller (SLFREF / PE) 266, and a reliable read / write training engine (RRW / TE) 268. The configuration registers 262 are programmed via the AXI bus and store configuration information for controlling the operation of various blocks within the memory controller 200. Accordingly, the configuration registers 262 have outputs connected to these blocks not shown in detail in FIG. 2. The self refresh controller 266 is an engine that enables manual generation of refreshes in addition to the automatic generation of refreshes by the refresh logic block 232. The reliable read / write training engine 268 provides a continuous memory access stream to a memory or I / O device for purposes such as DDR interface read latency training and loopback testing.
[0032] The memory channel controller 210 includes circuitry that enables selection of memory access for dispatch to the associated memory channel. To make the desired arbitration decision, the address generator 222 decodes the address information into pre-decoded information including rank, row address, column address, bank address, and bank group within the memory system, and the command queue 220 stores the pre-decoded information. The configuration register 262 stores configuration information for determining how the address generator 222 decodes the received address information. The arbiter 238 uses the decoded address information, the timing eligibility information indicated by the timing block 234, and the active page information indicated by the page table 236 to efficiently schedule memory access while complying with other criteria such as QoS requirements. For example, the arbiter 238 implements a priority for access to open pages and hides overhead access to one bank by interleaving it with read and write accesses to another bank to avoid the overhead of the precharge and activation commands required to change memory pages. In particular, during normal operation, the arbiter 238 may determine to keep pages open in different banks until these pages need to be precharged before selecting different pages.
[0033] FIG. 3 is a block diagram showing a memory controller 300 that can be used as the memory controller 173 of FIG. 1 according to some embodiments. The memory controller 300 generally includes a front-end interface stage 310, a DRAM command queue stage 320, an arbiter stage 330, a back-end queue stage 340, and a refresh logic circuit 350. The memory controller 300 includes other elements similar to those shown in FIG. 2, but those circuit blocks are not important for understanding the disclosed embodiments and will not be further described.
[0034] The front - end interface stage 310 is a circuit that includes front - end interface circuits 311 and 312, each labeled "FEI". The front - end interface circuit 311 has an upstream port connected to the first downstream port of the data fabric 160 and a downstream port. In the embodiment of FIG. 3, the upstream port uses an interface known as a scalable data port (SDP), and thus the upstream port is labeled "SDP PC0", and the downstream port makes a memory access request for pseudo - channel 0 and is thus labeled "PC0". The front - end interface circuit 312 has an upstream port connected to the second downstream port of the data fabric 160 labeled "SDP PC1" and a downstream port labeled "PC1".
[0035] The DRAM command queue stage 320 is a circuit that includes DRAM command queues 321 and 322, each labeled "DCQ". The DRAM command queue 321 has an upstream port connected to the downstream port of the front - end interface circuit 311 and a downstream port, also labeled "PC0". The DRAM command queue 322 has an upstream port connected to the downstream port of the front - end interface circuit 312 and a downstream port, also labeled "PC1".
[0036] The arbiter stage 330 is a circuit that includes arbiters 331 and 332, each labeled "ARB", and a pseudo-channel arbiter 333 labeled "PCARB". Arbiter 331 has a first upstream port connected to the downstream port of the DRAM command queue 321, a second upstream port, and a downstream port similarly labeled "PC0". Arbiter 332 has a first upstream port connected to the downstream port of the DRAM command queue 322, a second upstream port, and a downstream port similarly labeled "PC1". The pseudo-channel arbiter 333 has a first upstream port connected to the downstream port of arbiter 331, a second upstream port connected to the downstream port of arbiter 332, a first downstream port labeled "SLOT0", and a second downstream port labeled "SLOT1".
[0037] The back-end queue stage 340 is a circuit that includes back-end queues 341 and 342, each labeled "BEQ", and command replay queues 343 and 344, each labeled "REC". The back-end queue 341 has a first upstream port connected to the first downstream port of the pseudo-channel arbiter 333, a second upstream port, and a downstream port connected to the physical interface circuit 174, and provides a signal for the first phase labeled "PHASE0". The back-end queue 342 has a first upstream port connected to the second downstream port of the pseudo-channel arbiter 333, a second upstream port, and a downstream port connected to the physical interface circuit 174, and provides a signal for the second phase labeled "PHASE1". The command replay queue 343 has a downstream port bidirectionally connected to the second upstream port of the back-end queue 341. The command replay queue 344 has a downstream port bidirectionally connected to the second upstream port of the back-end queue 342.
[0038] The refresh logic circuit 350 has a first output connected to the second upstream port of the arbiter 331 and a second output connected to the second upstream port of the arbiter 332. In the embodiment shown in FIG. 3, the refresh logic circuit 350 generates refresh accesses to each of the arbiters 331 and 332 based on the requirements of the memory 180. The refresh logic circuit 350 includes a state machine for various power-down, refresh, and termination resistance (ZQ) calibration cycles that are generated separately from the normal read and write memory access requests received from the memory access agent. For example, when the memory is in precharge power-down, the refresh logic must be periodically activated to execute a refresh cycle. The refresh logic circuit 350 periodically generates an auto-refresh command to prevent data errors caused by charge leakage from the storage capacitors of the memory cells within the DRAM chip. In addition, the refresh logic circuit 350 periodically calibrates ZQ to prevent mismatches in the on-die termination resistance due to thermal variations within the system. Also, the refresh logic circuit 350 determines when to put the DRAM device into different power-down modes. In other embodiments, the memory controller 300 can include a separate refresh logic block for each pseudo-channel.
[0039] The memory controller 300 utilizes the capabilities of the data fabric 160 to identify virtual channels of access and separate requests. The memory controller 300 implements a parallel virtual channel pipeline that receives normalized requests from the data fabric 160. Each normalized request includes a normalized address of a physical address space associated with a particular virtual channel. Each of the parallel virtual channel pipelines decodes its respective normalized address without considering one or more virtual channel bits of the memory access request and maps the normalized request to its own physical address space. Thus, the front-end interface circuit 311, command queue 321, and arbiter 331 form a virtual channel pipeline circuit for PC0, and the front-end interface circuit 312, command queue 322, and arbiter 332 form a virtual channel pipeline circuit for PC1.
[0040] The memory controller 300 operates according to the UCLK signal. In some embodiments, the UCLK signal is half the frequency of the MEMCLK signal. For example, in some embodiments, the UCLK signal is 1 gigahertz (1 GHz) and the MEMCLK signal is 2 GHz. This slower UCLK speed enables the arbiters 331 and 332 to resolve timing dependencies and eligibility within a single UCLK cycle. The arbiters 331 and 332 rearrange accesses for efficiency using normal criteria during each UCLK cycle so that the back-end queues 341 and 342 can send commands to both PC0 and PC1 in the same UCLK cycle. For an activate (ACT) command, each back-end queue extends the command for both phases. Thus, if an ACT command is issued from PHASE1, PHASE0 of the next command cycle cannot be filled with any command.
[0041] Memory controller 300 adds a special pseudo-channel arbiter 333 that arbitrates between requests to the pseudo-channels to identify the best access pattern based on the minimum timing required for the memory. The pseudo-channel arbiter 333 takes into account the timing constraints in each pseudo-channel and preferentially places accesses in one of two command timing slots. Arbiters 331 and 332 indicate in which phase of the MEMCLK signal the command turns on. For example, if the read activation command (from ACT to RD) time is filled with odd clocks and the ACT command starts at PHASE0, the pseudo-channel arbiter 333 issues RD from PHASE1. If a conflict occurs, a different pseudo-channel than the previous winner in the pseudo-channel arbitration is selected. Next, back-end queues 341 and 342 send commands to the physical interface circuit 174 and multiplex them on the HBM3 command bus according to the selected phase of the UCLK signal based on those slots.
[0042] Memory controller 300 implements shared circuit blocks for special operations. For example, the refresh logic circuit 350 that controls power-down, refresh, and ZQ calibration cycles is shared between both pseudo-channel pipelines for efficiency. Similarly, other blocks corresponding to the register file and blocks within the power controller 250 (not shown in FIG. 3) are commonly shared between the pseudo-channels. However, each of the command replay queues 343 and 344 is dedicated to a specific pseudo-channel and sends a replay in a dedicated command slot. The command replay queues 343 and 344 support only dependent replays, and in response to a replay, the command replay queues stop read and write operations in both pseudo-channels and in the power controller shared between the channels.
[0043] The physical interface circuit 174 serializes two inputs received in one UCLK cycle to a higher MEMCLK rate.
[0044] By executing parallel pipelines, the memory controller 300 operates at a slower speed required to resolve arbitration decisions and uses a lower power supply voltage to reduce power consumption, but adds an acceptable increase in die area.
[0045] FIG. 4 is a block diagram showing a memory controller 400 that can be used as another memory controller 173 of FIG. 1 according to some embodiments. The memory controller 400 is configured in the same manner as the memory controller 300 of FIG. 3, except that it further includes a serializer 410 that provides memory commands to the physical interface circuit 174 based on serializing access to a first command phase known as PHASE0 and a second command phase known as PHASE1 based on the rearrangement of virtual channel requests by the virtual channel arbiter 333 of FIG. 1. The memory controller 400 is useful when the interface between the memory controller 173 and the physical interface circuit 174 is standardized to include a serialized virtual channel access pattern. For example, a specification known as the DDR-PHY interface (DFI) specification is published by Cadence Design Systems, Inc., and enables the reuse of memory controller and PHY designs among different integrated circuit manufacturers. The memory controller 400 is compatible with future DFI specifications assuming a future DFI specification where virtual channel interleaving is performed by the memory controller.
[0046] The data processor, SOC, memory controller, or a part thereof described in this specification can be embodied in one or more integrated circuits, any of which can be read by a program and used directly or indirectly to manufacture an integrated circuit, and can be described or represented by a computer-accessible data structure in the form of a database or other data structure. For example, this data structure may be a behavioral level description or a register transfer level (RTL) description of hardware functions in a high-level design language (HDL) such as Verilog or VHDL. The description can be read by a synthesis tool that can synthesize the description to generate a netlist containing a list of gates from a synthesis library. The netlist contains a set of gates that also represent the functions of the hardware including the integrated circuit. The netlist may then be placed and routed to generate a data set that describes the geometric shapes to be applied to the mask. Next, the mask may be used in various semiconductor manufacturing processes to manufacture the integrated circuit. Alternatively, a database on a computer-accessible storage medium may, if desired, be a netlist (with or without a synthesis library) or a data set, or Graphic Data System (GDS) II data.
[0047] While specific embodiments have been described, various modifications to these embodiments will be apparent to those skilled in the art. For example, a pseudo-channel aware data fabric can decode a single address bit to separate accesses to different pseudo-channels or can decode a more complex memory mapping function. The memory controller can use the same arbitration rules for each pseudo-channel to simplify the design or to enable access to be made in different ways to two pseudo-channels. The memory controller can perform separate arbitration in each pseudo-channel and then reorder the outputs of each pseudo-channel based on timing eligibility criteria to improve utilization of the pseudo-channel bus. The refresh logic and related non-memory access functions may be integrated into a single refresh logic block for all pseudo-channels or may be separated into similar blocks for each pseudo-channel. Further, serialization of accesses can be performed either in the PHY or in the memory controller.
[0048] Accordingly, the appended claims are intended to cover all modifications of the disclosed embodiments that fall within the scope of the disclosed embodiments.
Claims
1. A data processor for accessing a memory having a first pseudo-channel and a second pseudo-channel, comprising: at least one memory access agent for generating a memory access request; a memory controller for selectively using a first pseudo-channel pipeline circuit and a second pseudo-channel pipeline circuit to provide a memory command to the memory in response to a normalized request; a data fabric for selectively converting the memory access request into the normalized request for the first pseudo-channel and the second pseudo-channel. The data processor.
2. The data fabric converts the memory access request into the normalized request by decoding the memory access request into pseudo-channel bits and a normalized address, and provides the normalized address to a selected pseudo-channel pipeline circuit of the first pseudo-channel pipeline circuit and the second pseudo-channel pipeline circuit. The data processor according to Claim 1.
3. The data fabric comprises: a pseudo-channel decoder circuit for routing the memory access request to the first pseudo-channel pipeline circuit in response to a predetermined bit of the access address being in a first logic state, and routing the memory access request to the second pseudo-channel pipeline circuit in response to the predetermined bit of the access address being in a second logic state. The data processor according to Claim 2.
4. The second pseudo-channel pipeline circuit operates independently of the first pseudo-channel pipeline circuit. The data processor according to Claim 1.
5. Each of the first pseudo-channel pipeline circuit and the second pseudo-channel pipeline circuit comprises: a front-end interface circuit coupled to the data fabric for converting a normalized address into a decoded address of a decoded memory access request; a command queue coupled to the front-end interface circuit for storing the decoded memory access request. An arbiter that selects a memory access request from among the decoded memory access requests of the command queue according to a predetermined criterion and provides the selected memory access request to its own output, The data processor according to claim 4. **Claim 6** The arbiter of the first pseudo-channel pipeline circuit uses the same criterion as the arbiter of the second pseudo-channel pipeline circuit. The data processor according to claim 5. **Claim 7** The memory controller, selects an output between the outputs of the arbiters of each of the first pseudo-channel pipeline circuit and the second pseudo-channel pipeline circuit based on memory access timing eligibility, and provides a memory access command to a first phase output for a first phase and a second phase output for a second phase, and includes a pseudo-channel arbiter for The data processor according to claim 5. **Claim 8** A physical interface circuit having a first phase input coupled to the first phase output of the pseudo-channel arbiter, a second phase input coupled to the second phase output of the pseudo-channel arbiter, and an output coupled to the memory, The physical interface circuit serializes the first memory access command of the first phase and the second memory access command of the second phase on a common command bus of both pseudo-channels. The data processor according to claim 7. **Claim 9** A first back-end queue for the first phase, having a first input coupled to the first phase output of the pseudo-channel arbiter, a second input for receiving a replay command for the first pseudo-channel, and an output, A second back-end queue for the second phase, having a first input coupled to the second phase output of the pseudo-channel arbiter, a second input for receiving a replay command for the second pseudo-channel, and an output, The data processor according to claim 7. **Claim 10** Each of the front-end interface circuits, the command queue, and the arbiter of the first decoding and command arbitration circuit and the second decoding and command arbitration circuit operate according to a memory controller clock signal. The physical interface circuit provides the memory command to the memory using a memory clock signal, wherein the memory clock signal has a higher frequency than the memory controller clock signal, The data processor according to claim 5.
11. A memory controller for providing a memory command to a physical interface circuit of a memory having a first pseudo-channel and a second pseudo-channel, comprising a first pseudo-channel pipeline circuit and a second pseudo-channel pipeline circuit, each of the first pseudo-channel pipeline circuit and the second pseudo-channel pipeline circuit, a front-end interface circuit for converting a normalized address into a decoded address of a decoded memory access request, a command queue coupled to the front-end interface circuit for storing the decoded memory access request, an arbiter for selecting a memory access request from the decoded memory access requests in the command queue according to a predetermined criterion and providing the selected memory access request to its output, Memory controller.
12. A back-end queue stage having a first back-end queue for a first phase of a command bus to the memory and a second back-end queue for a second phase of the command bus to the memory, The memory controller according to claim 11.
13. Each of the first back-end queue and the second back-end queue has a first input for receiving a replay command for each of the first pseudo-channel and the second pseudo-channel, The memory controller according to claim 12.
14. The front-end interface circuit, the command queue and the arbiter of each of the first pseudo-channel and the second pseudo-channel operate according to a memory controller clock signal, The back-end queue receives a memory access command from the arbiter using the memory controller clock signal and provides the memory access command to the first-phase input and the second-phase input of the physical interface circuit using a memory clock signal, The memory clock signal has a higher frequency than the memory controller clock signal. The memory controller of claim 12.
15. A serializer having a first input coupled to the output of the first back-end queue, a second input coupled to the output of the second back-end queue, and an output coupled to the first input of the physical interface circuit. The memory controller of claim 12.
16. The first back-end queue has an output coupled to the first input of the physical interface circuit. The second back-end queue has an output coupled to the second input of the physical interface circuit. Before the physical interface circuit provides the memory access commands of the first virtual channel and the second virtual channel to the memory, the physical interface circuit serializes the output of the first back-end queue and the output of the second back-end queue. The memory controller of claim 12.
17. A method for a data processor to provide commands to a memory having a first virtual channel and a second virtual channel, comprising: Generating a memory access request; Selectively routing the memory access request between an upstream port and a downstream port of a data fabric, and selectively routing a memory access response between the downstream port and the upstream port of the data fabric; Decoding the memory access request of the data fabric according to any one of a plurality of virtual channels including a first virtual channel and a second virtual channel; Processing a first memory access request of the first virtual channel in a first decoding and command arbitration circuit; Processing a second memory access request of the second virtual channel in a second decoding and command arbitration circuit independent of the first decoding and command arbitration circuit. Method.
18. Selectively arbitrating the outputs of the first decoding and command arbitration circuit and the second decoding and command arbitration circuit at a first phase port and a second phase port based on a memory timing reference. The method of claim 17.
19. Using a physical interface circuit to couple commands from the first phase port and the second phase port to a physical interface to the memory, The method of claim 18. **Claim 20** Processing the first memory access request of the first pseudo-channel and the second memory access request of the second pseudo-channel at a memory controller clock rate, Operating the physical interface circuit at a memory clock rate, the memory clock rate being higher than the memory controller clock rate, The method of claim 19.
Citation Information
Patent Citations
Adaptive memory access management
US11360897B1
Storage device and method of operating the same
US20170031631A1
Memory controller with virtual controller mode
US20180018105A1
Memory controller with flexible address decoding
US20180019006A1
Multiple algorithmic pattern generator testing of a memory device
US20200194090A1