Memory controller with pseudo channel support
The memory controller architecture addresses inefficiencies in HBM systems by implementing separate pseudo channel pipelines and arbiters for efficient command scheduling, enhancing performance and power efficiency.
Patent Information
- Application Number
- JP2024575326
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-06-24
- Filing Date
- 2023-06-09
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2043-06-09
AI Technical Summary
The challenge of efficiently scheduling memory access commands across pseudo channels in high-bandwidth memory (HBM) systems, which operate semi-independently but share a command bus, leading to inefficiencies in memory controller operations.
A memory controller architecture that includes separate pseudo channel pipeline circuits for each pseudo channel, with an arbiter that selectively routes and prioritizes memory access requests based on predetermined criteria, allowing for independent command execution and efficient scheduling.
Enhances memory access efficiency by enabling independent command execution across pseudo channels, optimizing bandwidth utilization and reducing power consumption while maintaining compatibility with existing standards.
Smart Images

Figure 0007802968000001 
Figure 0007802968000002 
Figure 0007802968000003
Abstract
Description
[Background technology]
[0001] Modern dynamic random-access memory (DRAM) provides high memory bandwidth by increasing the speed of data transmission on the bus connecting the DRAM to one or more data processors, such as a graphics processing unit (GPU) or central processing unit (CPU). DRAM is typically inexpensive and dense, allowing for large amounts of DRAM to be integrated per device. Most DRAM chips sold today conform to various double data rate (DDR) DRAM standards promoted by the Joint Electron Devices Engineering Council (JEDEC). Typically, several DDR DRAM chips are combined on a single printed circuit board to form a memory module that is relatively fast but also offers scalability.
[0002] Recently, a new form of memory known as high-bandwidth memory (HBM) has emerged. HBM promises to increase memory speed and bandwidth by integrating vertically stacked memory dies, allowing for a larger number of data signal lines and shorter signal traces, enabling faster operation. A key feature of newer HBM memory is known as pseudo channel mode. Pseudo channel mode divides the channel into two individual sub-channels that operate semi-independently. The pseudo channels share a command bus but execute commands independently. While pseudo channel mode allows for more flexibility, it also creates challenges for the memory controller to efficiently schedule accesses across each pseudo channel at high command rates. [Brief explanation of the drawings]
[0003] [Figure 1]FIG. 1 is a block diagram of a data processing system according to some embodiments. [Figure 2] 1 is a block diagram of a memory controller known in the prior art, according to some embodiments. [Figure 3] 2 is a block diagram illustrating a memory controller that can be used as the memory controller of FIG. 1 according to some embodiments. [Figure 4] 2 is a block diagram illustrating another memory controller that can be used as the memory controller of FIG. 1 according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0004] In the following description, the use of the same reference numerals in different figures indicates similar or identical items. Unless otherwise stated, the word "coupled" and its related verb include both direct and indirect electrical connections by means known in the art, and unless otherwise stated, any reference to a direct connection also refers to alternative embodiments using a suitable form of indirect electrical connection.
[0005] A data processor accesses a memory having a first pseudo channel and a second pseudo channel, the data processor including at least one memory access agent that generates memory access requests, a memory controller that selectively uses the first pseudo channel pipeline circuit and the second pseudo channel pipeline circuit to provide memory commands to the memory in response to the normalized requests, and a data fabric that selectively converts the memory access requests into the normalized requests for the first pseudo channel and the second pseudo channel.
[0006] The memory controller provides memory commands to a physical interface circuit for a memory having a first pseudo channel and a second pseudo channel, the memory controller comprising a first pseudo channel pipeline circuit and a second pseudo channel pipeline circuit, respectively, comprising: a front-end interface circuit for converting normalized addresses into decoded addresses of decoded memory access requests; a command queue coupled to the front-end interface circuit for storing the decoded memory access requests; and an arbiter for selecting among the decoded memory access requests from the command queue according to predetermined criteria and providing the selected memory access request at its output.
[0007] A method for a data processor to provide commands to a memory having a first pseudo channel and a second pseudo channel includes generating memory access requests and memory access responses. The memory access requests are selectively routed between upstream and downstream ports of a data fabric. The memory access responses are selectively routed between downstream and upstream ports of the data fabric. Decoding the memory access requests in the data fabric according to one of a plurality of pseudo channels including the first pseudo channel and the second pseudo channel. The first memory access request of the first pseudo channel is decoded in a first decoding and command arbitration circuit, and the second memory access request of the second pseudo channel is decoded in a second decoding and command arbitration circuit independent of the first decoding and command arbitration circuit.
[0008] 1 is a block diagram of a data processing system 100 according to some embodiments. Data processing system 100 includes a data processor in the form of a system-on-chip (SOC) 110 and memory 180 in the form of low-power high-bandwidth memory, version 3 (HBM3). In other embodiments, memory 180 may be HBM, version 4 (HBM4) memory or other types of memory that implement pseudo channels. Many other components of an actual data processing system are typically present but are not relevant to understanding this disclosure and are not shown in FIG. 1 for ease of illustration.
[0009] SOC 110 generally includes system management unit (SMU) 111, system management network (SMN) 112, central processing unit (CPU) core complex 120 labeled "CCX," graphics controller 130 labeled "GFX," real-time client subsystem 140, memory / client subsystem 150, data fabric 160, memory channel 170, and peripheral component interface express (PCIe) subsystem 190. As will be appreciated by those skilled in the art, SOC 110 may not have all of these elements present in all embodiments and may further have additional elements included therein.
[0010] SMU 111 is bidirectionally connected to major components within SOC 110 via SMN 112. SMN 112 forms the control fabric for SOC 110. SMU 111 is a local controller that controls the operation of resources on SOC 110 and synchronizes communication between them. SMU 111 manages the power-up sequencing of various processors on SOC 110 and controls multiple off-chip devices via reset, enable, and other signals. SMU 111 includes one or more clock sources (not shown), such as phase-locked loops (PLLs), to provide clock signals to each of the components of SOC 110. SMU 111 can also receive measured power consumption values from the CPU cores in CPU core complex 120 and graphics controller 130 to manage power and determine appropriate P-states for the various processors and other functional blocks.
[0011] CPU core complex 120 includes a set of CPU cores, each of which is bidirectionally connected to SMU 111 via SMN 112. Each CPU core may be a single core that shares only a last-level cache with other CPU cores, or may be combined with some, but not all, of the other cores in a cluster.
[0012] Graphics controller 130 is bidirectionally connected to SMU 111 via SMN 112. Graphics controller 130 is a high-performance graphics processing unit capable of performing graphics operations such as vertex processing, fragment processing, shading, and texture blending in a highly integrated and parallel manner. Graphics controller 130 requires periodic access to external memory to perform its operations. In the embodiment shown in FIG. 1, graphics controller 130 shares a common memory subsystem with the CPU cores in CPU core complex 120, an architecture known as a unified memory architecture. Because SOC 110 includes both a CPU and a GPU, it is also referred to as an accelerated processing unit (APU).
[0013] Real-time client subsystem 140 includes a set of real-time clients, such as representative real-time clients 142 and 143, and a memory management hub 141, labeled "MM HUB." Each real-time client is bidirectionally connected to SMU 111 via SMN 112 and to memory management hub 141. A real-time client can be any type of peripheral controller that requires the periodic movement of data, such as an image signal processor (ISP), an audio coder-decoder (CODEC), a display controller that renders and rasterizes objects generated by graphics controller 130 for display on a monitor, etc.
[0014] Memory / client subsystem 150 includes a set of memory elements or peripheral controllers, such as representative memory / client devices 152 and 153, and a system and input / output hub 151, labeled "SYSHUB / IOHUB." Each memory / client device is bidirectionally connected to SMU 111 via SMN 112 and to system and input / output hub 151. A memory / client device is a circuit that stores or requires access to data aperiodically, such as non-volatile memory, static random-access memory (SRAM), external disk controllers such as a Serial Advanced Technology Attachment (SATA) interface controller, a universal serial bus (USB) controller, a system management hub, etc.
[0015] The data fabric 160 is an interconnect that controls the flow of traffic in the SOC 110. The data fabric 160 is bidirectionally connected to the SMU 111 via the SMN 112, and to the CPU core complex 120, graphics controller 130, memory management hub 141, and system and input / output hub 151. The data fabric 160 includes a crossbar switch for routing memory-mapped access requests and responses between various devices in the SOC 110. The data fabric also includes a system memory map defined by the basic input / output system (BIOS) to determine the destination of memory accesses based on the system configuration, as well as buffers for each virtual connection. The data fabric also includes a pseudo-channel decoder circuit 161 for providing memory access requests to two downstream ports. The pseudo channel decoder circuit 161 has an input for receiving a memory access request including a multi-bit address labeled "ADDRESS", a first output connected to a first downstream port, a second output connected to a second downstream port, and a control input for receiving bits of the ADDRESS labeled "PC".
[0016] Memory channel 170 is a circuit that controls the transfer of data to and from memory 180. Memory channel 170 includes last level cache 171 labeled "LLC0," last level cache 172 labeled "LLC1," memory controller 173, and physical interface circuit 174 labeled "PHY" connected to memory 180. Last level cache 171 is bidirectionally connected to SMU 111 via SMN 112 and has an upstream port connected to a first downstream port of data fabric 160 and a downstream port. Last level cache 172 is bidirectionally connected to SMU 111 via SMN 112 and has an upstream port connected to a second downstream port of data fabric 160 and a downstream port. Memory controller 173 has a first upstream port connected to the downstream port of last level cache 171, a second upstream port connected to the downstream port of last level cache 172, a first downstream port, and a second downstream port. The physical interface circuit 174 has a first upstream port bidirectionally connected to a first downstream port of the memory controller 173, a second upstream port bidirectionally connected to a second downstream port of the memory controller 173, and a downstream port bidirectionally connected to the memory 181.The downstream port of physical interface circuit 174 conforms to the HBM3 standard and includes a set of differential pairs of command clock signals for each channel and / or pseudo channel collectively labeled "CK," a set of differential pairs of write data strobe signals for each channel and / or pseudo channel collectively labeled "WDQS," a set of differential pairs of read data strobe signals for each channel and / or pseudo channel collectively labeled "RDQS," a set of row command signals for each channel and / or pseudo channel collectively labeled "ROW COM," a set of column command signals for each channel and / or pseudo channel collectively labeled "COL COM," a set of data signals for pseudo channel 0 for each channel labeled "DQ[31:0]," and a set of data signals for pseudo channel 1 for each channel labeled "DQ[63:32]."
[0017] Peripheral Component Interface Express (PCIe) subsystem 190 includes PCIe controller 191 and PCIe physical interface circuit 192. PCIe controller 191 is bidirectionally connected to SMU 111 via SMN 112, and has upstream and downstream ports bidirectionally connected to the system and input / output hub 151. PCIe physical interface circuit 192 has upstream ports bidirectionally connected to PCIe controller 191 and downstream ports bidirectionally connected to a PCIe fabric, not shown in FIG. 1 . PCIe controller 191 can form a PCIe root complex of the PCIe system for connection to a PCIe network, including PCIe switches, routers, and devices.
[0018] In operation, SOC 110 integrates a complex assortment of computing and storage devices on a single chip, including CPU core complex 120 and graphics controller 130. Most of these controllers are well known and will not be described further. SOC 110 includes multiple internal buses for conducting data between these circuits at high speeds. For example, CPU core complex 120 accesses data via a high-speed 32-bit bus through an upstream port of data fabric 160. Data fabric 160 multiplexes accesses between any of multiple memory access agents connected to its upstream port and memory access responders connected to its downstream port. Due to the large number of memory access agents and memory access responders, the number of internal bus lines is similarly large, and a crossbar switch within data fabric 160 multiplexes these wide buses to form virtual connections between memory access requesters and memory access responders.
[0019] The various processing nodes also maintain their own cache hierarchies. In a typical configuration, CPU core complex 120 includes four CPU cores, each with its own dedicated level 1 (L1) and level 2 (L2) caches, and a level 3 (L3) cache shared among the four CPU cores in the cluster. In this example, last level caches 171 and 172 form a level 4 (L4) cache, but operate as last level caches in the cache hierarchy, regardless of the internal organization of the cache hierarchy within CPU core complex 120. In one example, last level caches 171 and 172 implement inclusive caches, where any cache line stored in any higher level cache within SOC 110 is also stored in last level caches 171 and 172. In another example, last level caches 171 and 172 are victim caches, containing cache lines, each of which contained data requested by a data processor at an earlier point in time but eventually became the most recently used cache line and has been evicted from all higher level caches.
[0020] Data processing system 100 uses HBM3, an emerging type of memory that offers the opportunity to increase memory bandwidth over other types of DDR SDRAM. HBM uses vertically stacked memory dies that communicate using short through-silicon-via (TSV) technology, which reduces the physical distance between the processor and memory and enables faster operating speeds with larger bus widths. In particular, HBM3 supports "pseudo channels," which allow different portions of the memory die to be accessed independently using separate data paths with common command and address paths. Because command and address signals are provided at a lower frequency than data signals, command and address information can be sent from the memory controller to the HBM3 memory stack between two pseudo channels in an interleaved manner. In other embodiments, data processing system 100 can use HBM, version 4 (HBM4), memory, or other types of memory that implement pseudo channels.
[0021] As described further below, SOC 110 uses an architecture in which data fabric 160 is pseudo-channel aware. For example, in one embodiment, physical addresses provided by memory accessing agents such as CPU core complex 120 or graphics controller 130 are mapped to different pseudo-channels based on a single address bit. This feature allows data fabric 160 to segregate accesses based on the accessed pseudo-channel by examining a single bit of the physical address and routing (demultiplexing) them to specific downstream ports based on the accessed pseudo-channel. Memory controller 173 receives normalized addresses on each of its upstream ports that do not include the PC bit.
[0022] This pseudo-channel awareness allows the memory controller 173 to independently select accesses for each pseudo-channel, so that the accesses are reordered and prioritized based on the particular access patterns in each pseudo-channel. Prioritization can be achieved by an arbiter for each pseudo-channel using the same arbitration rules, so that the arbitration and other circuitry associated with each pseudo-channel can be easily replicated without requiring redesign.
[0023] Physical interface circuit 174 then receives accesses on upstream ports corresponding to different timing slots and provides them to the HBM3 die in a deserialized manner using HBM3 pseudo channel interleaving. As described below, deserialization can occur in either PHY 174 or memory controller 173 as shown in FIG. 1.
[0024] 2 is a block diagram of a memory controller 200 known in the prior art. The memory controller 200 includes a memory channel controller 210 and a power controller 250. The memory channel controller 210 generally includes an interface 212, a queue 214, a command queue 220, an address generator 222, a content addressable memory (CAM) 224, a replay queue 230, a refresh logic block 232, a timing block 234, a page table 236, an arbiter 238, an error correction code (ECC) check block 242, an ECC generation block 244, and a data buffer (DB) 246.
[0025] Interface 212 has a first bidirectional connection to the data fabric via an external bus and has an output. In memory controller 200, this external bus conforms to the Advanced Extensible Interface Version 4 specified by ARM Holdings, PLC of Cambridge, UK, known as "AXI4," although other embodiments may be other types of interfaces. Interface 212 translates memory access requests from a first clock domain known as the FCLK (or MEMCLK) domain to a second clock domain internal to memory controller 200, known as the UCLK domain. Similarly, queue 214 provides memory access requests from the UCLK domain to the DFICLK domain associated with the DFI interface.
[0026] Address generator 222 decodes addresses of memory access requests received from the data fabric via the AXI4 bus. Memory access requests include access addresses in the physical address space expressed as normalized addresses. Address generator 222 converts the normalized addresses into a format that can be used to address actual memory devices in the memory system and to efficiently schedule associated accesses. This format includes a region identifier that associates the memory access request with a particular rank, row address, column address, bank address, and bank group. At startup, the system BIOS interrogates memory devices in the memory system to determine their size and configuration and programs a set of configuration registers associated with address generator 222. Address generator 222 uses the configuration stored in the configuration registers to convert the normalized addresses into the appropriate format. Command queue 220 is a queue of memory access requests received from memory access agents in data processing system 100, such as CPU cores in CPU core complex 120 and graphics controller 130. Command queue 220 stores address fields decoded by address generator 222 and other address information that allows arbiter 238 to efficiently select memory accesses, including access type and quality of service (QoS) identifiers. CAM 224 contains information for enforcing ordering rules, such as write after write (WAW) and read after write (RAW) ordering rules.
[0027] Replay queue 230 is a temporary queue for storing memory accesses selected by arbiter 238 that are waiting for a response, such as an address and command parity response, a write cyclic redundancy check (CRC) response for DDR4 DRAM, or a write and read CRC response for GDDR5 DRAM. Replay queue 230 accesses ECC check block 242 to determine whether the returned ECC is correct or indicates an error. Replay queue 230 allows the access to be replayed in the event of a parity or CRC error on one of these cycles.
[0028] The refresh logic block 232 includes state machines for various power-down, refresh, and termination resistor (ZQ) calibration cycles that are generated separately from normal read and write memory access requests received from memory access agents. For example, when a memory rank is in precharge power-down, the refresh logic must be periodically activated to perform refresh cycles. The refresh logic block 232 periodically generates auto-refresh commands to prevent data errors caused by charge leakage from the storage capacitors of memory cells in the DRAM chip. In addition, the refresh logic block 232 periodically calibrates ZQ to prevent on-die termination resistor mismatches due to thermal changes in the system. The refresh logic block 232 also determines when to place the DRAM device in different power-down modes.
[0029] Arbiter 238 is bidirectionally connected to command queue 220 and is the heart of memory channel controller 210. Arbiter 238 improves efficiency by intelligently scheduling accesses to improve memory bus utilization. Using timing block 234, arbiter 238 enforces the proper timing relationships by determining whether a particular access in command queue 220 is eligible for issue based on DRAM timing parameters. For example, each DRAM has a "t RCThere is a minimum specified time between activate commands to the same bank, known as a "minimum time between activate commands to the same bank." Timing block 234 maintains a set of counters that determine eligibility based on this and other timing parameters defined in the JEDEC specification, and is bidirectionally connected to replay queue 230. Page table 236 maintains state information about active pages in each bank and rank of the memory channel for arbiter 238, and is bidirectionally connected to replay queue 230.
[0030] In response to a write memory access request received from interface 212, ECC generation block 244 calculates an ECC according to the write data. DB 246 stores the write data and ECC for the received memory access request. The data buffer outputs the combined write data / ECC to queue 214 when arbiter 238 selects the corresponding write access for dispatch to the memory channel.
[0031] Power controller 250 includes an interface 252 to an Advanced Extensible Interface, version 1 (AXI), an APB interface 254, and a power engine 260. Interface 252 has a first bidirectional connection to the SMN, including an input for receiving an event signal labeled "EVENT_n," shown separately in FIG. 5, and an output. APB interface 254 has an input connected to the output of interface 252 and an output for connecting to the PHY via the APB. Power engine 260 has an input connected to the output of interface 252 and an output connected to the input of queue 214. Power engine 260 includes a set of configuration registers 262, a microcontroller (μC) 264, a self refresh controller (SLFREF / PE) 266, and a reliable read / write training engine (RRW / TE) 268. Configuration registers 262 are programmed via the AXI bus and store configuration information for controlling the operation of various blocks within memory controller 200. Accordingly, configuration registers 262 have outputs connected to these blocks that are not shown in detail in FIG. 2. Self-refresh controller 266 is an engine that allows for manual generation of refreshes in addition to the automatic generation of refreshes by refresh logic block 232. Reliable read / write training engine 268 provides a continuous stream of memory accesses to memory or I / O devices for purposes such as DDR interface read latency training and loopback testing.
[0032] The memory channel controller 210 includes circuitry that enables it to select memory accesses for dispatch to the associated memory channel. To make the desired arbitration decisions, the address generator 222 decodes address information into pre-decoded information, including rank, row address, column address, bank address, and bank group within the memory system, and the command queue 220 stores the pre-decoded information. The configuration registers 262 store configuration information that determines how the address generator 222 decodes the received address information. The arbiter 238 uses the decoded address information, timing eligibility information indicated by the timing block 234, and active page information indicated by the page table 236 to efficiently schedule memory accesses while adhering to other criteria, such as QoS requirements. For example, the arbiter 238 implements prioritization of accesses to open pages to avoid the overhead of precharge and activation commands required to change memory pages, and hides overhead accesses to one bank by interleaving them with read and write accesses to another bank. In particular, during normal operation, the arbiter 238 may decide to keep pages open in different banks until those pages need to be precharged before selecting a different page.
[0033] Figure 3 is a block diagram illustrating a memory controller 300 that can be used as memory controller 173 of Figure 1, according to some embodiments. Memory controller 300 generally includes a front-end interface stage 310, a DRAM command queue stage 320, an arbiter stage 330, a back-end queue stage 340, and a refresh logic circuit 350. Memory controller 300 includes other elements similar to those shown in Figure 2, but those circuit blocks are not important to understanding the disclosed embodiments and will not be described further.
[0034] Front-end interface stage 310 is a circuit that includes front-end interface circuits 311 and 312, each labeled "FEI." Front-end interface circuit 311 has an upstream port connected to a first downstream port of data fabric 160 and a downstream port. In the embodiment of FIG. 3, the upstream port uses an interface known as a scalable data port (SDP) and is therefore labeled "SDP PC0," while the downstream port makes memory access requests to pseudo channel 0 and is therefore labeled "PC0." Front-end interface circuit 312 has an upstream port connected to a second downstream port of data fabric 160 labeled "SDP PC1," and a downstream port labeled "PC1."
[0035] DRAM command queue stage 320 is a circuit that includes DRAM command queues 321 and 322, each labeled "DCQ." DRAM command queue 321 has an upstream port connected to the downstream port of front-end interface circuit 311 and a downstream port similarly labeled "PC0." DRAM command queue 322 has an upstream port connected to the downstream port of front-end interface circuit 312 and a downstream port similarly labeled "PC1."
[0036] Arbiter stage 330 is a circuit including arbiters 331 and 332, each labeled "ARB," and pseudo-channel arbiter 333, labeled "PCARB." Arbiter 331 has a first upstream port connected to the downstream port of DRAM command queue 321, a second upstream port, and a downstream port similarly labeled "PC0." Arbiter 332 has a first upstream port connected to the downstream port of DRAM command queue 322, a second upstream port, and a downstream port similarly labeled "PC1." Pseudo-channel arbiter 333 has a first upstream port connected to the downstream port of arbiter 331, a second upstream port connected to the downstream port of arbiter 332, a first downstream port labeled "SLOT0," and a second downstream port labeled "SLOT1."
[0037] The back-end queue stage 340 is a circuit including back-end queues 341 and 342, each labeled "BEQ," and command replay queues 343 and 344, each labeled "REC." The back-end queue 341 has a first upstream port connected to a first downstream port of the pseudo channel arbiter 333, a second upstream port, and a downstream port connected to the physical interface circuit 174, and provides signals for a first phase labeled "PHASE 0." The back-end queue 342 has a first upstream port connected to a second downstream port of the pseudo channel arbiter 333, a second upstream port, and a downstream port connected to the physical interface circuit 174, and provides signals for a second phase labeled "PHASE 1." The command replay queue 343 has a downstream port bidirectionally connected to the second upstream port of the back-end queue 341. The command replay queue 344 has a downstream port bidirectionally connected to the second upstream port of the back-end queue 342.
[0038] The refresh logic circuit 350 has a first output connected to a second upstream port of the arbiter 331 and a second output connected to a second upstream port of the arbiter 332. In the embodiment shown in FIG. 3, the refresh logic circuit 350 generates refresh accesses to each of the arbiters 331 and 332 based on the requirements of the memory 180. The refresh logic circuit 350 includes state machines for various power-down, refresh, and termination resistor (ZQ) calibration cycles that are generated separately from normal read and write memory access requests received from memory access agents. For example, when the memory is in precharge power-down, the refresh logic must be periodically activated to perform refresh cycles. The refresh logic circuit 350 periodically generates auto-refresh commands to prevent data errors caused by charge leakage from the storage capacitors of memory cells in the DRAM chip. In addition, the refresh logic circuit 350 periodically calibrates the ZQ to prevent on-die termination resistor mismatches due to thermal changes in the system. The refresh logic circuit 350 also determines when to place the DRAM device in different power-down modes. In other embodiments, memory controller 300 may include a separate refresh logic block for each pseudo channel.
[0039] Memory controller 300 leverages the capabilities of data fabric 160 to identify the pseudo channel of an access and separate the requests. Memory controller 300 implements parallel pseudo channel pipelines that receive normalized requests from data fabric 160. Each normalized request includes a normalized address in the physical address space associated with a particular pseudo channel. Each of the parallel pseudo channel pipelines decodes its respective normalized address and maps the normalized request to its own physical address space without considering one or more pseudo channel bits of the memory access request. Thus, front-end interface circuit 311, command queue 321, and arbiter 331 form a pseudo channel pipeline circuit for PC0, and front-end interface circuit 312, command queue 322, and arbiter 332 form a pseudo channel pipeline circuit for PC1.
[0040] Memory controller 300 operates according to the UCLK signal. In some embodiments, the UCLK signal is half the frequency of the MEMCLK signal. For example, in some embodiments, the UCLK signal is 1 gigahertz (1 GHz) and the MEMCLK signal is 2 GHz. This slower UCLK speed allows arbiters 331 and 332 to resolve timing dependencies and eligibility during a single UCLK cycle. Arbiters 331 and 332 reorder accesses for efficiency using regular criteria during each UCLK cycle so that back-end queues 341 and 342 can send commands to both PC0 and PC1, respectively, in the same UCLK cycle. For activate (ACT) commands, each back-end queue extends the command for both phases. Therefore, if an ACT command is issued from PHASE 1, PHASE 0 of the next command cycle cannot be filled with any commands.
[0041] The memory controller 300 adds a special pseudo channel arbiter 333 that arbitrates between requests for pseudo channels to identify the best access pattern based on the minimum timing required for the memory. The pseudo channel arbiter 333 considers the timing constraints of each pseudo channel and prioritizes accesses into one of two command timing slots. Arbiters 331 and 332 indicate which phase of the MEMCLK signal the command should occur in. For example, if a read activation command (ACT to RD) time is filled with an odd number of clocks and the ACT command is initiated in PHASE 0, the pseudo channel arbiter 333 will issue RD from PHASE 1. If a conflict occurs, a pseudo channel different from the previous winner of the pseudo channel arbitration is selected. The back-end queues 341 and 342 then send commands to the physical interface circuit 174 to multiplex them onto the HBM3 command bus according to the selected phase of the UCLK signal based on their slot.
[0042] Memory controller 300 implements shared circuit blocks for specialized operations. For example, refresh logic circuit 350, which controls power-down, refresh, and ZQ calibration cycles, is shared between both pseudo channel pipelines for efficiency. Similarly, other blocks corresponding to blocks in register file and power controller 250 (not shown in FIG. 3 ) are commonly shared between the pseudo channels. However, each command replay queue 343 and 344 is dedicated to a specific pseudo channel and sends replays in dedicated command slots. Command replay queues 343 and 344 support only dependent replays; in response to a replay, the command replay queues stall read and write operations in both pseudo channels and in the power controller shared between the channels.
[0043] The physical interface circuit 174 serializes the two inputs received in one UCLK cycle to the higher MEMCLK rate.
[0044] By implementing parallel pipelines, memory controller 300 operates at a slower speed than required to resolve arbitration decisions and uses a lower power supply voltage to reduce power consumption, but adds an acceptable increase in die area.
[0045] FIG. 4 is a block diagram illustrating a memory controller 400 that can be used as another memory controller 173 of FIG. 1 , according to some embodiments. Memory controller 400 is configured similarly to memory controller 300 of FIG. 3 , except that memory controller 400 further includes a serializer 410 that provides memory commands to physical interface circuit 174 based on serializing accesses to a first command phase, known as PHASE 0, and a second command phase, known as PHASE 1, based on reordering of pseudo channel requests by pseudo channel arbiter 333 of FIG. 1 . Memory controller 400 is useful when the interface between memory controller 173 and physical interface circuit 174 is standardized to include a serialized pseudo channel access pattern. For example, a specification known as the DDR-PHY Interface (DFI) Specification is published by Cadence Design Systems, Inc., and allows memory controller and PHY designs to be reused between different integrated circuit manufacturers. Memory controller 400 supports interoperability with future DFI standards, provided that such standards assume pseudo channel interleaving is performed by the memory controller.
[0046] The data processor, SOC, memory controller, or portions thereof described herein may be embodied in one or more integrated circuits, any of which may be described or represented by a computer-accessible data structure in the form of a database or other data structure that may be read by a program and used directly or indirectly to fabricate the integrated circuit. For example, the data structure may be a behavioral or register transfer level (RTL) description of the hardware functionality in a high-level design language (HDL) such as Verilog or VHDL. The description may be read by a synthesis tool that can synthesize the description to generate a netlist that includes a list of gates from a synthesis library. The netlist includes a set of gates that also represent the functionality of the hardware comprising the integrated circuit. The netlist may then be placed and routed to generate a data set that describes the geometric shapes to be applied to a mask. The mask may then be used in various semiconductor manufacturing processes to fabricate the integrated circuit. Alternatively, the database on the computer-accessible storage medium may be a netlist (with or without a synthesis library) or a data set, or Graphic Data System (GDS) II data, if desired.
[0047] While specific embodiments have been described, various modifications to these embodiments will be apparent to those skilled in the art. For example, a pseudo-channel-aware data fabric can decode a single address bit to separate accesses to different pseudo channels or to decode more complex memory mapping functions. The memory controller can use the same arbitration rules for each pseudo channel to simplify the design or to allow two pseudo channels to be accessed differently. The memory controller can perform separate arbitration for each pseudo channel but then reorder the outputs of each pseudo channel based on timing eligibility criteria to improve utilization of the pseudo channel bus. Refresh logic and related non-memory access functions may be combined into a single refresh logic block for all pseudo channels or separated into similar blocks for each pseudo channel. Furthermore, access serialization can occur in either the PHY or the memory controller.
[0048] Therefore, it is intended that the appended claims cover all modifications of the disclosed embodiments that fall within the scope of the disclosed embodiments.
Claims
1. 1. A data processor for accessing a memory having a first pseudo channel and a second pseudo channel, comprising: at least one memory access agent for generating memory access requests; a memory controller for providing memory commands to the memory in response to normalized requests using selectively a first pseudo channel pipeline circuit and a second pseudo channel pipeline circuit, each of the first pseudo channel pipeline circuit and the second pseudo channel pipeline circuit reordering and prioritizing accesses based on particular access patterns in the respective pseudo channel; a data fabric coupled between the at least one memory access agent and the memory controller for selectively converting the memory access requests to the normalized requests for the first pseudo channel pipeline circuit and the second pseudo channel pipeline circuit; a serializer operable to serialize memory commands from the first pseudo channel pipeline circuit and the second pseudo channel pipeline circuit. Data processor.
2. the data fabric converts the memory access request into the normalized request by decoding the memory access request into pseudo channel bits and a normalized address that does not include the pseudo channel bits, and provides the normalized address to a selected one of the first pseudo channel pipeline circuit and the second pseudo channel pipeline circuit.
10. The data processor of claim 1.
3. 1. A memory controller for providing memory commands to a physical interface circuit of a memory having a first pseudo channel and a second pseudo channel, comprising: a first pseudo channel pipeline circuit and a second pseudo channel pipeline circuit; a serializer; Each of the first pseudo channel pipeline circuit and the second pseudo channel pipeline circuit comprises: a front-end interface circuit for receiving a normalized address of each of the first pseudo channel pipeline circuit and the second pseudo channel pipeline circuit, the normalized address not indicating a pseudo channel, and for converting the normalized address into a decoded address of a decoded memory access request; a command queue coupled to the front-end interface circuit for storing the decoded memory access requests; an arbiter for selecting a memory access request from among the decoded memory access requests in the command queue according to predetermined criteria and providing the selected memory access request at its output; the serializer is operable to serialize memory commands from the first pseudo channel pipeline circuit and the second pseudo channel pipeline circuit; each of the first pseudo channel pipeline circuit and the second pseudo channel pipeline circuit sorting and prioritizing accesses based on a particular access pattern in the respective pseudo channel; Memory controller.
4. a back-end queue stage having a first back-end queue for a first phase of a command bus to the memory and a second back-end queue for a second phase of the command bus to the memory; The memory controller of claim 3.
5. each of the first backend queue and the second backend queue having a first input for receiving a replay command for each of the first pseudo channel and the second pseudo channel; The memory controller of claim 4.
6. the front-end interface circuit, the command queue, and the arbiter of each of the first pseudo channel and the second pseudo channel operate according to a memory controller clock signal; the back-end queue receives memory access commands from the arbiter using the memory controller clock signal and provides the memory access commands to the first phase input and the second phase input of the physical interface circuit using a memory clock signal; the memory clock signal has a higher frequency than the memory controller clock signal; The memory controller of claim 4.
7. the serializer having a first input coupled to an output of the first back-end queue, a second input coupled to an output of the second back-end queue, and an output coupled to a first input of the physical interface circuit; The memory controller of claim 4.
8. the first back-end queue has an output coupled to a first input of the physical interface circuit; the second back-end queue has an output coupled to a second input of the physical interface circuit; the physical interface circuit serializes the output of the first back-end queue and the output of the second back-end queue before providing the memory commands of the first pseudo channel and the second pseudo channel to the memory. The memory controller of claim 4.
9. 1. A method for use in a data processing system having a memory implementing a pseudo channel including a first pseudo channel and a second pseudo channel, the method comprising: generating a memory access request; routing the memory access request to a selected one of a first pseudo channel pipeline circuit and a second pseudo channel pipeline circuit; selecting a memory access request from among the memory access requests of the first pseudo channel pipeline circuit using a first arbiter, and selecting a memory access request from among the memory access requests of the second pseudo channel pipeline circuit using a second arbiter; providing a memory command to a memory in response to the memory access request selected by the first pseudo channel pipeline circuit and the second pseudo channel pipeline circuit; method.
10. 1. A method for a data processor to provide commands to a memory having a first pseudo channel and a second pseudo channel, comprising: generating a memory access request; selectively routing the memory access requests between upstream and downstream ports of a data fabric and selectively routing memory access responses between the downstream and upstream ports of the data fabric; decoding the memory access request of the data fabric according to one of a plurality of pseudo channels including the first pseudo channel and the second pseudo channel based on a pseudo channel address of the memory access request; processing a first memory access request on the first pseudo channel in a first decoding and command arbitration circuit; processing second memory access requests of the second pseudo channel in a second decoding and command arbitration circuit, wherein processing the first memory access requests and processing the second memory access requests includes reordering and prioritizing accesses based on particular access patterns in each pseudo channel; providing memory commands to the memory based on serializing the memory commands processed by the first decoding and command arbitration circuit and the memory commands processed by the second decoding and command arbitration circuit; method.
11. 1. A data processing system comprising: a memory implementing a pseudo channel including a first pseudo channel and a second pseudo channel; a data processor coupled to the memory; The data processor generating a memory access request; routing the memory access request to a selected one of a first pseudo channel pipeline circuit and a second pseudo channel pipeline circuit; selecting a memory access request from among the memory access requests of the first pseudo channel pipeline circuit using a first arbiter, and selecting a memory access request from among the memory access requests of the second pseudo channel pipeline circuit using a second arbiter; providing a memory command to a memory in response to the memory access request selected by the first pseudo channel pipeline circuit and the second pseudo channel pipeline circuit; configured to: Data processing system.
12. the data processor provides the memory command to the memory using a common address and data path of the first pseudo channel and the second pseudo channel.
12. The data processing system of claim 11.
13. the data processor independently selects memory access requests in the first pseudo channel pipeline circuit and the second pseudo channel pipeline circuit using the same arbitration rules; 12. The data processing system of claim 11.
14. the data processor independently selects memory access requests in the first pseudo channel pipeline circuit and the second pseudo channel pipeline circuit using different arbitration rules; 12. The data processing system of claim 11.
15. the data processor comprising a serializer operable to serialize memory commands from the first pseudo channel pipeline circuit and the second pseudo channel pipeline circuit; 12. The data processing system of claim 11.
Citation Information
Patent Citations
Adaptive memory access management
US11360897B1
Storage device and method of operating the same
US20170031631A1
Memory controller with virtual controller mode
US20180018105A1
Memory controller with flexible address decoding
US20180019006A1
Multiple algorithmic pattern generator testing of a memory device
US20200194090A1