Memory controller with multiple command subqueues and corresponding arbiters

The use of multiple smaller command queues and dedicated arbiters in DDR memory controllers addresses delays in conventional systems by segregating read and write requests, enhancing processing efficiency and speed.

JP7897223B2Active Publication Date: 2026-07-29ADVANCED MICRO DEVICES INC
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
ADVANCED MICRO DEVICES INC
Filing Date
2021-08-24
Publication Date
2026-07-29

AI Technical Summary

Technical Problem

Conventional DDR memory controllers face delays due to long evaluation times for prioritizing 64 entries and difficulty in handling JEDEC timing dependencies and page state information, leading to inefficiencies in processing memory read and write commands.

Method used

Implementing multiple smaller command queues and dedicated arbiters for each queue to independently select memory access requests based on predetermined criteria, improving operating speed by segregating read and write requests into separate queues.

Benefits of technology

Enhances the operating speed of memory controllers and data processing systems by reducing evaluation times and optimizing command processing through segregated command queues and arbiters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007897223000001
    Figure 0007897223000001
  • Figure 0007897223000002
    Figure 0007897223000002
  • Figure 0007897223000003
    Figure 0007897223000003
Patent Text Reader

Abstract

The memory controller includes a memory channel controller that uses multiple groups of command queues and arbiter pairs. Each arbiter is coupled to a respective command queue and selects memory access commands from each command queue according to predetermined criteria. Each arbiter independently selects from among the memory access requests in each command queue based on the predetermined criteria and sends the selected memory access requests to a selector that functions as a second-level arbiter that sends the requests to a memory sub-channel.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - Reference to Related Applications) This application claims priority to U.S. Patent Application No. 17 / 085,304, filed October 30, 2020, which claims priority to U.S. Provisional Patent Application No. 63 / 069,352, filed August 24, 2020, the entire disclosure of which is expressly incorporated herein by reference.

Background Art

[0002] Computer systems typically use inexpensive and high - density dynamic random access memory (DRAM) chips for main memory. Most DRAM chips sold today are compatible with various double - data - rate (DDR) DRAM standards promulgated by the Joint Electron Devices Engineering Council (JEDEC). DDR DRAM provides both high performance and low - power operation by offering various low - power modes.

[0003] Modern DDR memory controllers maintain queues for storing pending memory access requests and allow for the selection of these requests in no particular order relative to the order in which they were generated or stored, in order to improve efficiency. For example, a memory controller can retrieve multiple memory access requests for the same row within a given rank of memory from the queue based on page hit checks, and issue them sequentially to the memory system, thus avoiding the overhead of precharging the current row and activating another row. Some DDR memory controllers employ a single command queue, such as a 64-entry command queue, and an arbiter that mediates between all 64 command queue entries, each containing a memory access request. Data processing systems employing high-density dynamic random access memory, such as cloud computing servers, desktop computers, laptop computers, mobile devices, printers, and other devices, require higher performance capabilities than ever before.

[0004] When accompanied by the following diagrams in which similar symbols represent similar elements, the embodiments will be more readily understood with consideration of the following description. [Brief explanation of the drawing]

[0005] [Figure 1] This is a block diagram of a data processing system according to several embodiments. [Figure 2] This is a block diagram of an Accelerated Processing Unit (APU) suitable for use in the data processing system shown in Figure 1. [Figure 3] This is a block diagram of a memory controller and associated physical interface (PHY) suitable for use in the APU shown in Figure 2, according to several embodiments. [Figure 4] This is a block diagram of another memory controller and associated PHY suitable for use in the APU of Figure 2, according to several embodiments. [Figure 5]This is a block diagram of a memory controller according to several embodiments. [Figure 6] This is a partial block diagram of a memory controller employing multiple command subqueues, according to several embodiments. [Figure 7] This is a block diagram of another example of a memory controller employing separate write command subqueues and separate read command subqueues according to several embodiments. [Figure 8] Block diagram showing another memory controller using different groups of read / write subqueues, according to several embodiments. [Figure 9] This is a block diagram of another memory controller employing multiple configurations of the memory controller architecture shown in Figure 7. [Figure 10] This flowchart shows an example of a method for controlling a memory system having multiple memory channels, according to several embodiments. [Modes for carrying out the invention]

[0006] In the following description, the use of the same reference numerals in different drawings indicates the same or identical items. Unless otherwise noted, the word “combined” and its associated verb forms include both direct and indirect electrical connections by means known in the art, and unless otherwise noted, any description of a direct connection also means an alternative embodiment using a preferred form of indirect electrical connection.

[0007] It has been found that arbiters in conventional DDR memory controllers can take too long to evaluate all 64 entries for priority, JEDEC timing dependencies, and page state information. Timing is difficult and causes delays when processing memory read and write commands (also referred to herein as read access, read request, write access, and write request). As described below, in some embodiments, the memory controller includes a memory channel controller that uses multiple groups of command queues and arbiter pairs so that more, but smaller, command queues are employed. Multiple command subqueues are used as command queues for a channel or subchannel. In some embodiments, each arbiter is coupled to its respective command queue to select memory access commands from each command queue according to predetermined criteria such as DDR timing criteria and other criteria. The arbiter independently selects from among the memory access requests in each command queue based on predetermined criteria and sends the selected memory access requests to a selector that acts as a second-level arbiter to send the requests to the subchannel. In some embodiments, using multiple smaller command queues and corresponding dedicated arbiters instead of a single larger command queue and arbiter improves the operating speed of the memory controller and the corresponding data processing system.

[0008] According to some embodiments, a method for controlling a memory system having multiple memory channels includes selecting a memory access request in a first command subqueue, selecting a memory access request in a second command subqueue, selecting a memory access request from among the first and second memory access requests, and sending the selected memory access request to a memory channel.

[0009] According to some embodiments, a method for controlling a memory system having multiple memory channels includes receiving a memory access request. The method includes decoding each of the memory access requests. The method also includes storing the decoded memory access requests in a first command subqueue or a second command subqueue. The method also includes selecting from a plurality of memory access requests in the first command subqueue using predetermined criteria in order to provide a first memory access request selected from the first command subqueue. The method also includes selecting from a plurality of memory access requests in the second command subqueue using predetermined criteria in order to provide a second memory access request selected from the second command subqueue. The method also includes selecting a preferred memory access request from the first memory access request provided from the first command subqueue and the second memory access request provided from the second command subqueue. The method also includes sending the thus selected preferred memory access request to one of the plurality of memory channels according to the subchannel.

[0010] In some embodiments, the method includes sorting memory access requests into different command subqueues such that a first command subqueue contains only read requests and a second command subqueue contains only write requests. In certain embodiments, the method includes decoding memory access requests into banks, ranks, and subchannels of multiple subchannels of a memory device in a memory system, and storing the banks, ranks, and subchannels in one of the multiple command subqueues. In some embodiments, the method includes selecting preferred memory access requests by selecting the oldest memory access request from among the first memory access requests provided from the first command subqueue and the second memory access requests from the second command subqueue.

[0011] In some embodiments, the memory controller includes a memory channel controller, which includes a first command subqueue configured to store memory access requests and a corresponding first arbiter coupled to the first command subqueue to select memory access commands from the first command subqueue. The memory also includes a second command subqueue configured to store memory access requests and a corresponding second arbiter coupled to the second command subqueue to select memory access commands from the second command subqueue. The memory also includes command queue entry logic for placing memory access requests in the first and second command subqueues. The memory also includes a first selector capable of selecting a memory request from either the first or second command subqueue and sending the selected memory access request to at least one of a plurality of subchannels. In some embodiments, the arbiter selects memory access commands based on predetermined criteria. In some embodiments, the first selector is coupled to both the first and second arbiters.

[0012] In some embodiments, the command queue entry logic sorts memory access requests into different command subqueues such that a first command subqueue contains only read requests and a second command subqueue contains only write requests. In other embodiments, the command queue entry logic transfers entries from the first command subqueue to the second command subqueue.

[0013] In a particular embodiment, the memory controller includes shared timing logic and a shared page table shared between a first arbiter and a second arbiter.

[0014] In some embodiments, the memory control logic includes a third command subqueue for storing memory access requests and a corresponding third arbiter coupled to the third command subqueue for selecting memory access commands from the third command subqueue according to predetermined criteria. The memory also includes a fourth command subqueue for storing memory access requests and a corresponding fourth arbiter coupled to the fourth command subqueue for selecting memory access commands from the fourth command subqueue according to predetermined criteria. The memory also includes a second selector coupled to both the third and fourth arbiters and operable to select a memory request from either the third or fourth command subqueue. The memory also includes a third selector operablely coupled to the first and second selectors and operable to select a memory request from either the first or second selector and transmit the selected memory access request to at least one of a plurality of subchannels.

[0015] In certain embodiments, the memory controller is operably coupled to a first command subqueue, a second command subqueue, a third command subqueue, and a fourth command subqueue, and includes command queue entry logic that is operable to sort memory access requests into different command queues such that the first and second command queues contain only read requests, and the third and fourth command queues contain only write requests.

[0016] In some embodiments, the memory control logic includes a third command subqueue for storing memory access requests and a corresponding third arbiter coupled to the third command subqueue for selecting memory access commands from the third command subqueue according to predetermined criteria. The memory control logic also includes a fourth command subqueue for storing memory access requests and a corresponding fourth arbiter coupled to the fourth command subqueue for selecting memory access commands from the fourth command subqueue according to predetermined criteria. The memory control logic also includes a second selector coupled to both the third and fourth arbiters, capable of selecting a memory request from either the third or fourth command subqueue and transmitting the selected memory access request to the corresponding subchannel.

[0017] In one embodiment, a data processing system includes a plurality of memory access agents for providing memory access requests. The data processing system also includes a plurality of memory channels. The data processing system further includes a memory controller coupled to the plurality of memory access agents and the plurality of memory channels and having a memory channel controller. The memory channel controller includes a first command subqueue for storing memory access requests, and a corresponding first arbiter coupled to the first command subqueue for selecting a memory access command from the first command subqueue. The memory channel controller also includes a second command subqueue for storing memory access requests, and a corresponding second arbiter coupled to the second command subqueue for selecting a memory access command from the second command subqueue. In some embodiments, the arbiter selects a memory access command based on a predetermined criterion. The memory channel controller is also operably coupled to the first command subqueue and the second command subqueue and includes command queue entry logic operable to place memory access requests into the first command subqueue and the second command subqueue. The memory channel controller is also coupled to both the first and second arbiters and includes a first selector for selecting a memory request from either the first command subqueue or the second command subqueue and transmitting the selected memory access request to at least one of the plurality of subchannels.

[0018] In some examples, the command queue entry logic sorts memory access requests into different command subqueues such that the first command subqueue contains only read requests and the second command subqueue contains only write requests. In some embodiments, the command queue entry logic transfers entries from the first command subqueue to the second command subqueue.

[0019] In certain embodiments, the data processing system includes shared timing logic and a shared page table shared between a first arbiter and a second arbiter.

[0020] In some embodiments, the data processing system includes a third command subqueue for storing memory access requests and a corresponding third arbiter coupled to the third command subqueue to select a memory access command from the third command subqueue according to a predetermined criterion. The data processing system also includes a fourth command subqueue for storing memory access requests and a corresponding fourth arbiter coupled to the fourth command subqueue to select a memory access command from the fourth command subqueue according to a predetermined criterion. The data processing system also includes a second selector coupled to both the third and fourth arbiters and operable to select a memory request from either the third command subqueue or the fourth command subqueue. The data processing system also includes a third selector operably coupled to the first and second selectors, operable to select a memory request from either the first selector or the second selector, and transmit the selected memory access request to at least one of a plurality of subchannels.

[0021] In certain embodiments, the memory controller includes command queue entry logic operably coupled to the first command subqueue, the second command subqueue, the third command subqueue, and the fourth command subqueue, and sorting memory access requests into different command queues such that the first command subqueue and the second command queue include only read requests, and the third command subqueue and the fourth command queue include only write requests.

[0022] In some embodiments, the data processing system includes a third command subqueue for storing memory access requests and a corresponding third arbiter coupled to the third command subqueue for selecting memory access commands from the third command subqueue according to predetermined criteria. The data processing system also includes a fourth command subqueue for storing memory access requests and a corresponding fourth arbiter coupled to the fourth command subqueue for selecting memory access commands from the fourth command subqueue according to predetermined criteria. The data processing system also includes a second selector coupled to both the third and fourth arbiters, capable of selecting a memory request from either the third or fourth command subqueue and transmitting the selected memory access request to a corresponding subchannel.

[0023] Figure 1 is a non-limiting exemplary block diagram showing a data processing system 100 in several embodiments. The data processing system 100 generally includes a data processor 110 in the form of an accelerated processing unit (APU), a memory system 120, a peripheral component interconnect express (PCIe) system 150, a universal serial bus (USB) system 160, and a disk drive 170. The data processor 110 acts as the central processing unit (CPU) of the data processing system 100, providing various buses and interfaces useful in modern computer systems. These interfaces include two double data rate (DDRx) memory channels, a PCIe root complex for connection to a PCIe link, a USB controller for connection to a USB network, and an interface to a Serial Advanced Technology Attachment (SATA) mass storage device.

[0024] The memory system 120 includes memory channels 130 and 140. Memory channel 130 includes a set of dual inline memory modules (DIMMs) connected to the DDRx bus 132, including representative DIMMs 134, 136, and 138 corresponding to individual ranks in this embodiment. Similarly, memory channel 140 includes a set of DIMMs connected to the DDRx bus 142, including representative DIMMs 144, 146, and 148. For example, a DDR5 dual inline memory module (DIMM) has two independent 32-bit channels called "subchannels". From the perspective of using a DRAM controller architecture, a single controller operates two separate 32-bit channels independently, and in this case, from the controller's perspective, the two channels are also called subchannels.

[0025] The PCIe system 150 includes a PCIe switch 152 connected to the PCIe root complex in the data processor 110, as well as PCIe devices 154, 156, and 158. PCIe device 156 is connected to the basic input / output system (BIOS) memory 157. The system BIOS memory 157 can be any of various non-volatile memory types, such as read-only memory (ROM) or electrically erasable programmable ROM (EEPROM).

[0026] The USB system 160 includes a USB hub 162 connected to a USB master in the data processor 110, and representative USB devices 164, 166, and 168 connected to the USB hub 162, respectively. The USB devices 164, 166, and 168 can be devices such as a keyboard, mouse, and flash EEPROM port.

[0027] The disk drive 170 is connected to the data processor 110 via a SATA bus and provides mass storage for the operating system, application programs, application files, etc.

[0028] The data processing system 100 is suitable for use in modern computing applications by providing memory channels 130 and 140. Each of the memory channels 130 and 140 can be connected to the latest DDR memory technologies, such as DDR version four (DDR4), low power DDR4 (LPDDR4), graphics DDR version five (gDDR5), and high bandwidth memory (HBM), allowing for compatibility with future memory technologies. These memories provide high bus bandwidth and high-speed operation. At the same time, they offer a low-power mode to conserve power for battery-powered applications such as laptop computers, and also provide built-in thermal monitoring.

[0029] Figure 2 is a block diagram of an APU 200 suitable for use in the data processing system 100 of Figure 1. The APU 200 generally includes a central processing unit (CPU) core complex 210, a graphics core 220, a display engine 230, a memory management hub 240, a data fabric 250, a set of peripheral controllers 260, a set of peripheral bus controllers 270, a system management unit (SMU) 280, and a set of memory controllers 290 (memory controllers 292 and 294).

[0030] The CPU core complex 210 includes CPU cores 212 and 214. In this embodiment, the CPU core complex 210 includes two CPU cores, but in other embodiments, the CPU core complex 210 may include any number of CPU cores. Each of the CPU cores 212 and 214 is bidirectionally connected to the system management network (SMN) forming the control fabric and to the data fabric 250, and can provide memory access requests to the data fabric 250. Each of the CPU cores 212 and 214 may be a single core, or further, they may be a core complex consisting of two or more single cores sharing specific resources such as a cache.

[0031] The graphics core 220 is a high-performance graphics processing unit (GPU) capable of performing graphics processing such as vertex processing, fragment processing, shading, and texture blending in a highly integrated parallel manner. The graphics core 220 is bidirectionally connected to the SMN and the data fabric 250 and can provide memory access requests to the data fabric 250. In this regard, the APU200 can support either an integrated memory architecture in which the CPU core complex 210 and the graphics core 220 share the same memory space, or a memory architecture in which the CPU core complex 210 and the graphics core 220 share a portion of the memory space, while the graphics core 220 also uses private graphics memory that is not accessible by the CPU core complex 210.

[0032] The display engine 230 renders and rasterizes objects generated by the graphics core 220 for display on the monitor. The graphics core 220 and the display engine 230 are bidirectionally connected to a common memory management hub 240 for uniform translation to appropriate addresses in the memory system 120, and the memory management hub 240 is bidirectionally connected to a data fabric 250 to generate such memory accesses and receive read data returned from the memory system.

[0033] The data fabric 250 includes a crossbar switch for routing memory access requests and memory responses between any memory access agent and the memory controller 290 (memory controllers 292 and 294). The data fabric also includes a system memory map defined by the BIOS for determining the destination of memory accesses based on the system configuration, as well as buffers for each virtual connection.

[0034] The peripheral controller 260 includes a USB controller 262 and a SATA interface controller 264, each of which is bidirectionally connected to the system hub 266 and the SMN bus. These two controllers are merely typical examples of peripheral controllers that may be used in the APU200.

[0035] The peripheral bus controller 270 includes a system controller or "Southbridge" (SB) 272 and a PCIe controller 274, each of which is bidirectionally connected to an input / output (I / O) hub 276 and the SMN bus. The I / O hub 276 is also bidirectionally connected to a system hub 266 and a data fabric 250. Thus, for example, a CPU core can program registers in the USB controller 262, SATA interface controller 264, SB 272, or PCIe controller 274 via access routed by the data fabric 250 through the I / O hub 276.

[0036] The SMU280 is a local controller that controls the operation of resources on the APU200 and synchronizes communication between them. The SMU280 manages the power-up sequencing of various processors on the APU200 and controls multiple off-chip devices via reset, enable, and other signals. The SMU280 includes one or more clock sources (not shown in Figure 2), such as a phase-locked loop (PLL), to provide clock signals to each of the components of the APU200. The SMU280 can also manage power for various processors and other functional blocks and receive power consumption values ​​measured from the CPU cores 212, 214 and the graphics core 220 to determine appropriate power states.

[0037] Furthermore, the APU200 also implements various system monitoring and power saving functions. In particular, one system monitoring function is thermal monitoring. For example, if the APU200 becomes hot, the SMU280 can reduce the frequency and voltage of CPU cores 212, 214 and / or graphics core 220. If the APU200 becomes too hot, the SMU can be completely shut down. The SMU280 can also receive thermal events from external sensors via the SMN bus, and the SMU280 can reduce the clock frequency and / or power supply voltage accordingly.

[0038] Figure 3 is a block diagram of a memory controller 300 and associated physical interface (PHY) 330 suitable for use with the APU 200 of Figure 2, according to several embodiments. The memory controller 300 includes a memory channel 310 and a power engine 320. The memory channel 310 includes a host interface 312, a memory channel controller 314, and a physical interface 316. The host interface 312 connects the memory channel controller 314 to the data fabric 250 bidirectionally via a scalable data port (SDP). The physical interface 316 connects the memory channel controller 314 to the PHY 330 bidirectionally via a bus and conforms to the DDR-PHY Interface Specification (DFI). The power engine 320 is bidirectionally connected to the SMU 280 via an SMN bus, to the PHY 330 via an Advanced Peripheral Bus (APB), and to the memory channel controller 314. The PHY330 has bidirectional connections to memory channels such as memory channel 130 or memory channel 140 in Figure 1. The memory controller 300 is an instantiation of a memory controller for a single memory channel using a single memory channel controller 314 and has a power engine 320 for controlling the operation of the memory channel controller 314 in a manner further described below.

[0039] Figure 4 is a block diagram of another memory controller 400 and associated PHYs 440 and 450 suitable for use with the APU 200 of Figure 2, according to several embodiments. The memory controller 400 includes memory channels 410 and 420 and a power engine 430. The memory channel 410 includes a host interface 412, a memory channel controller 414, and a physical interface 416. The host interface 412 connects the memory channel controller 414 to the data fabric 250 bidirectionally via an SDP. The physical interface 416 connects the memory channel controller 414 to the PHY 440 bidirectionally and conforms to the DFI specification. The memory channel 420 includes a host interface 422, a memory channel controller 424, and a physical interface 426. The host interface 422 connects the memory channel controller 424 to the data fabric 250 bidirectionally via another SDP. The physical interface 426 connects the memory channel controller 424 to the PHY 450 bidirectionally and conforms to the DFI specification. The power engine 430 is bidirectionally connected to the SMU 280 via the SMN bus, to the PHYs 440 and 450 via the APB, and to the memory channel controllers 414 and 424. The PHY 440 has bidirectional connections to memory channels such as memory channel 130 in Figure 1. The PHY 450 has bidirectional connections to memory channels such as memory channel 140 in Figure 1. The memory controller 400 is an instantiation of a memory controller having two memory channel controllers, and uses the power engine 430 to control the operation of both the memory channel controller 414 and the memory channel controller 424 in a manner further described below.

[0040] Figure 5 shows a non-limiting example of a block diagram of a memory controller 500 according to several embodiments. The memory controller 500 generally includes a memory channel controller 510 and a power controller 550. The memory channel controller 510 generally includes an interface 512, a queue 514, a first command subqueue 520, an address generator 522, content addressable memory (CAM) 524, 529 for each of the respective command subqueue / arbiter pairs, a replay queue 530, a refresh logic block 532, a timing block 534, a page table 536, a corresponding arbiter 538, an error correction code (ECC) check block 542, an ECC generation block 544, and a data buffer (DB) 546. In some embodiments, the memory controller 500 further includes a second command subqueue 521, a corresponding arbiter 539, and a selector 541, as will be further described below. Each command subqueue 525, 527 includes a corresponding arbiter, and is referred herein to as a subqueue / arbiter pair, command subqueue / arbiter group, or command subqueue / arbiter module. In some embodiments, the command subqueue / arbiter modules 525, 527 are duplicated to allow depth expansion as required by a given architecture, and a second level of arbitration via selector 541 is configured to handle additional outputs from the expanded number of modules. In some embodiments, the command queue entry logic 523 places memory access requests into different command subqueues (e.g., sorting and / or transferring between subqueues). In some embodiments, this includes storing only read requests in one command subqueue and only write requests in another command subqueue. In other embodiments, the command queue entry logic 523 combines both read and write requests into the same command subqueue.In certain embodiments, the command queue entry logic 523 moves an entry from one command subqueue to another.

[0041] In some embodiments, each command subqueue 520, 521 stores 32 entries, and therefore the total number of entries from both command subqueues is 64. However, any appropriate number of command subqueues and command subqueue entry sizes may be employed. In some embodiments, functions within various blocks can be combined with other blocks as needed. In some embodiments, the command queue entry logic 523 may be incorporated as part of an address generator or combined with other blocks. In one embodiment, the command queue entry logic is implemented as one or more state machines. However, any appropriate logic may be used. In some embodiments, the command queue entry logic 523 provides subqueue control such that in-order memory access requests are used and one subqueue pushes entries to another subqueue. For example, the command queue entry logic 523 includes a feedback structure from one subqueue to another indicating that an entry may be transferred, and a structure for placing the entry into the other command subqueue.

[0042] Interface 512 has a first bidirectional connection to the data fabric 250 via an external bus and has an output. In the memory controller 500, this external bus conforms to the highly extensible interface version 4 specified by ARM Holdings, PLC of Cambridge, UK, known as "AXI4," although in other embodiments it may be a different type of interface. Interface 512 translates memory access requests from a first clock domain known as the FCLK (or MEMCLK) domain to a second clock domain inside the memory controller 500 known as the UCLK domain. Similarly, queue 514 grants memory access from the UCLK domain to the DFICLK domain associated with the DFI interface.

[0043] The address generator 522 decodes the addresses of memory access requests received from the data fabric 250 via the AXI4 bus. Each memory access request includes an access address in the physical address space, represented in a normalized format. The address generator 522 translates the normalized addresses into a format that can be used to address actual memory devices in the memory system 120 and to efficiently schedule associated accesses. This format includes a region identifier that associates the memory access request with a specific rank, row address, column address, bank address, and bank group. At startup, the system BIOS queries the memory devices in the memory system 120 to determine their size and configuration and programs a set of configuration registers associated with the address generator 522. The address generator 522 uses the configurations stored in the configuration registers to translate the normalized addresses into the appropriate format. Command subqueues 520 and 521, respectively, are queues of memory access requests received from memory access agents in the data processing system 100, such as CPU cores 212, 214, and graphics core 220, provided by the command queue entry logic 523. Command subqueue 520 stores the address field decoded by address generator 522, as well as other address information that enables arbiter 538 to efficiently select memory accesses, including access type and quality of service (QoS) identifiers. Similarly, command subqueue 521 stores the address field decoded by address generator 522, as well as other address information that enables the corresponding arbiter 539 to efficiently select memory accesses, including access type and quality of service (QoS) identifiers. CAM 524 and CAM 529 each contain information for implementing ordering rules, such as write-after-write (WAW) and read-after-write (RAW) ordering rules.

[0044] The replay queue 530 is a temporary queue for storing memory accesses selected by arbiters 538 and 539, awaiting responses such as address and command parity responses, write cyclic redundancy check (CRC) responses for DDR4 DRAM, or write and read CRC responses for GDDR5 DRAM. In some embodiments, for each command queue / arbiter pair, the replay queue 530 accesses the ECC check block 542 to determine whether the returned ECC is correct or indicates an error. The replay queue 530 allows the access to be replayed in the event of a parity or CRC error in one of these cycles. In other embodiments, the replay mechanism is instantiated for each command subqueue / arbiter pair.

[0045] The refresh logic 532 includes a state machine for various power-down, refresh, and termination resistance (ZQ) calibration cycles, which are generated separately from normal read and write memory access requests received from the memory access agent. For example, when a memory rank is in pre-charge power-down mode, the refresh control logic must be invoked periodically to perform a refresh cycle. The refresh logic 532 periodically generates refresh commands to prevent data errors caused by charge leakage from the storage capacitors of the memory cells within the DRAM chip. Furthermore, the refresh logic 532 periodically calibrates the ZQ to prevent mismatches in on-die termination resistances due to thermal changes in the system. The refresh logic 532 also determines when to put the DRAM device into different power-down modes.

[0046] Arbiter 538 is bidirectionally connected to command subqueue 520, and arbiter 539 is bidirectionally connected to command subqueue 521. Each arbiter improves efficiency by intelligently scheduling access from smaller command queues compared to conventional systems in order to improve memory bus utilization. Arbiters 538 and 539 each implement appropriate timing relationships by using timing block 534 to determine, based on DRAM timing parameters, whether a particular access in command subqueue 520 and / or command subqueue 521 is eligible to issue. For example, each DRAM is "t RC It has a minimum specified time between activation commands to the same bank, known as "t". Timing block 534 maintains a set of counters that determine eligibility based on this timing parameter and other timing parameters defined in the JEDEC specification and is bidirectionally connected to replay queue 530. For example, each DRAM has "t RCD It has a minimum specified time between the activation command (or row command) and the column command, known as "". Arbiters 538 and 539 use counters in timing block 534 to determine the eligibility of each CMD. Page table 536 maintains state information about active pages in each bank and rank of the memory channel for arbiters 538 and 539 and is bidirectionally connected to replay queue 530.

[0047] The ECC generation block 544 calculates the ECC according to the write data in response to a write memory access request received from interface 512. DB 546 stores the write data and ECC related to the received memory access request. Selector 541 outputs the combined write data / ECC to queue 514 when it selects the corresponding write access from either command subqueue 520 or command subqueue 521 based on the command selected by each arbiter 538 or arbiter 539 for dispatch to the memory channel, as will be further described below.

[0048] The power controller 550 generally includes an interface 552 to an Advanced Extensible Interface, Version 1 (AXI), an APB interface 554, and a power engine 560. Interface 552 has a first bidirectional connection to the SMN, including an input and an output for receiving an event signal labeled “EVENT_n”, shown separately in Figure 5. The APB interface 554 has an input connected to the output of interface 552 and an output for connecting to the PHY via APB. The power engine 560 has an input connected to the output of interface 552 and an output connected to the input of queue 514. The power engine 560 includes a set of configuration registers 562, a microcontroller (μC) 564, a self refresh controller (SLFREF / PE) 566, and a reliable read / write training engine (RRW / TE) 568. The configuration register 562 is programmed via the AXI bus and stores configuration information for controlling the operation of various blocks within the memory controller 500. Therefore, the configuration register 562 has outputs connected to these blocks, which are not shown in detail in Figure 5. The self-refresh controller 566 is an engine that enables manual generation of refreshes in addition to automatic generation of refreshes by the refresh logic 532. The reliable read / write training engine 568 provides a continuous memory access stream to memory or I / O devices for purposes such as DDR interface read latency training and loopback testing.

[0049] The memory channel controller 510 includes circuitry that enables the selection of memory access for dispatch to the associated memory channel. To make the desired arbitration decision, the address generator 522 decodes the address information into pre-decoded information including rank, row address, column address, bank address, and bank group in the memory system, and the command subqueues 520 and 521 store the pre-decoded information. The configuration register 562 stores configuration information for determining how the address generator 522 decodes the received address information. For entries in the command subqueue 520, the arbiter 538 uses the decoded address information, timing eligibility information indicated by the timing block 534, and active page information indicated by the page table 536 to efficiently provide a "winning" memory access to the selector 541 while complying with other criteria such as QoS requirements. For example, arbiter 538 implements priority access to open pages to avoid the overhead of precharge and activation commands required to change memory pages, hiding overhead access to one bank by interleaving read and write access to another bank. In particular, during normal operation, arbiter 538 may decide to keep pages open in different banks until they need to be precharged before these pages select different pages. Arbiter 539 operates similarly to arbiter 538 and provides a “winning” memory access to selector 541. Selector 541 selects a preferred memory access request 543 from among the memory access requests provided from the first command subqueue 520 and the memory access requests provided from command subqueue 521. Selector 541 sends the thus selected preferred memory access request to one of a plurality of memory channels according to the subchannel. Selector 541 includes a multiplexer circuit in one embodiment. However, any suitable logic may be used.

[0050] The address generator 522 sends the decrypted memory access request, including the decrypted subchannel number, to the command subqueue 520 and the subcommand queue 521. The command subqueue 520 stores the decrypted memory access request in an entry within the command subqueue 520 having a first field for storing the decrypted subchannel number and a second field for storing the remainder of the decrypted memory access request as described above. Similarly, the command subqueue 521 stores the decrypted memory access request in an entry within the command subqueue 521 having a first field for storing the decrypted subchannel number and a second field for storing the remainder of the decrypted memory access request.

[0051] Arbiter 538 is bidirectionally connected to command subqueue 520 and implements appropriate timing relationships by using timing block 534 to determine whether a particular access in command subqueue 520 is eligible for issuance based on DRAM timing parameters. Arbiter 538 selects eligible memory access requests from command subqueue 520 according to predetermined criteria it uses. Similarly, arbiter 539 is bidirectionally connected to command subqueue 521 and implements appropriate timing relationships by using timing block 534 to determine whether a particular access in command subqueue 521 is eligible for issuance based on DRAM timing parameters. Arbiter 538 selects eligible memory access requests from command subqueue 521 according to predetermined criteria it uses. Examples of these predetermined criteria are described above and may be modified between embodiments.

[0052] Figures 6–9 show various non-exclusive examples of command subqueues and corresponding arbiter configurations. However, these are only a few examples, and it will be recognized that other configurations are conceivable. Although not always shown, the arbiters described herein (including the arbiter in Figure 5) acquire priority information, PGT information, timing information, and other information required to comply with JEDEC specifications or any other appropriate standards.

[0053] Figure 6 shows an example of a memory controller having a memory channel controller, which includes a group of command subqueues and corresponding arbiters. The group of command subqueues and corresponding arbiters are shown as command subqueue / arbiter modules 525 and 527. The corresponding arbiters select memory access commands from their respective command subqueues according to predetermined criteria, such as the criteria described above. The command queue entry logic 523 communicates with the command subqueues and places memory access requests in command subqueues 520 and 527. In one embodiment, the command queue entry logic 523 is implemented as one or more state machines. In some examples, the command queue entry logic includes programmable control registers programmed by the CPU or GPU to configure the command queue entry logic 523 to sort memory access requests, such as read and write requests, among the command subqueues in a specific manner. In one embodiment, the command queue entry logic 523 sorts read requests into one command subqueue and write requests into different command subqueues. In other embodiments, reads and writes are mixed within the command subqueues. In some embodiments, entries are transferred between command subqueues. In some embodiments, programmable control registers are used to set thresholds for various levels of arbitration, as further described below.

[0054] Each arbiter 538, 539, in one embodiment, arbitrates based on JEDEC specification criteria and evaluates page table (PGT) information 600, timing information 602, and priority information 604 (e.g., low, medium, high, urgent) to select winning memory access requests 606, 608 to be provided to selector 541. For example, PGT information 600 indicates whether a DRAM page is open (i.e., activated) or closed (i.e., precharged), and therefore no pages on the bank are open. Timing information 602, in one embodiment, is as described above. rc , t rcd This refers to the following. Priority information 604 indicates the priority level of the request, whether it is low, medium, high, or urgent.

[0055] A page hit (PH) means that the required page is already "active" in the DRAM device's sense amplifier and that line is directly readable or writable. This is the lowest latency scenario. A page miss (PM) means that the required page is not open in the DRAM sense amplifier and therefore needs to be "activated" and then a page hit for access. rcd Waiting. This is a medium latency response. Page contention (PC) means that because the desired page is not the currently open page in its DRAM bank, an existing page must be precharged (Trp wait), then the new page must be activated (Trp wait), then a page hit occurs and it can be accessed. This has the highest latency overhead.

[0056] The first level of arbitration is performed by arbiters 538 and 539, respectively, to select a winning entry and pass it to the second level of arbitration logic 612. For example, arbiters 538 and 539 select memory access commands in each subqueue that meet specified criteria, such as the command with the highest priority, based on priority information 604 from all entries in each command subqueue. Timing dependencies are resolved for the command as indicated by timing information 602 (timing OK information, etc.), detected for the command, and there must also be page table hits as indicated by page table information 600. If no command meets this criterion, a winning command is selected using other criteria, such as the oldest command in the subqueue or any other suitable criterion.

[0057] Selector 541 includes a multiplexer 610 and a second-level arbitration logic 612. The multiplexer 610 selects one of the winning memory access requests from one of the two command subqueues based on criteria such as timing thresholds and other criteria determined by the second-level arbitration logic 612. In one embodiment, the second-level arbitration logic 612 is implemented as one or more state machines, but any suitable logic including a programmed processor or any other suitable logic can be used. Thus, in this embodiment, the second-level arbitration logic 612 selects one winner from the winners of each instance of the command subqueue and the corresponding arbiter group based on generated or pre-stored criteria such as timing information, page hit information and other information. For example, page hits take precedence over page misses. In some embodiments, if one of the winning memory access requests 606 or 608 is a page hit, for example, if the other memory access request is for a memory access request that requires activating and precharging a row of memory (e.g., a page miss), then that memory access request is selected as the preferred memory access request 543. In some embodiments, if both memory access requests have page hits, the second-level arbitration logic 612 selects the memory access request with the highest priority. In some embodiments, if both memory access requests 606 and 608 have page hits and both have the same priority level, the second-level arbitration logic 612 selects the older memory access request as the preferred memory access request. These are examples, and it will be recognized that any appropriate selection criteria may be used.

[0058] In some embodiments, entries are transferred between command subqueues as indicated by dashed arrows such as arrow 614. In some embodiments, command subqueues are operated sequentially such that the oldest entries are at the end of each subqueue. The command queue entry logic 523 in this embodiment provides feedback from one command subqueue to the other to inform the other that an entry may be transferred to a command subqueue that has an open entry. In other embodiments, non-sequential operation is provided.

[0059] Figure 7 shows another exemplary embodiment 701 in which multiple command subqueues / arbiter modules 525, 527 are configured as write subqueues and additional command subqueues 700, 702 (e.g., command subqueues / arbiter modules) are configured as read subqueues. For example, embodiment 701 uses dedicated write command subqueues and dedicated read command subqueues. In this example, the second-level arbitration logic 612 uses criteria such as timing criteria to determine which write access requests are selected and which read requests are selected to be output to a third selector 704, which acts as selector 541 in this embodiment. Thus, in this embodiment, there are four command subqueues and four corresponding arbiters, two second-level arbiter logic sections, selectors 706 and 708, and a third selector 704. Selector 704 selects memory requests from either the first selector 706 or selector 708 and sends the selected memory access requests as preferred selected memory access requests to one or more subchannels. In this embodiment, the selectors are 706 and 708.

[0060] In this embodiment, the selector 704 uses write and read thresholds stored in the control register to select whether a read or write request is selected as a preferred memory access request via the multiplexer 710. In some embodiments, read requests take precedence over write requests. The write and read thresholds are set to avoid collisions on the data bus when switching the bus between read and write operations. For example, the read threshold is the number of consecutive reads that should be performed before the bus is switched to a write operation, and the write threshold is the number of consecutive writes that may occur before the bus is switched to perform a read operation. Also in this embodiment, the command queue entry logic 523 sorts the memory access requests received from the address generator 522 into different command subqueues 525, 527, 700, and 702 so that command subqueue 525 contains only write requests and command subqueue 700 receives only read requests. Similarly, the command queue entry logic 523 sorts the memory access requests so that only write requests are stored in command subqueue 527 and only read requests are stored in command subqueue 702. In this embodiment, arbiters 538 and 539 are configured to be of the same type, evaluating timing information related to write requests and other necessary information as described above, while arbiters 714 and 716 are configured to arbitrate for read requests and take into account the timing information necessary to properly process memory access read requests.

[0061] Figure 8 shows another exemplary memory channel controller configuration in which read and write queues are employed for each subchannel, such that each command subqueue 800, 802 stores both read and write memory access requests instead of only read or write memory access requests, and each of the command subqueues 804, 806 also stores both read and write memory access requests.

[0062] Figure 9 shows another exemplary embodiment that includes multiple instances of the structure shown in Figure 7, dedicated to each subchannel of memory. Thus, in this example, there are dedicated read command subqueues and dedicated write command subqueues for each subchannel. For example, there are multiple command write subqueues 900, 902 for subchannel 0 and multiple command subqueues 904, 906 for subchannel 1. Similarly, there are multiple command read subqueues 908, 910 dedicated to subchannel 0 and command write subqueues 912, 914 dedicated to subchannel 1. Each command subqueue has a corresponding dedicated arbiter, as shown.

[0063] Figure 10 shows an exemplary method 1000 for controlling a memory system having multiple memory channels. In certain embodiments, the method is performed by the structures shown in Figures 5 to 9. In some embodiments, the method includes receiving memory access requests by an address generator 522, etc. 1002 and decoding each of the memory access requests by the address generator 522, etc. In embodiments employing DDRx memory, the method includes decoding the address to a bank, rank, and one of the subchannels of the memory device in the memory system, as shown in block 1004. The method includes storing the decoded address in one of the command subqueues. In embodiments using DDRx memory, the method includes storing the memory access request, including the bank, rank, and subchannel, in at least a first command subqueue or at least a second command subqueue by a command queue entry logic 523, etc. 1006. When memory that does not use bank and rank designation is used, operations 1004 and 1006 do not need to be performed. In some embodiments, storage includes sorting memory access requests into different command subqueues such that a first command subqueue contains only read requests and a second command subqueue contains only write requests, but other embodiments use other sorting criteria. In some embodiments, storage includes transferring entries between command subqueues. This method includes arbiter 539, etc., selecting from a plurality of memory access requests in the first command subqueue using predetermined criteria 1008 in order to provide a first memory access request selected from the first command subqueue, and arbiter 538, etc., selecting from a plurality of memory access requests in the second command subqueue using predetermined criteria 1010 in order to provide a second memory access request selected from the second command subqueue.This method includes selecting a preferred memory access request from a first memory access request provided from a first command subqueue and a second memory access request from a second command subqueue using a selector 541, etc. 1012, and sending the thus selected preferred memory access request to one of a plurality of memory channels according to the subchannel 1014. It will be recognized that the methods provided herein are merely examples, the operations may be combined, their order may be changed, and other variations may be made depending on the desired operation.

[0064] One or more embodiments utilize dedicated command subqueues and corresponding arbiters that are smaller than conventional command queues. For example, in some embodiments of this specification, instead of a larger 64-entry command queue, multiple 32-entry command queues and corresponding arbiters dedicated to serving each of the smaller command queues are used. Arbitration and timing determination are performed much faster (e.g., higher clock frequencies are used), improving the operating speed of the memory controller and improving the data throughput of the integrated circuit including the memory controller.

[0065] The memory controller 500 in Figure 5 may be implemented in various combinations of hardware and software. This hardware circuit may include a priority encoder, a finite state machine, a programmable logic array (PLA), etc. In some embodiments, other functional blocks can be executed by a data processor under software control. Some of the software components may be stored in a computer-readable storage medium for execution by at least one processor and may correspond to instructions stored in non-temporary computer memory or a computer-readable storage medium. In various embodiments, the non-temporary computer-readable storage medium may include a magnetic or optical disk storage device, a solid-state storage device such as flash memory, or other non-volatile memory devices. Executable instructions stored in the non-temporary computer-readable storage medium may be source code, assembly language code, object code, or other instruction forms that can be interpreted and / or otherwise executed by one or more processors.

[0066] The memory controller 500 in Figure 5, or any part thereof, may be described or represented by a computer-accessible data structure in the form of a database or other data structure that can be read by a program and used directly or indirectly to manufacture an integrated circuit. For example, this data structure may be an operational-level description or register-transfer-level (RTL) description of hardware functions in a high-level design language (HDL) such as Verilog or VHDL. The description can be read by a synthesis tool that can synthesize the description to generate a netlist containing a list of gates from a synthesis library. The netlist contains a set of gates that also represent the functions of hardware, including the integrated circuit. The netlist can then be arranged and routed to generate a dataset that describes the geometric shapes applied to a mask. The mask can then be used in various semiconductor manufacturing processes to manufacture the integrated circuit. Alternatively, the database on a computer-accessible storage medium may, as desired, be a netlist (with or without a synthesis library), a dataset, or Graphic Data System (GDS) II data.

[0067] While specific embodiments have been described, various modifications to these embodiments will be apparent to those skilled in the art. For example, the memory controller 500 may interface with other types of memory besides DDRx, such as high-bandwidth memory (HBM), RAMbus DRAM (RDRAM), and others, as well as different types of DIMMs. The memory controller can be integrated into network controllers, hard drive controllers, and other devices. The illustrated embodiments illustrate memory addressing and control signals useful in DDR memory, but these vary depending on the type of memory used. Furthermore, the memory access control in Figure 6 can be extended to more than two virtual channels.

[0068] Therefore, the attached claims are intended to cover all modifications of the disclosed embodiments that fall within the scope of the disclosed embodiments.

[0069] While features and elements are described above in specific combinations, each feature or element can be used alone without other features and elements, or in various combinations with or without other features and elements. In some embodiments, the devices described herein may be implemented in computer programs, software, or firmware embedded in a non-temporary computer-readable storage medium for implementation by a general-purpose computer or processor. Examples of computer-readable storage media include read-only memory (ROM), random-access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (e.g., internal hard disks and removable disks), magneto-optical media, and optical media (e.g., CD-ROM disks and digital versatile disks (DVDs)).

[0070] In the above detailed descriptions of various embodiments, references have been made to the accompanying drawings illustrating specific preferred embodiments that form part thereof and can carry out the invention. These embodiments are described in sufficient detail to enable those skilled in the art to carry out the invention, and it should be understood that other embodiments may be utilized and logical, mechanical, and electrical modifications may be made without departing from the scope of the invention. In order to avoid details that are not necessary to enable those skilled in the art to carry out the invention, the descriptions may omit certain information known to those skilled in the art. Furthermore, many other various embodiments incorporating the teachings of this disclosure can be readily constructed by those skilled in the art. Accordingly, the invention is not intended to be limited to the specific forms described herein, but rather to encompass such alternative forms, modifications, and equivalents that can reasonably be included within the scope of the invention. Accordingly, the above detailed descriptions should not be construed as restrictive, and the scope of the invention is defined solely by the appended claims. The above detailed descriptions of embodiments and examples described herein are presented for illustrative and explanatory purposes only, and not as limitations. For example, the operations described may be performed in any suitable order or method. Accordingly, the present invention is intended to encompass any modifications, variations, or equivalents that fall within the scope of the fundamental principles disclosed above and claimed herein.

[0071] The above detailed description and the examples described herein are provided for illustrative and explanatory purposes only, and not for limitation.

Claims

1. A method for controlling a memory system having multiple memory channels or subchannels, The memory access request is placed in at least a first and second read command subqueue of the read command queue, and at least a first and second write command subqueue of the write command queue. The process involves mediating read commands between the first read command subqueue and the second read command subqueue, and mediating commands between the first write command subqueue and the second write command subqueue, This includes sending a memory access request selected from among the arbitrated commands to at least one of the memory channels or subchannels, method.

2. This includes sorting memory access requests into different command subqueues such that the first read command subqueue and the second read command subqueue contain only read requests, and the first write command subqueue and the second write command subqueue contain only write requests. The method according to claim 1.

3. Decoding each of the aforementioned memory access requests to the banks, ranks, and subchannels of multiple subchannels of a memory device within the memory system, The memory access request, including the bank, the rank, and the subchannel, is stored in at least one of the first read command subqueue and the second read command subqueue, or the first write command subqueue and the second write command subqueue. The method according to claim 1.

4. The process includes transferring memory access requests between the first read command subqueue and the second read command subqueue, and selecting the memory access requests in each of the first and second read command subqueues includes using predetermined criteria. The method according to claim 1.

5. This includes selecting a preferred memory access request by selecting the oldest memory access request from among the oldest memory request in the first read command subqueue and the oldest memory request in the second read command subqueue. The method according to claim 1.

6. It is a memory controller, A command queue entry logic configured to sort memory access requests into different command subqueues, wherein the first and second read command subqueues of a read command queue contain read requests, and the first and second write command subqueues of a write command queue contain write requests; A first arbiter, which selects memory access commands, is coupled to the first read command subqueue and the second read command subqueue, A second arbiter, which is coupled to the first write command subqueue and the second write command subqueue, selects memory access commands, A first selector capable of selecting a memory request from either the first arbiter or the second arbiter and transmitting the selected memory access request to a memory channel or subchannel, Memory controller.

7. The command queue entry logic is configured to sort memory access requests into different command subqueues such that the first read command subqueue and the second read command subqueue contain only read requests, and the first write command subqueue and the second write command subqueue contain only write requests. The memory controller according to claim 6.

8. The command queue entry logic is configured to transfer entries from a first command subqueue to a second command subqueue, and the command queue entry logic is operably coupled to the first read command subqueue, the second read command subqueue, the first write command subqueue, and the second write command subqueue, and the first selector is operably coupled to both the first arbiter and the second arbiter. The memory controller according to claim 6.

9. The first arbiter and the second arbiter each determine the eligibility of a memory access command, comprising: a shared timing logic shared between the first arbiter and the second arbiter; and a shared page table shared between the first arbiter and the second arbiter, which stores information about active pages used to select a memory request provided to the first selector from the first arbiter or the second arbiter. The memory controller according to claim 6.

10. The command queue entry logic described above is: Sorting received memory access requests, decoded into banks, ranks, and subchannels of multiple subchannels of a memory device within a memory system, A memory access request including the bank, the rank, and the subchannel is stored in at least one of the first read command subqueue and the second read command subqueue, or the first write command subqueue and the second write command subqueue. It is configured to do, The memory controller according to claim 6.

11. A data processing system, Multiple memory access agents for providing memory access requests, Multiple memory channels, A memory controller comprising any one of claims 6 to 10, Data processing system.

12. The read command queue is configured to store memory access requests for multiple ranks of a memory channel, and the write command queue is configured to store memory access requests for multiple ranks of the same memory channel. The method according to claim 1.

13. The read command queue is configured to store memory access requests for multiple ranks of a memory channel, and the write command queue is configured to store memory access requests for multiple ranks of the same memory channel. The memory controller according to claim 6.