Writing bank group masks during arbitration

By introducing a command queue and an arbitrator into the memory controller, combining the memory bank group tracking circuit, optimizing the memory access sequence and managing the memory bank group, the problems of low memory access efficiency and high cost in the prior art are solved, and efficient and low-cost memory access is achieved.

CN117083588BActive Publication Date: 2025-05-06ADVANCED MICRO DEVICES INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202280025683.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-03-31
Filing Date
2022-03-17
Publication Date
2025-05-06
Estimated Expiration
2042-03-17

AI Technical Summary

Technical Problem

When optimizing memory access efficiency, existing DDRDRAM memory controllers have difficulty in dealing with complex memory characteristics without increasing size and cost, and it is difficult to effectively manage the delays caused by bank groups and on-chip ECC computing.

Method used

A memory controller is designed that includes a command queue and an arbitrator that tracks the bank group number of previously written requests through a bank group tracking circuit and allows these requests to be issued after a specified period, thereby optimizing the memory access sequence and reducing the write delay of the bank group.

Benefits of technology

By optimizing the memory access sequence and managing the memory bank group, the memory access efficiency is significantly improved, the memory controller size and cost are reduced, and the delay caused by on-chip ECC calculations is effectively managed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117083588B_ABST
    Figure CN117083588B_ABST
Patent Text Reader

Abstract

A memory controller includes an arbiter for selecting memory requests from a command queue for transmission to a DRAM memory. The arbiter includes a bank group tracking circuit that tracks bank group numbers of three or more previous write requests selected by the arbiter. The arbiter also includes a selection circuit that selects a request to be issued from the command queue and prevents selection of a write request and associated activate command for the tracked bank group number unless no other write requests in the command queue are eligible. The bank group tracking circuit indicates that a previous write request and associated activate command are eligible to be issued after a number of clock cycles of a minimum write-to-write timing period of the bank group corresponding to the previous write request have passed.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Computer systems typically use inexpensive and high-density dynamic random access memory (DRAM) chips as main memory. Most DRAM chips sold today are compatible with various double data rate (DDR) DRAM standards issued by the Joint Electron Device Engineering Council (JEDEC). DDR DRAM uses a conventional DRAM memory cell array with high-speed access circuits to achieve high transfer rates and improve the utilization of the memory bus.

[0002] The memory controller is a digital circuit that manages the flow of data to and from the DRAM over the memory bus. The memory controller receives memory access requests from the host system, stores them in a queue, and dispatches them to the DRAM in an order selected by an arbiter. Over time, JEDEC has specified DDR DRAM with additional features and complexity, making it difficult for DRAM memory controllers to optimize memory access efficiency without incurring excessive size and cost and requiring a complete redesign of the previous memory controller. BRIEF DESCRIPTION OF THE DRAWINGS

[0003] Figure 1 An accelerated processing unit (APU) and a memory system known in the prior art are shown in block diagram form;

[0004] Figure 2 A block diagram is shown in accordance with some embodiments suitable for use in a similar Figure 1 A memory controller used in an APU of an APU;

[0005] Figure 3 A memory bank group structure of a memory according to the prior art is shown in the form of a block diagram;

[0006] Figure 4 A block diagram is shown according to some embodiments. Figure 2 A portion of a memory controller;

[0007] Figure 5 In block diagram form, a Figure 2 A portion of a memory controller;

[0008] Figure 6 A flowchart of a process for processing a write command according to some embodiments is shown; and

[0009] Figure 7 A flow diagram of a process for arbitrating commands according to some embodiments is shown.

[0010] In the following description, the same reference numerals are used in different drawings to indicate similar or identical items. Unless otherwise specified, the word "couple" and its associated verb forms include both direct connection and indirect electrical connection by means known in the art, and unless otherwise specified, any description of direct connection also means an alternative embodiment using an appropriate form of indirect electrical connection. DETAILED DESCRIPTION

[0011] The memory controller includes a command queue and an arbiter. The command queue has an input for receiving a memory access request for a memory channel, and a plurality of entries for storing a predetermined number of memory access requests. The arbiter is used to select a memory request from the command queue for transmission to a DRAM memory connected to the DRAM channel, and includes a memory bank group tracking circuit that tracks the memory bank group numbers of three or more previous write requests selected by the arbiter. The selection circuit selects a request to be issued from the command queue and prevents the selection of a write request and an associated activation command of the tracked memory bank group number. The memory bank group tracking circuit indicates that a previous write request and an associated activation command are eligible to be issued after a specified period has passed.

[0012] The method includes receiving a plurality of memory access requests at a memory controller and placing them in a command queue to await transmission to a DRAM memory including a plurality of memory bank groups. Selecting a write request from the command queue using an arbitrator for transmission to the DRAM memory. Tracking a memory bank group number for three or more previous write requests selected by the arbitrator. Preventing subsequent write requests and associated activation commands from selecting the tracked memory bank group number. Enabling the previous write request and associated activation command to be issued after a specified period has elapsed.

[0013] A data processing system includes a memory channel connected to a DRAM memory and a memory controller connected to the memory channel. The memory controller includes a command queue and an arbiter. The command queue has an input for receiving a memory access request to the memory channel, and a plurality of entries for storing a predetermined number of memory access requests. The arbiter selects a memory request from the command queue for transmission to the DRAM memory, and includes a memory bank group tracking circuit that tracks the memory bank group numbers of three or more previous write requests selected by the arbiter; and a selection circuit that prevents the selection of a write request and an associated activation command of the tracked memory bank group number. The memory bank group tracking circuit indicates that the previous write request and the associated activation command are eligible to be issued after a specified period has passed.

[0014] Figure 1An accelerated processing unit (APU) 100 and a memory system 130, known in the art, are shown in block diagram form. The APU 100 is an integrated circuit suitable for use as a processor in a host data processing system and generally includes a central processing unit (CPU) core complex 110, a graphics core 120, a set of display engines 122, a data fabric 125, a memory management hub 140, a set of peripheral controllers 160, a set of peripheral bus controllers 170, and a system management unit (SMU) 180.

[0015] The CPU core complex 110 includes a CPU core 112 and a CPU core 114. In this example, the CPU core complex 110 includes two CPU cores, but in other embodiments, the CPU core complex 110 may include any number of CPU cores. Each of the CPU cores 112 and 114 is bidirectionally connected to a system management network (SMN) (which forms a control fabric) and a data fabric 125, and can provide memory access requests to the data fabric 125. Each of the CPU cores 112 and 114 can be a unitary core, or can also be a core complex with two or more unitary cores that share certain resources such as a cache.

[0016] The graphics core 120 is a high-performance graphics processing unit (GPU) that can perform graphics operations such as vertex processing, fragment processing, shading, texture blending, etc. in a highly integrated and parallel manner. The graphics core 120 is bidirectionally connected to the SMN and the data texture 125, and can provide memory access requests to the data texture 125. In this regard, the APU 100 may support a unified memory architecture in which the CPU core complex 110 and the graphics core 120 share the same memory space, or a memory architecture in which the CPU core complex 110 and the graphics core 120 share a portion of the memory space while the graphics core 120 also uses a private graphics memory that the CPU core complex 110 cannot access.

[0017] Display engine 122 renders and rasterizes objects generated by graphics core 120 for display on a monitor. Graphics core 120 and display engine 122 are bidirectionally connected to a common memory management hub 140 via data texture 125 for unified translation to appropriate addresses in memory system 130.

[0018] The data fabric 125 includes a crossbar switch for routing memory access requests and memory responses between any memory access agents and the memory management hub 140. The data fabric also includes a system memory map defined by the basic input / output system (BIOS) for determining the destination of memory accesses based on the system configuration, and a buffer for each virtual connection.

[0019] Peripheral controllers 160 include a Universal Serial Bus (USB) controller 162 and a Serial Advanced Technology Attachment (SATA) interface controller 164, each of which is bidirectionally connected to a system hub 166 and the SMN bus. These two controllers are merely examples of peripheral controllers that may be used with APU 100.

[0020] The peripheral bus controller 170 includes a system controller or "South Bridge" (SB) 172 and a peripheral component interconnect express (PCIe) controller 174, each of which is bidirectionally connected to an input / output (I / O) hub 176 and the SMN bus. The I / O hub 176 is also bidirectionally connected to the system hub 166 and the data texture 125. Thus, for example, the CPU core can program registers in the USB controller 162, the SATA interface controller 164, the SB 172, or the PCIe controller 174 through accesses routed through the I / O hub 176 through the data texture 125. The software and firmware of the APU 100 are stored in a system data drive or system BIOS memory (not shown), which can be any of a variety of non-volatile memory types, such as read-only memory (ROM), flash electrically erasable programmable ROM (EEPROM), etc. Typically, the BIOS memory is accessed through the PCIe bus, and the system data drive is through the SATA interface.

[0021] The SMU 180 is a local controller that controls the operation of resources on the APU 100 and synchronizes communications between these resources. The SMU 180 manages the power-up sequencing of the various processors on the APU 100 and controls a number of off-chip devices via reset, enable, and other signals. The SMU 180 includes one or more clock sources (not shown), such as a phase-locked loop (PLL), to provide clock signals to each component of the APU 100. The SMU 180 also manages the power of the various processors and other functional blocks, and may receive measured power consumption values ​​from the CPU cores 112 and 114 and the graphics core 120 to determine the appropriate power state.

[0022] In this embodiment, a memory management hub 140 and its associated physical interfaces (PHYs) 151 and 152 are integrated with the APU 100. The memory management hub 140 includes memory channels 141 and 142 and a power engine 149. The memory channel 141 includes a host interface 145, a memory channel controller 143, and a physical interface 147. The host interface 145 connects the memory channel controller 143 bidirectionally to the data texture 125 via a serial presence detection link (SDP). The physical interface 147 connects the memory channel controller 143 bidirectionally to the PHY 151 and complies with the DDR PHY interface (DFI) specification. The memory channel 142 includes a host interface 146, a memory channel controller 144, and a physical interface 148. The host interface 146 connects the memory channel controller 144 bidirectionally to the data texture 125 via another SDP. The physical interface 148 connects the memory channel controller 144 bidirectionally to the PHY 152 and complies with the DFI specification. Power engine 149 is bidirectionally connected to SMU 180 via the SMN bus, to PHY 151 and PHY 152 via APB, and also bidirectionally connected to memory channel controllers 143 and 144. PHY 151 has a bidirectional connection to memory channel 131. PHY 152 has a bidirectional connection to memory channel 133.

[0023] The memory management hub 140 is an instantiation of a memory controller having two memory channel controllers, and uses a shared power engine 149 to control the operation of both memory channel controller 143 and memory channel controller 144 in a manner that will be further described below. Each of the memory channels 141 and 142 can be connected to prior art DDR memories, such as fifth generation DDR (DDR5), fourth generation DDR (DDR4), low power DDR4 (LPDDR4), fifth generation graphics DDR (GDDR5), and high bandwidth memory (HBM), and can be adapted to future memory technologies. These memories provide high bus bandwidth and high-speed operation. At the same time, they also provide low power modes to save power for battery-powered applications such as laptops, and also provide built-in thermal monitoring.

[0024] Memory system 130 includes memory channel 131 and memory channel 133. Memory channel 131 includes a set of dual inline memory modules (DIMMs) connected to DDRx bus 132, including representative DIMMs 134, 136, and 138, which correspond to separate memory ranks in this example. Similarly, memory channel 133 includes a set of DIMMs connected to DDRx bus 129, including representative DIMMs 135, 137, and 139.

[0025] The APU 100 operates as the central processing unit (CPU) of the host data processing system and provides various buses and interfaces available in modern computer systems. These interfaces include two double data rate (DDRx) memory channels, a PCIe root complex for connecting to a PCIe link, a USB controller for connecting to a USB network, and an interface to a SATA mass storage device.

[0026] APU 100 also implements various system monitoring and power saving functions. Specifically, one system monitoring function is thermal monitoring. For example, if APU 100 becomes hot, SMU 180 may reduce the frequency and voltage of CPU cores 112 and 114 and / or graphics core 120. If APU 100 becomes too hot, it may be shut down completely. SMU 180 may also receive thermal events from external sensors via the SMN bus, and in response, SMU 180 may reduce the clock frequency and / or the power supply voltage.

[0027] Figure 2 The block diagram shows a suitable Figure 1 The memory controller 200 used in the APU of the APU of the present invention. The memory controller 200 generally includes a memory channel controller 210 and a power controller 250. The memory channel controller 210 generally includes an interface 212, a memory interface queue 214, a command queue 220, an address generator 222, a content addressable memory (CAM) 224, a replay control logic 231 including a replay queue 230, a refresh control logic block 232, a timing block 234, a page table 236, an arbiter 238, an error correction code (ECC) check circuit 242, an ECC generation block 244, a data buffer 246, and a refresh control logic 232.

[0028] The interface 212 has a first bidirectional connection to the data fabric through an external bus and has an output. In the memory controller 200, the external bus is compatible with the Advanced eXtensible Interface version four (referred to as AXI4) specified by ARM Holdings, PLC of Cambridge, England, but may be other types of interfaces in other embodiments. The interface 212 converts memory access requests from a first clock domain referred to as the FCLK (or MEMCLK) domain to a second clock domain referred to as the UCLK domain internal to the memory controller 200. Similarly, the memory interface queue 214 provides memory access from the UCLK domain to the DFICLK domain associated with the DFI interface.

[0029] The address generator 222 decodes the address of the memory access request received from the data texture via the AXI4 bus. The memory access request includes an access address represented in a normalized format in the physical address space. The address generator 222 converts the normalized address into a format that can be used to address the actual memory device in the memory system 130 and efficiently schedule the related access. The format includes a region identifier that associates the memory access request with a specific storage column, row address, column address, storage body address, and storage body group. At startup, the system BIOS queries the memory devices in the memory system 130 to determine their size and configuration, and programs a set of configuration registers associated with the address generator 222. The address generator 222 uses the configuration stored in the configuration register to convert the normalized address into the appropriate format. The command queue 220 is a queue of memory access requests received from memory access agents in the APU 100, such as the CPU cores 112 and 114 and the graphics core 120. Command queue 220 stores address fields decoded by address generator 222 and other address information, including access type and quality of service (QoS) identification, that allows arbiter 238 to efficiently select memory accesses. CAM 224 includes information to implement ordering rules such as write-after-write (WAW) and read-after-write (RAW) ordering rules.

[0030] Error correction code (ECC) generation block 244 determines the ECC for write data to be sent to memory. This ECC data is then added to the write data in data buffer 246. ECC check circuit 242 checks the received ECC against the incoming ECC.

[0031] The replay queue 230 is a temporary queue for storing selected memory accesses selected by the arbiter 238, which are awaiting responses, such as address and command parity responses. The replay control logic 231 accesses the ECC check circuit 242 to determine whether the returned ECC is correct or indicates an error. The replay control logic 231 initiates and controls a replay sequence in which the access is replayed if a parity error or ECC error occurs in one of these cycles. The replayed commands are placed in the memory interface queue 214.

[0032] The refresh control logic 232 includes a state machine for various power-down, refresh and terminal resistance (ZQ) calibration cycles, which are generated separately from the normal read and write memory access requests received from the memory access agent. For example, if the memory storage column is in pre-charge power-down, the memory storage column must be periodically awakened to run the refresh cycle. The refresh control logic 232 generates refresh commands periodically and in response to specified conditions to prevent data errors caused by charge leakage from the storage capacitor of the memory cell in the DRAM chip. The refresh control logic 232 includes an activation counter 248, which has a counter for each memory area in this embodiment, and the counter counts the rolling number of activation commands sent to the memory area through the memory channel. The storage area is a memory storage body in some embodiments, and a memory sub-storage body in other embodiments, as discussed further below. In addition, the refresh control logic 232 periodically calibrates the ZQ to prevent mismatches of on-chip terminal resistances caused by thermal changes in the system.

[0033] The arbiter 238 is bidirectionally coupled to the command queue 220 and is the heart of the memory channel controller 210, performing intelligent scheduling of accesses to improve memory bus usage. In this embodiment, the arbiter 238 includes a bank group tracking circuit 235 for tracking the bank group numbers of multiple recently issued write commands and "masking" these bank groups by preventing commands from being dispatched to these bank groups for a specified period of time under certain conditions, as further described below. The arbiter 238 uses a timing block 234 to enforce correct timing relationships by determining whether certain accesses in the command queue 220 are eligible to be issued based on DRAM timing parameters. For example, each DRAM has a minimum specified time between activation commands, called "t RC ". The timing block 234 maintains a set of counters that determine eligibility based on the timing parameters and other timing parameters specified in the JEDEC specification, and the timing block is bidirectionally connected to the replay queue 230. The page table 236 maintains status information about the active pages in each memory bank and memory column of the memory channel of the arbiter 238, and is bidirectionally connected to the replay queue 230.

[0034] In response to receiving a write memory access request from the interface 212, the ECC generation block 244 calculates the ECC based on the write data. The data buffer 246 stores the write data and ECC for the received memory access request. When the arbiter 238 selects the corresponding write access for dispatching to the memory channel, the data buffer outputs the combined write data / ECC to the memory interface queue 214.

[0035] The power controller 250 generally includes an interface 252 to the Advanced eXtensible Interface version 1 (AXI), an Advanced Peripheral Bus (APB) interface 254, and a power engine 260. The interface 252 has a first bidirectional connection to the SMN, which includes a first bidirectional connection for receiving Figure 2 The APB interface 254 has an input connected to the output of the interface 252, and an output for connecting to the PHY via the APB. The power engine 260 has an input connected to the output of the interface 252, and an output connected to the input of the memory interface queue 214. The power engine 260 includes a set of configuration registers 262, a microcontroller (μC) 264, a self-refresh controller (SLFREF / PE) 266, and a reliable read / write timing engine (RRW / TE) 268. The configuration registers 262 are programmed via the AXI bus and store configuration information to control the operation of various blocks in the memory controller 200. Therefore, the configuration registers 262 have outputs connected to these blocks, which are in Figure 2 SLFREF / PE 266 is an engine that allows refreshes to be manually generated in addition to the refreshes automatically generated by refresh control logic 232. Reliable read / write timing engine 268 provides a continuous memory access stream to the memory or I / O device for purposes such as DDR interface maximum read latency (MRL) training and loopback testing.

[0036] The memory channel controller 210 includes circuitry that allows it to select memory accesses for scheduling to associated memory channels. In order to make the desired arbitration decision, the address generator 222 decodes the address information into pre-decoding information, which includes storage columns, row addresses, column addresses, memory bank addresses, and memory bank groups in the memory system, and the command queue 220 stores the pre-decoding information. The configuration register 262 stores configuration information to determine the manner in which the address generator 222 decodes the received address information. The arbitrator 238 uses the decoded address information, the timing qualification information indicated by the timing block 234, and the active page information indicated by the page table 236 to efficiently schedule memory accesses while complying with other criteria such as quality of service (QoS) requirements. For example, the arbitrator 238 implements priority for accessing open pages to avoid the overhead of precharge and activation commands required to change memory pages, and hides the overhead access to one memory bank by interleaving the overhead access to one memory bank with the read and write access to another memory bank. Particularly during normal operation, arbiter 238 typically keeps pages open in different memory banks until the pages need to be precharged and then selects a different page. In some embodiments, arbiter 238 determines eligibility for command selection based at least on the corresponding value of activation counter 248 for the target memory region of the corresponding command.

[0037] Figure 3 A bank group structure 300 of a DDR5 memory according to the prior art is shown in block diagram form. In general, the DDR5 standard doubles the maximum number of bank groups relative to DDR4 memory while providing more total banks, thereby improving overall system efficiency by allowing more pages to be open at any given time. The DDR5 bank group scheme allows DDR5 DRAM to have a relatively small size by sharing common circuits between memory arrays of four banks instead of two banks. However, the DDR5 standard also includes on-chip ECC calculations, which causes the data write circuits of the bank groups to take longer cycles per write than existing standards like DDR4.

[0038] The depicted bank group structure 300 shows eight bank groups in a DDR5 memory, with the bank groups labeled "BGA" through "BGH." (The DDR5 standard allows for different bank groupings of the 32 banks of a DDR5 memory device). The bank group block 310 shows the structure of the bank group BGA. The bank group block includes four banks "A0," "A1," "A2," and "A3" and their shared input / output circuits. The shared circuits include a column decoder ("COLDEC"), a write driver, and an I / O sense amplifier 312, as well as a serializer / deserializer and ECC encoder / decoder block 314.

[0039] In operation, the ECC encoder / decoder causes delays between write commands to any bank in a bank group because the ECC must be calculated and stored when writing data. In DDR5, this ECC is stored in 8 additional memory bits for every 128 bits of data stored to the DRAM. For each write command, the ECC encoder calculates the ECC for the write data and then stores it in the additional memory bits. The time required for the ECC encoder to function increases the "minimum write-to-write" time within the same bank group to a much longer time than that required by previous standards that did not include on-chip ECC. For example, while standards without such on-chip ECC have same bank write-to-write times of approximately 9 clock cycles, 10 clock cycles, or 11 clock cycles, on-chip ECC increases that time by 4-8 times, with a total same bank write-to-write delay of approximately 32 clock cycles to 64 clock cycles, depending on the particular implementation. In addition, if the write data size is smaller than the codeword used for ECC (typically 128 bits), the DRAM performs a "read-modify-write" (RMW) operation because the new ECC code must be calculated for the entire codeword, not just the modified portion. Such use of the RMW command also increases the minimum write-to-write time. When the number of bits written is equal in size to the number of bits in the ECC codeword, RMW is not required.

[0040] Figure 4 A block diagram is shown according to some embodiments. Figure 2 A portion 400 of the memory controller 200. The depicted portion 400 is suitable for use with a memory controller for DDR5, GDDR5, and other similar DRAM types, where bank groups and on-chip ECC calculations are performed by circuitry shared by all banks in a bank group. The portion 400 includes a command queue 410, an arbiter 420 with associated bank group tracking circuitry 425, and a multiplexer 440. The command queue 410 stores memory access requests received from the interface 212. The memory access request can be a read cycle or a write cycle to any bank in the memory and is generated by any of the memory access agents in the APU 100. In Figure 4 In the example of , command queue 410 has a total of 16 entries for all memory banks and memory ranks, each entry storing an access address, data in the case of a write cycle, and a tag indicating its relative age. Arbiter 420 is bidirectionally connected to each entry in command queue 410 for reading its attributes, has an additional input for receiving an access when it is dispatched, has an output for providing protocol commands, such as ACT and PRE commands, and has a control output for selecting one of the 16 entries to send to interface 212. Multiplexer 440 has 16 inputs connected to the corresponding entries of command queue 410, an additional input connected to the output of arbiter 420, a control input connected to the output of arbiter 420, and an output for providing a dispatched access to the memory PHY.

[0041] The bank group tracking circuit 425 has an input connected to the output of the arbiter 420, and a plurality of outputs connected to corresponding inputs of the arbiter 420. The bank group tracking circuit 425 includes a plurality of timers 426, each labeled "T", each connected to a corresponding bank group number entry 428, and bank group tracking logic for implementing the tracking function. In operation, when a write command is selected and dispatched to the memory interface queue, the bank group tracking circuit 425 receives an indicator from the arbiter 420. The information received in the indicator includes the bank number or bank group number of the selected write command. If a bank number is provided, the bank group tracking logic will provide the bank group number for the bank group containing the identified bank number. In response to each indicator, the bank group number for the selected command is added to one of the bank group number entries 428, and the corresponding timer 426 is activated to track the time since the dispatch of the write command by counting a number of clock cycles during which the corresponding bank group number has been tracked. Each entry 428 is connected to the arbiter 420 so that the arbiter checks which memory bank group numbers are being tracked when selecting a new write command to be dispatched to the memory. Each entry is tracked until a number of clock cycles corresponding to the minimum write-to-write timing cycle of the memory bank group counted by the corresponding timer 426 have passed, after which the memory bank group number is removed from its corresponding entry 428 and is no longer tracked. The number of memory bank group numbers tracked may vary depending on the length of the minimum write-to-write timing cycle of the memory bank group, which is typically selected based on the speed at which the DRAM chip calculates the on-chip ECC code and writes it to the DRAM together with the associated write data. For example, if the minimum write-to-write time cycle is above 40 clock cycles, four write commands may be sent to the DRAM during such time, so the memory bank group number tracking circuit has entries for tracking four commands. For longer minimum write-to-write time cycles, more memory bank group number entries are used. Preferably, sufficient entries are provided for the memory bank group tracking circuit to track commands within the maximum allowed write-to-write time cycle.

[0042] The command queue 410 stores accesses received from the interface 212 and assigns tags to indicate their relative age. The arbiter 420 determines which pending access in the command queue 410 is scheduled and dispatched to the memory interface queue 214 based on a set of policies, such as timing eligibility, age, fairness, and activity type. When selecting a command from the command queue 410, the arbiter 420 also considers the tracked memory bank group number from the memory bank group tracking circuit 425. Specifically, the arbiter 420 prevents the selection of a write request and an associated activation command for the tracked memory bank group number unless all other write requests in the command queue 410 to another memory bank group are not eligible.

[0043] Arbiter 420 includes Figure 4 420 to indicate the open pages in each memory bank and memory column of the memory system. Generally speaking, the arbitrator 420 can improve the efficiency of the memory system bus by scheduling multiple accesses to the same row together and delaying earlier accesses to different rows in the same memory bank. Therefore, the arbitrator 420 improves efficiency by selectively postponing access to rows different from the currently activated row. The arbitrator 420 also uses the aging tag of the entry to limit the waiting time for access. Therefore, when the access to another page has been waiting for a certain time, the arbitrator 420 will interrupt a series of accesses to the open pages in the memory. The arbitrator 420 also schedules accesses to other memory banks between the ACT and PRE commands to a given memory bank to hide the overhead.

[0044] As discussed above, based on the tracked memory bank group number in the entry 428 of the memory bank group tracking circuit 425, the arbiter 420 prevents the selection of a write request and an associated activation command for the tracked memory bank group number unless no other write request in the command queue for another memory bank group is eligible. If a request is ignored (not selected) because its memory bank group number is tracked, the request will become eligible after the tracking period expires. The memory bank group tracking circuit 425 indicates that the previous write request is eligible to be issued after the specified period has passed. In this embodiment, the specified period is based on the number of clock cycles corresponding to the minimum write-to-write timing period of the memory bank group of the previous write request. The indication includes at least removing the memory bank group number 428 from its corresponding entry. In some embodiments, the memory bank group tracking circuit 425 is also connected to the command queue 410 through an interface to mark or annotate each write command in the command queue 410 with an eligibility tag or a mask tag (to indicate ineligibility) based on whether the memory bank group number of the corresponding command is currently tracked in one of the memory bank group entries 428. In these embodiments, the bank group tracking circuit 425 updates the tag in the command queue 410 when a bank group number is added to or removed from the entry 428 .

[0045] Figure 5 According to some embodiments, Figure 22 is a block diagram of a portion 500 of a memory controller 200. The portion 500 depicted is suitable for use with DDR5, GDDR5, and other similar DRAM-type memory controllers with on-chip ECC calculations. The portion 500 includes an arbiter 238 and a set of control circuits 560 associated with the operation of the arbiter 238. The arbiter 238 includes a set of sub-arbiters 505 and a final arbiter 550. The sub-arbiters 505 include sub-arbiters 510, sub-arbiters 520, and sub-arbiters 530. The sub-arbiters 510 include a page hit arbiter 512 labeled "PH ARB" and an output register 514. The page hit arbiter 512 has a first input connected to the command queue 220, a second input connected to the memory bank group tracking circuit 235, a third input connected to the timing block 234, and an output. The register 514 has a data input connected to the output of the page hit arbiter 512, a clock input for receiving a UCLK signal, and an output. Sub-arbiter 520 includes a page conflict arbiter 522 labeled "PC ARB" and an output register 524. Page conflict arbiter 522 has a first input connected to command queue 220, a second input connected to memory bank group tracking circuit 235, a third input connected to timing block 234, and an output. Register 524 has a data input connected to the output of page conflict arbiter 522, a clock input for receiving a UCLK signal, and an output. Sub-arbiter 530 includes a page miss arbiter 532 labeled "PM ARB" and an output register 534. Page miss arbiter 532 has a first input connected to command queue 220, a second input connected to memory bank group tracking circuit 235, a third input connected to timing block 234, and an output. Register 534 has a data input connected to the output of page miss arbiter 532, a clock input for receiving a UCLK signal, and an output. The final arbitrator 550 has a first input connected to the output of the refresh control logic 232, a second input from the page close predictor 562, a third input connected to the output of the output register 514, a fourth input connected to the output of the output register 524, a fifth input connected to the output of the output register 534, a first output for providing a first arbitration winner marked “CMD1” to the queue 214, and a second output for providing a second arbitration winner marked “CMD2” to the queue 214.

[0046] The control circuit 560 includes the Figure 2The timing block 234 and the page table 236, the page close predictor 562, the current mode register 502 and the bank group tracking circuit 235 are described. The timing block 234 has an input connected to the page table 236, and inputs and outputs connected to the page hit arbiter 512, the page conflict arbiter 522 and the page miss arbiter 532. The page table 236 has an input connected to the output of the replay queue 230, an output connected to the input of the replay queue 230, an output connected to the input of the command queue 220, an output connected to the input of the timing block 234, and an output connected to the input of the page close predictor 562. The page close predictor 562 has an input connected to one output of the page table 236, an input connected to the output of the output register 514, and an output connected to a second input of the final arbiter 550. The bank group tracking circuit 235 has an input connected to the command queue 220 , an input and an output connected to the final arbiter 550 , and inputs and outputs connected to the page hit arbiter 510 , the page conflict arbiter 520 , and the page miss arbiter 530 .

[0047] Each of the page hit arbiter 512, the page conflict arbiter 522, and the page miss arbiter 532 has an input connected to the output of the timing block 234 to determine the timing qualifications of the commands in the command queue 220 that fall into these corresponding categories. The timing block 234 includes an array of binary counters that counts the duration associated with a specific operation for each memory bank in each storage column. The number of timers required to determine the state depends on the timing parameters, the number of memory banks of a given memory type, and the number of storage banks supported by the system on a given memory channel. The number of timing parameters implemented depends on the type of memory implemented in the system. For example, compared with other DDRx memory types, DDR5 memory and GDDR5 memory require more timers to comply with more timing parameters. By including an array of general timers implemented as binary counters, the timing block 234 can be scaled and reused for different memory types.

[0048] A page hit is a read or write cycle to open a page. However, in this embodiment, page hit tracking for write commands is also controlled by the bank group tracking process. If the bank group number of a candidate command is currently tracked by the bank group tracking circuit 235, then a page hit for a write command is not eligible to be selected unless no other candidates are available. The page hit arbitrator 512 arbitrates between accesses in the command queue 220 to open a page. Timing eligibility parameters tracked by the timer in the timing block 234 and checked by the page hit arbitrator 512 include, for example, the row address strobe (RAS) to column address strobe (CAS) delay time (t RCD ) and CAS waiting time (t CL ). For example, t RCDSpecifies the minimum amount of time that must pass before read access to a page after it has been opened in a RAS cycle. The page hit arbiter 512 selects the sub-arbitration winner based on the assigned access priority. In one embodiment, the priority is a 4-bit, unique hot value, thus indicating the priority among four values, however, it should be apparent that this four-level priority scheme is only an example. If the page hit arbiter 512 detects two or more requests of the same priority level, the oldest entry wins.

[0049] A page conflict is an access to a row in a bank when another row in the bank is currently active. The page conflict arbitrator 522 arbitrates between accesses to pages in the command queue 220 that conflict with pages currently open in the corresponding bank and rank. The page conflict arbitrator 522 selects the sub-arbitration winner that results in the issuance of a precharge command. The timing eligibility parameters tracked by the timer in the timing block 234 and checked by the page conflict arbitrator 522 include, for example, the number of active precharge command cycles (t RAS ). The page conflict sub-arbitration for write commands also considers the bank group number of the candidate write command. If the bank group number of the candidate command is currently tracked by the bank group tracking circuit 235, the candidate write command is not eligible to be selected unless no other candidates are available. The page conflict arbitrator 522 selects the sub-arbitration winner based on the assigned access priority. If the page conflict arbitrator 522 detects two or more requests of the same priority level, the oldest entry wins.

[0050] A page miss is an access to a memory bank that is in a precharged state. The page miss arbiter 532 arbitrates between accesses to precharged memory banks in the command queue 220. Timing eligibility parameters tracked by a timer in the timing block 234 and checked by the page miss arbiter 532 include, for example, the precharge command period (t RP ). If there are two or more requests that are page misses at the same priority level, the oldest entry wins.

[0051] Each sub-arbitrator outputs a priority value for their respective sub-arbitration winner. The final arbitrator 550 compares the priority values ​​of the sub-arbitration winners from each of the page hit arbitrator 512, the page conflict arbitrator 522, and the page miss arbitrator 532. The final arbitrator 550 determines the relative priority between the sub-arbitration winners by performing a set of relative priority comparisons that consider two sub-arbitration winners at a time. The sub-arbitrators may include a set of logic for arbitrating commands for read and write for each mode so that when the current mode changes, a set of available candidate commands as sub-arbitration winners are quickly available.

[0052] After determining the relative priorities between the three sub-arbitration winners, the final arbitrator 550 then determines whether these sub-arbitration winners conflict (i.e., whether they point to the same memory bank and memory column). When there is no such conflict, the final arbitrator 550 selects at most two sub-arbitration winners with the highest priority. When there is a conflict, the final arbitrator 550 follows the following rules. When the priority value of the sub-arbitration winner of the page hit arbitrator 512 is higher than the priority value of the sub-arbitration winner of the page conflict arbitrator 522 and both point to the same memory bank and memory column, the final arbitrator 550 selects the access indicated by the page hit arbitrator 512. When the priority value of the sub-arbitration winner of the page conflict arbitrator 522 is higher than the priority value of the sub-arbitration winner of the page hit arbitrator 512 and both point to the same memory bank and memory column, the final arbitrator 550 selects the winner based on several additional factors. In some cases, the page close predictor 562 causes the page to be closed at the end of the access indicated by the page hit arbitrator 512 by setting the automatic precharge attribute.

[0053] Within the page hit arbiter 512, the priority is initially set by the request priority from the memory access agent, but the priority is dynamically adjusted based on the access type (read or write) and the access sequence. In general, the page hit arbiter 512 assigns a higher implicit priority to reads, but implements a priority boost mechanism to ensure that writes make progress toward completion.

[0054] Whenever the page hit arbiter 512 selects a read or write command, the page close predictor 562 determines whether to send a command with an automatic precharge (AP) attribute. The automatic precharge attribute is set with a predefined address bit during the read or write cycle, and the automatic precharge attribute causes the DDR device to close the page after the read or write cycle is completed, which avoids the need for the memory controller to send a separate precharge command for the memory bank later. The page close predictor 562 considers other requests that already exist in the command queue 220 to access the same memory bank as the selected command accesses. If the page close predictor 562 converts the memory access into an AP command, the next access to the page will be a page miss.

[0055] By using different sub-arbiters for different memory access types, each arbiter can be implemented with simpler logic than if arbitration were required between all access types (page hits, page misses, and page conflicts), although embodiments including a single arbiter are contemplated. Thus, the arbitration logic can be simplified and the size of arbiter 238 can be kept relatively small.

[0056] In other embodiments, arbitrator 238 may include a different number of sub-arbiters. In yet another embodiment, arbitrator 238 may include two or more sub-arbiters of a specific type. For example, arbitrator 238 may include two or more page hit arbiters, two or more page conflict arbiters, and / or two or more page miss arbiters.

[0057] In operation, the arbiter 238 selects memory access commands from the command queue 220 and the refresh control logic 232 by considering the memory bank group number tracked by the memory bank group tracking circuit 235, the page state of each entry, the priority of each memory access request, and the correlation between these requests. The priority is related to the quality of service or QoS of the request received from the AXI4 bus and stored in the command queue 220, but can be changed based on the type of memory access and the dynamic operation of the arbiter 238. The arbiter 238 includes three sub-arbiters that operate in parallel to address the mismatch between the processing and transmission limitations of existing integrated circuit technologies. The winners of the corresponding sub-arbitrations are presented to the final arbiter 550. In some embodiments, each winner is marked to indicate whether it has a currently tracked memory bank number, but is still selected because no other suitable write commands are available. The final arbitrator 550 selects between the three sub-arbitration winners and selects the refresh operation from the refresh control logic 232 , and may further modify the read or write command to a read or write with an auto-precharge command as determined by the page close predictor 562 .

[0058] Based on the tracked bank group number in the entry 428 of the bank group tracking circuit 235, the arbiter 238 prevents selection of a write request and associated activate command for the tracked bank group number unless no other write requests in the command queue are eligible. If a request is ignored (not selected) because its bank group number is tracked, the request will become eligible after the tracking period expires. The bank group tracking circuit 235 indicates that a previous write request is eligible to be issued after a number of clock cycles of the minimum write-to-write timing period of the bank group corresponding to the previous write request have passed by removing the restricted bank group number from its tracking entry.

[0059] Figure 6 A flowchart 600 is shown of a process for operating a bank group tracking circuit to track write commands according to some embodiments. The process is suitable for being performed by the bank group tracking circuit 425 ( Figure 4 ), memory bank group tracking circuit 235 ( Figure 5) or another suitable digital logic circuit connected to the arbiter. When the arbiter selects a write request from the command queue for transmission to the DRAM memory, the process starts at box 602. An associated activation command (ACT) for the selected write command is dispatched before the write command to activate the row for the write command. Intermediate commands can be dispatched between the associated ACT and the selected write command. The selected write command is de-allocated from the command queue and sent to the memory interface queue or memory PHY for transmission to the DRAM memory. The memory bank group tracking circuit monitors the dispatched write command or is notified by the arbiter or memory interface queue that the write command is dispatched. In response to such notification, at box 604, the memory bank group tracking circuit adds the memory bank group number of the dispatched write command to its memory bank group tracking entry. At box 604, a timer can be started or input with a timer tracking circuit to count the clock cycles that have passed after the write command is dispatched.

[0060] For the tracked storage body group number, the process at box 606 includes preventing the search for subsequent write requests and their associated activation commands for the tracked storage body group number unless all other write requests in the command queue are not eligible. Typically, the selection process performed at the arbiter circuit loops through the untracked storage body group numbers and checks whether the write request for each corresponding storage body group is eligible to be selected until a qualified write request is found. At box 608, after a number of clock cycles corresponding to the minimum write-to-write timing cycle of the storage body group of the previous write request have passed, the process removes the tracking of the storage body group number. If any previous write request has been passed for selection of a specific storage body group, the box makes such previous write request eligible to be issued after the minimum minimum write-to-write timing cycle of the storage body group has passed since the previous write request was issued to the storage body group.

[0061] Typically, the process involves tracking the bank group numbers of three or more previous write requests selected by the arbitrator. The number of bank group numbers tracked varies depending on the length of the bank group's minimum write-to-write timing cycle, which is typically controlled by the speed at which the DRAM chip calculates the on-chip ECC code and writes it to the DRAM along with the associated write data.

[0062] Figure 7 Flowchart 700 shows a process for arbitrating commands according to some embodiments. The depicted process is suitable for use with, for example, Figure 5 At block 702, the process is performed for the sub-arbiter that is selecting a candidate request.

[0063] At block 704, the process applies the subarbiter selection policy when selecting candidate requests, and also applies the bank group number mask for the tracked bank group number when checking the command queue for eligible write requests. As described above, a write command to a tracked bank group is not eligible, or "masked", unless no other write commands are eligible for selection.

[0064] At block 706, the process provides candidate write requests from each sub-arbitrator. Multiple sub-arbitrators of the arbitrator typically perform the process, and each sub-arbitrator will provide candidate requests, such as Figure 5 At block 708, the final arbiter selects the deallocated write request from the command queue and transmits it to the DRAM memory. Because the memory bank group number mask is applied at the sub-arbiter, the final arbiter does not need to consider the memory bank group tracking process.

[0065] Figure 2 The memory controller 200 or any part thereof (such as the arbiter 238 and the refresh control logic 232) can be described or represented by a computer accessible data structure in the form of a database or other data structure that can be read by a program and used directly or indirectly to manufacture the integrated circuit. For example, the data structure can be a behavioral level description or a register transfer level (RTL) description of the hardware functionality in a high-level design language (HDL) such as Verilog or VHDL. The description can be read by a synthesis tool, which can synthesize the description to produce a netlist including a list of gates from a synthesis library. The netlist includes a set of gates, which also represents the functionality of the hardware comprising the integrated circuit. The netlist can then be placed and routed to produce a data set describing the geometry to be applied to the mask. The mask can then be used in various semiconductor manufacturing steps to produce the integrated circuit. Alternatively, the database on the computer accessible storage medium can be a netlist (with or without a synthesis library) or a data set (as required) or a graphic data system (GDS) II data.

[0066] Although specific embodiments have been described, various modifications to these embodiments will be apparent to those skilled in the art. For example, the internal architecture of the memory channel controller 210 and / or the power engine 250 may vary in different embodiments. The memory controller 200 may be interfaced to other types of memory other than DDRx, such as high bandwidth memory (HBM), RAMbus DRAM (RDRAM), etc. Although the illustrated embodiments show each memory storage rank corresponding to a separate DIMM or SIMM, in other embodiments, each module may support multiple storage ranks. Still other embodiments may include other types of DRAM modules or DRAM not included in a particular module, such as DRAM mounted to a host motherboard. Therefore, the appended claims are intended to cover all modifications of the disclosed embodiments that fall within the scope of the disclosed embodiments.

Claims

1. A method comprising: receiving a plurality of memory access requests at a memory controller and placing the plurality of memory access requests in a command queue to await transmission to a dynamic random access memory (DRAM) comprising a plurality of memory bank groups; selecting, using an arbiter, a write request from the plurality of memory access requests in the command queue for transmission to the DRAM; tracking memory bank group numbers of three or more previous write requests selected by the arbitrator from the plurality of memory access requests; preventing subsequent write requests and associated activate commands from selecting a corresponding memory bank group of the tracked memory bank group number; and A write request that has been communicated to select one of the tracked memory bank groups and an associated one of the associated activate commands are made eligible to be issued after a specified period has elapsed.

2. The method according to claim 1, further comprising: When a new write request is selected, the untracked bank group numbers are looped through and the write request for each corresponding bank group is checked to see if it qualifies for selection until a qualifying write request is found.

3. The method of claim 1, further comprising counting, with at least one timer, a number of clock cycles during which each of the memory bank group numbers has been tracked.

4. The method of claim 1 , further comprising selecting a memory access request from the plurality of memory access requests in the command queue based on the tracked memory bank group number and a set of policies including timing eligibility, aging, fairness, and activity type.

5. The method according to claim 1, further comprising: selectively selecting a candidate memory request from the plurality of memory access requests in the command queue based on first class accesses, second class accesses, and information provided by a memory bank group tracking circuit, wherein each class of accesses corresponds to a different page state of a memory bank in the memory; Request selections from the candidate memory requests are then arbitrated. 6 . The method of claim 1 , wherein the designated period is a plurality of clock periods corresponding to a minimum write-to-write timing period of the memory bank group of the previous write request.

7. A memory controller comprising: a command queue circuit having an input for receiving a memory access request for a memory channel and a plurality of entries for holding a predetermined number of the memory access requests; and an arbiter circuit for selecting the memory access request from the command queue for transmission to a dynamic random access memory (DRAM) coupled to a DRAM channel, the arbiter comprising: a bank group tracking circuit that tracks bank group numbers of three or more previous write requests selected by the arbiter circuit from among the memory access requests; and a selection circuit that selects a request to be issued from the memory access requests in the command queue circuit and prevents selection of a write request and an associated activate command for the tracked memory bank group number, Wherein the memory bank group tracking circuit indicates that a write request that has been communicated to select one of the tracked memory bank groups and an associated one of the associated activate commands are eligible to be issued after a specified period has elapsed.

8. The memory controller of claim 7, wherein when selecting a new write request, the selection circuit cycles through the untracked memory bank group numbers and checks whether the write request for each corresponding memory bank group is eligible to be selected until a qualifying write request is found.

9. The memory controller of claim 7, wherein the bank group tracking circuit comprises at least one timer circuit for counting a number of clock cycles during which each of the bank group numbers has been tracked.

10. The memory controller of claim 7, wherein the bank group tracking circuit comprises: A plurality of entries, each entry of which stores a memory bank group number; and a plurality of timers, each of which is associated with a corresponding one of the plurality of entries of the memory bank group tracking circuit and counts a number of clock cycles during which a corresponding one of the memory bank group numbers has been tracked.

11. The memory controller of claim 7 , wherein the selection circuit comprises a plurality of inputs, each of which receives a tracked memory bank group number, the selection circuit selecting a memory access request from the memory access requests in the command queue circuit based on the tracked memory bank group number and a set of policies including timing eligibility, aging, fairness, and activity type.

12. The memory controller according to claim 7, wherein the selection circuit comprises: at least one sub-arbiter circuit, the at least one sub-arbiter circuit configured to selectively select candidate memory requests from the command queue circuit according to a first type of access and a second type of access, wherein each type of access corresponds to a different page state of a bank in the memory, the at least one sub-arbiter circuit selecting a write request based on information provided by the bank group tracking circuit; and A final arbiter circuit arbitrates request selections from the candidate memory requests.

13. The memory controller of claim 7, wherein the designated cycle is a plurality of clock cycles corresponding to a minimum write-to-write timing cycle of a bank group number among the bank group numbers of a previous write request among the three or more previous write requests.

14. A data processing system comprising: a memory channel coupled to a dynamic random access memory (DRAM); and A memory controller, coupled to the memory channel, comprising: a command queue circuit having an input for receiving a memory access request for the memory channel and a plurality of entries for holding a predetermined number of the memory access requests; An arbiter circuit, the arbiter circuit being used to select a memory request to be issued from the memory access requests in the command queue circuit for transmission to the DRAM, the arbiter comprising: a bank group tracking circuit that tracks bank group numbers of three or more previous write requests selected by the arbiter circuit from among the memory access requests; and a selection circuit that prevents selection of a write request for the tracked memory bank group number from the memory access requests in the command queue circuit and prevents selection of an associated activate command, Wherein the memory bank group tracking circuit indicates that a write request that has been communicated to select one of the tracked memory bank groups and an associated one of the associated activate commands are eligible to be issued after a specified period has elapsed.

15. The data processing system of claim 14, wherein when selecting a new write request, the selection circuit cycles through the untracked memory bank group numbers and checks whether the write request for each corresponding memory bank group is eligible to be selected until a qualifying write request is found.

16. The data processing system of claim 14, wherein the memory bank group tracking circuit comprises at least one timer circuit for counting a number of clock cycles during which each of the memory bank group numbers has been tracked.

17. The data processing system of claim 14, wherein the bank group tracking circuit comprises: A plurality of entries, each entry of which stores a memory bank group number; and a plurality of timers, each of which is associated with a corresponding one of the plurality of entries in the memory bank group tracking circuit and counts a number of clock cycles during which a corresponding memory bank group number among the memory bank group numbers has been tracked.

18. The data processing system of claim 14 , wherein the selection circuit comprises a plurality of inputs, each of which receives a tracked memory bank group number, the selection circuit selecting a memory access request from among the memory access requests in the command queue circuit based on the tracked memory bank group number and a set of policies including timing eligibility, aging, fairness, and activity type.

19. The data processing system of claim 14, wherein the selection circuit comprises: at least one sub-arbiter circuit, the at least one sub-arbiter circuit selectively selecting candidate memory access requests from the command queue circuit according to a first type of access and a second type of access, wherein each type of access corresponds to a different page state of a bank in the memory, the at least one sub-arbiter circuit selecting a write request based on information provided by the bank group tracking circuit; and A final arbiter circuit arbitrates a request selection from the candidate memory access requests.

20. The data processing system of claim 14, wherein the designated period is a plurality of clock periods corresponding to a minimum write-to-write timing period of a bank group number of a previous write request among the three or more previous write requests among the bank group numbers.

Citation Information

Patent Citations

  • Memory controller with bank sorting and scheduling

    US20070156946A1