Channel and Subchannel Slot Ringing for Memory Controllers
The traffic throttle circuit in memory controllers addresses power consumption spikes by monitoring and limiting commands during low-activity states, enhancing power management and stability in server systems with multiple memory controllers.
Patent Information
- Application Number
- JP2024576455
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-29
- Filing Date
- 2023-06-22
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-06-22
AI Technical Summary
Modern server systems with multiple memory controllers experience harmful power consumption spikes due to sudden increases in memory traffic, which can occur over a short period, leading to voltage droops and inefficient power management.
Implementing a traffic throttle circuit in memory controllers that includes a monitoring circuit to measure read and write commands over a predetermined period and a throttle circuit to limit command issuance during low-activity states, thereby mitigating excessive power consumption.
The solution effectively manages power consumption by regulating the number of read and write commands, preventing voltage droops and ensuring stable operation of memory channels.
Smart Images

Figure 2025522593000001_ABST
Abstract
Description
Background Art
[0001] Computer systems generally use inexpensive and high-density dynamic random access memory (DRAM) chips for main memory. Most DRAM chips sold today are compatible with various double data rate (DDR) DRAM standards promoted by the Joint Electron Devices Engineering Council (JEDEC). DDR DRAM uses a conventional DRAM memory cell array with a high-speed access circuit to achieve a high transfer rate and improve the utilization of the memory bus. A DDR memory controller can interface with multiple DDR channels to accommodate more DRAM modules and exchange data with the memory faster using a single channel. Furthermore, modern server systems often include multiple memory controllers within a single data processor. For example, some modern server processors include eight or twelve memory controllers, each connected to a respective DDR channel.
[0002] Traffic on the memory channel may be throttled or slowed down to conserve power. Such throttling is performed over a relatively long time period with respect to the memory clock speed and is achieved, for example, by adjusting the memory clock and / or the memory controller clock. Traffic on the memory channel is often "bursty," i.e., a period of high traffic may follow a period of idle or low traffic. The memory controller may be placed in a low-power mode where it does not issue commands to the memory and may start issuing commands when the low-power mode ends. In particular, when multiple memory controllers are included, such a sudden increase in traffic can cause harmful spikes in the current consumed by the memory controller and the memory channel circuitry. These spikes can occur over a relatively short time period compared to the memory clock speed.
Brief Description of the Drawings
[0003]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Best Mode for Carrying Out the Invention
[0004] In the following description, the use of the same reference numerals in different drawings indicates similar or identical items. Unless otherwise noted, the word "coupled" and its related verbs include both direct connections and indirect electrical connections by means known in the art, and unless otherwise noted, any description of a direct connection also means an alternative embodiment using a suitable form of indirect electrical connection.
[0005] The memory controller is operable to retrieve commands from a command queue for dispatch to the memory. The memory controller includes an arbiter and a traffic throttle circuit that cooperates with one or more additional arbiters to mitigate an increase in excessive power consumption. The traffic throttle circuit includes a monitoring circuit and a throttle circuit. The monitoring circuit is for measuring the number of read and write commands selected by the arbiter and one or more additional arbiters over a first predetermined period. The throttle circuit limits the number of read and write commands issued by the arbiter during a second predetermined period in response to a low-activity state associated with the read and write commands.
[0006] The method cooperates with one or more additional arbiters to mitigate an increase in excessive power consumption in the arbiter. The method includes measuring the number of read and write commands selected by the arbiter and one or more additional arbiters over a first predetermined period. The method includes limiting the number of read and write commands issued by the arbiter during a second predetermined period in response to a low-activity state associated with the read and write commands.
[0007] The processing system includes a data processor and a plurality of memory channels that connect the data processor to memory. The memory controller is operable to retrieve commands from a command queue for dispatching to memory corresponding to any of the memory channels. The memory controller includes an arbiter and a traffic throttle circuit that cooperates with one or more additional arbiters for each memory channel to mitigate an increase in excessive power usage. The traffic throttle circuit includes a monitoring circuit and a throttle circuit. The monitoring circuit measures the number of read commands and write commands selected by the arbiter and the one or more additional arbiters over a first predetermined period. The throttle circuit is connected to the monitoring circuit and the arbiter and restricts the number of read commands and write commands issued by the arbiter during a second predetermined period in response to a low-activity state associated with the read commands and write commands.
[0008] FIG. 1 is a block diagram of a data processing system 100 according to the prior art. The data processing system 100 generally includes a data processor 110 in the form of an accelerated processing unit (APU), a memory system 120, a peripheral component interconnect express (PCIe) system 150, a universal serial bus (USB) system 160, and a disk drive 170. The data processor 110 operates as the central processing unit (CPU) of the data processing system 100 and provides various buses and interfaces useful in modern computer systems. These interfaces include two double data rate (DDRx) memory channels, a PCIe root complex for connection to a PCIe link, a USB controller for connection to a USB network, and an interface to a serial advanced technology attachment (SATA) mass storage device.
[0009] Memory system 120 includes memory channel 130 and memory channel 140. Memory channel 130 includes a set of dual in-line memory modules (DIMMs) connected to DDRx bus 132, which includes representative DIMMs 134, 136, 138 corresponding to individual ranks in this example. Similarly, memory channel 140 includes a set of DIMMs connected to DDRx bus 142, including representative DIMMs 144, 146, 148.
[0010] PCIe system 150 includes a PCIe switch 152 connected to a PCIe root complex in data processor 110, a PCIe device 154, a PCIe device 156, and a PCIe device 158. PCIe device 156 is connected to system basic input / output system (BIOS) memory 157. System BIOS memory 157 can be any of various non-volatile memory types such as read-only memory (ROM), flash electrically erasable programmable ROM (EEPROM), etc.
[0011] USB system 160 includes a USB hub 162 connected to a USB master in data processor 110 and representative USB devices 164, 166, 168 respectively connected to USB hub 162. USB devices 164, 166, 168 can be devices such as a keyboard, a mouse, a flash EEPROM port, etc.
[0012] Disk drive 170 is connected to data processor 110 via an SATA bus and provides a mass storage device for an operating system, application programs, application files, etc.
[0013] The data processing system 100 is suitable for use in up-to-date computing applications by providing memory channels 130 and 140. Each of the memory channels 130 and 140 can be connected to up-to-date technology DDR memories such as DDR version 4 (DDR4), low-power DDR4 (LPDDR4), graphics DDR version 5 (GDDR5), and high-bandwidth memory (HBM), and can be adapted to future memory technologies. These memories provide high bus bandwidth and high-speed operation. At the same time, those memories provide a low-power mode for saving power for battery-powered applications such as laptop computers, and also provide a built-in thermal monitoring function.
[0014] FIG. 2 is a block diagram of an APU 200 suitable for use in a system such as the data processing system 100 of FIG. 1. The APU 200 generally includes a central processing unit (CPU) core complex 210, a graphics core 220, a set of display engines 230, a memory management hub 240, a data fabric 250, a set of peripheral controllers 260, a set of peripheral bus controllers 270, a system management unit (SMU) 280, and a set of memory controllers 290.
[0015] The CPU core complex 210 includes CPU cores 212 and 214. In this example, the CPU core complex 210 includes two CPU cores, but in other embodiments, the CPU core complex 210 can include any number of CPU cores. Each of the CPU cores 212 and 214 is bidirectionally connected to a system management network (SMN) forming a control fabric and to the data fabric 250, and can provide a memory access request to the data fabric 250. Each of the CPU cores 212 and 214 can be a single core or a core complex having two or more single cores sharing specific resources such as a cache.
[0016] The Graphics Core 220 is a high-performance Graphics Processing Unit (GPU) that can execute graphics processing such as vertex processing, fragment processing, shading, texture blending, etc. in a highly integrated parallel manner. The Graphics Core 220 is bidirectionally connected to the SMN and to the Data Fabric 250, and can provide memory access requests to the Data Fabric 250. In this regard, the APU 200 can support either an integrated memory architecture in which the CPU Core Complex 210 and the Graphics Core 220 share the same memory space, or a memory architecture in which the CPU Core Complex 210 and the Graphics Core 220 share part of the memory space while the Graphics Core 220 also uses private graphics memory that cannot be accessed by the CPU Core Complex 210.
[0017] The Display Engine 230 renders and rasterizes the objects generated by the Graphics Core 220 for display on the monitor. The Graphics Core 220 and the Display Engine 230 are bidirectionally connected to a common Memory Management Hub 240 for uniform translation to appropriate addresses within the Memory System 120, and the Memory Management Hub 240 is bidirectionally connected to the Data Fabric 250 to generate such memory accesses and receive the read data returned from the Memory System.
[0018] The Data Fabric 250 includes a crossbar switch for routing memory access requests and memory responses between any memory access agent and the Memory Controller 290. The Data Fabric 250 also includes a system memory map defined by the BIOS for determining the destination of memory accesses based on the system configuration, and buffers for each virtual connection.
[0019] The peripheral controller 260 includes a USB controller 262 and a SATA interface controller 264, each of which is connected bidirectionally to the system hub 266 and the SMN bus. These two controllers are merely typical examples of peripheral controllers that can be used in the APU 200.
[0020] The peripheral bus controller 270 includes a system controller or "Southbridge" (SB) 272 and a PCIe controller 274, each of which is connected bidirectionally to the input / output (I / O) hub 276 and to the SMN bus. Also, the I / O hub 276 is connected bidirectionally to the system hub 266 and to the data fabric 250. Thus, for example, a CPU core can program the registers in the USB controller 262, the SATA interface controller 264, the SB 272, or the PCIe controller 274 through accesses routed by the data fabric 250 via the I / O hub 276.
[0021] The SMU 280 is a local controller that controls the operation of the resources on the APU 200 and synchronizes the communication between them. The SMU 280 manages the power-up sequencing of various processors on the APU 200 and controls a plurality of off-chip devices via reset signals, enable signals, and other signals. The SMU 280 includes one or more clock sources (not shown in FIG. 2), such as a phase-locked loop (PLL), to provide clock signals to each of the components of the APU 200. Also, the SMU 280 manages the power for various processors and other functional blocks and can receive measured power consumption values from the CPU cores 212, 214, and the graphics core 220 to determine appropriate power states.
[0022] The set 290 of memory controllers includes, in this embodiment, memory controllers 292, 294, 296, and 298. Each memory controller has an upstream bi-directional connection to the data fabric 250 and a downstream connection to a memory channel for accessing a memory such as a DDR memory, as will be further described below.
[0023] Figure 3 is a block diagram of a memory system 300 including a plurality of memory controllers according to some embodiments. The memory system 300 is generally implemented on an integrated circuit such as an APU, GPU, or CPU system-on-chip, similar to that of FIG. 2. The memory system 300 includes memory controllers 310, 320, 330 and DDR PHY interfaces 340, 342, 344 for interfacing with one or more external memory modules. Also, the memory system 300 generally includes a power distribution network (PDN) including a first power domain 302 labeled as a "VDD domain" in which the circuitry is powered from a power supply of voltage VDD and a second power domain 304 labeled as a "VDDP domain" in which the circuitry is powered from a power supply of voltage VDDP. Although three memory controllers are shown, in most embodiments in modern server systems, the system-on-chip can have many more DDR channels with many more memory controllers, such as 12, 16, or more. In some embodiments, such memory controllers are arranged within "quadrants" where groups of memory controllers adjust slotting between the groups. Memory controllers 310, 320, 330 form such groups in this embodiment.
[0024] Memory channel controllers 310, 320, 330 are generally supplied within the first power domain 302 and may include lower-level domains specific to each memory controller below the first power domain 302 in the PDN. Similarly, DDR PHY interfaces 340, 342, 344 are generally supplied in the second power domain 304 and may include lower-level domains below the second power domain 304 in the PDN that are specific to each DDR PHY interface. Generally, during operation, memory channel controllers 310, 320, 330 handle various traffic loads that exhibit "bursty" behavior where the memory channel controller is idle and then quickly becomes busy to satisfy memory access requests. Such behavior tends to cause voltage droops in one or both of the VDDP domain 304 and the VDD domain 302. In particular, when all memory controllers within a quadrant rise from idle almost simultaneously and begin transmitting read (RD) commands and write (WR) commands at their full bandwidth capacity, voltage droops on the order of nanoseconds (10 - 20 nanoseconds) tend to occur.
[0025] Each memory channel controller includes its respective one or more arbiters 312, 322, 332 and its respective traffic slot ring circuits 314, 324, 334. Each traffic slot ring circuit is bi-directionally connected to other traffic slot ring circuits to cooperate with arbiters of other memory controllers to mitigate an increase in excessive power usage. In this embodiment, each traffic slot circuit 314, 324, 334 includes a monitoring circuit for measuring the number of read and write commands picked up by its respective arbiter and other arbiters over a first predetermined period, and a throttle circuit coupled to the monitoring circuit and its respective arbiter. Each throttle circuit operates to limit the number of read and write commands issued by the arbiter during a second predetermined period in response to low-activity states associated with the read and write commands, and this can be repeated as further described below.
[0026] Figure 4 is a block diagram of a partial data processing system 400 including a dual-channel memory controller 410 suitable for use in an APU as shown in Figure 2. The dual-channel memory controller 410 connected to the data fabric 250 is shown, and the dual-channel memory controller 410 can communicate with the data fabric 250 using several memory agents present within the data processing system 400. The dual-channel memory controller 410 can control two sub-channels, as defined in the DDR5 specification for use with, for example, DDR5 DRAM, or those sub-channels defined in the High Bandwidth Memory 2 (HBM2) and HBM3 specifications. In some embodiments, instead of two sub-channels, the dual-channel memory controller 410 replaces two individual memory channel controllers and can control two DDRx channels together in a transparent manner to various memory addressing agents within the data fabric 250 and the APU 200, such that memory access commands can be sent and results received using a single memory controller interface 412. The dual-channel memory controller 410 generally includes an interface 412, an address decoder 422, and two instances of memory channel control circuits 423A and 423B respectively assigned to different memory channels. Each instance of the memory channel control circuits 423A, 423B includes a memory interface queue 414A, 414B, command queues 420A, 420B, content-addressable memory (CAM) 424A, 424B, timing blocks 434A, 434B, page tables 436A, 436B, arbiters 438A, 438B, ECC generation blocks 444A, 444B, data buffers 446A, 446B, and traffic throttle circuits 432A, 432B including window monitoring circuits 448A, 448B.
[0027] In other embodiments, only the command queues 430A, 430B, arbiters 438A, 438B, traffic throttle circuits 432A, 432B, and memory interface queues 414A, 414B are replicated for each memory channel or subchannel used, and the remaining illustrated circuits are adapted to be used with two channels. Further, the illustrated dual-channel memory controller 410 includes two instances of arbiters 438A, 438B, command queues 420A, 420B, and memory interface queues 414A, 414B to control two memory channels or subchannels, but other embodiments may include more instances, such as three or four or more, for communicating with DRAM on three or four channels or subchannels according to the credit management techniques of this specification. Further, although a dual-channel memory controller is used in this embodiment, other embodiments can use the slotting techniques of this specification with a single-channel memory controller.
[0028] Interface 412 has a first bidirectional connection to data fabric 250 via a communication bus and a second bidirectional connection to credit control circuit 421. In this embodiment, interface 412 uses scalable data port (SDP) links to establish several channels for communicating with data fabric 250, although other interface link standards are also suitable for use. For example, in another embodiment, the communication bus is compatible with Advanced eXtensible Interface version 4, known as "AXI4", specified by ARM Holdings, PLC of Cambridge, UK, although in other embodiments it may be other types of interfaces. Interface 412 converts memory access requests from a first clock domain known as the "FCLK" (or "MEMCLK") domain to a second clock domain inside dual-channel memory controller 410 known as the "UCLK" domain. Similarly, memory interface queue 414 provides memory access from the UCLK domain to the "DFICLK" domain associated with the DFI interface.
[0029] The address decoder 422 has a bidirectional link to the interface 412, a first output connected to the first command queue 420A (labeled as "command queue 0"), and a second output connected to the second command queue 420B (labeled as "command queue 1"). The address decoder 422 decodes the address of the memory access request received on the data fabric 250 via the interface 412. The memory access request includes an access address within the physical address space represented in a normalized format. Based on the access address, the address decoder 422 selects one of the memory channels having the associated command queue among the command queues 420A and 420B to process the request. The selected channel is identified to the credit control circuit 421 for each request so that a credit issuance decision can be made. The address decoder 422 converts the normalized address into a format that can be used to address the actual memory devices within the memory system 130 and to efficiently schedule the associated accesses. This format includes a region identifier that associates the memory access request with a specific rank, row address, column address, bank address, and bank group. At startup, the system BIOS queries the memory devices within the memory system 130 to determine their sizes and configurations and programs a set of configuration registers associated with the address decoder 422. The address decoder 422 uses the configuration stored in the configuration registers to convert the normalized address into an appropriate format. Each memory access request is loaded into the command queue 420A or 420B for the memory channel selected by the address decoder 422.
[0030] Each command queue 420A, 420B is a queue of memory access requests received from various memory access engines within the APU 200 such as the CPU cores 212 and 214 and the graphics core 220. Each command queue 420A, 420B is bi-directionally connected to respective arbiters 438A, 438B to select memory access requests issued via the associated memory channel from the command queues 420A, 420B. Each command queue 420A, 420B stores an address field decoded by the address decoder 422, as well as other address information that enables the respective arbiters 438A, 438B to efficiently select memory accesses including an access type and a quality of service (QoS) identifier. Each CAM 424A, 424B contains information for imposing ordering rules such as write-after-write (WAW) and read-after-write (RAW) ordering rules.
[0031] Each of arbiters 438A and 438B is bi-directionally connected to respective command queues 420A and 420B to select memory access requests satisfied by appropriate commands, and is bi-directionally connected to respective traffic throttling circuits 432A and 432B to receive throttling signals. Arbiters 438A and 438B generally improve the efficiency of their respective memory channels by intelligent scheduling of accesses to improve the use of the memory bus of the memory channel. Each arbiter 438A and 438B uses respective timing blocks 434A and 434B to impose appropriate timing relationships by determining whether specific accesses within respective command queues 420A and 420B are eligible for issuance based on DRAM timing parameters. Each page table 436A and 436B maintains state information regarding active pages in respective banks and ranks of respective memory channels for respective arbiters 438A and 438B, and is bi-directionally connected to respective replay queues 430A and 430B. Each arbiter 438A and 438B efficiently schedules memory accesses while complying with other criteria such as quality of service (QoS) requirements, using decoded address information, timing eligibility information indicated by timing blocks 434A and 434B, and active page information indicated by page tables 436A and 436B. In some embodiments, arbiters 438A and 438B determine whether the selected commands are permitted to be released based on signals from traffic throttling circuits 432A and 432B.
[0032] Each traffic slot ring circuit 432A, 432B is bidirectionally connected to its respective arbiter 438A, 438B and is bidirectionally connected to one or more other traffic slot ring circuits that may be within other memory controllers. As shown, traffic slot ring circuits 432A, 432B are bidirectionally connected to each other to provide signals labeled CASSENT to indicate to each other when a column address strobe (CAS) command is sent to the memory channel. The CASSENT signal is used in the slot ring process as further described below with respect to FIGS. 5 - 7. Each traffic slot ring circuit includes window monitoring circuits 448A, 448B for measuring the number of read and write commands selected by its respective arbiter 438A, 438B and one or more additional arbiters over a first predetermined period. Generally, as further described below, each traffic slot ring circuit 432A, 432B operates to limit the number of read and write commands issued by arbiters 438A, 438B during a second predetermined period in response to low activity states associated with the read and write commands. Each traffic slot ring circuit 432A, 432B can also send an "IDLE" signal to indicate an idle state, and a period of a specified length during which the CAS signal is not dispatched has been reached at the arbiter of the transmitting traffic slot ring circuit. In response to this IDLE signal, traffic slot ring circuits 432A, 432B that are currently in the slot ring mode adjust their own slot ring rate to allow more CAS commands to be sent.
[0033] Each error correction code (ECC) generation block 444A, 444B determines the ECC of the write data to be sent to the memory. In response to a write memory access request received from the interface 412, the ECC generation blocks 444A, 444B calculate the ECC according to the write data. The data buffers 446A, 446B store the write data and the ECC related to the received memory access request. When the respective arbiters 438A, 438B select the corresponding write access for dispatch to the memory channel, the data buffers 446A, 446B output the combined write data / ECC to the respective memory interface queues 414A, 414B.
[0034] FIG. 5 is a flowchart 500 of a process for throttling traffic in a memory controller according to some embodiments. This process is suitable for use with the memory system of FIG. 3, the memory controller of FIG. 4, or with other suitable systems including a traffic throttling circuit as shown in FIGS. 3, 4, or 6, using multiple memory channels or sub-channels.
[0035] The process starts at block 502, where the traffic slot ring circuit monitors some arbiters or sub-channel arbiters for some memory access commands (specifically, reads and writes) selected for dispatch to memory. At block 504, during such monitoring, the traffic slot ring circuit detects a low-activity state of data commands over a specified previous time window. The time window is measured as a configurable number of memory clock cycles, for example 64 cycles. The monitoring may be performed for two sub-channels of a single memory controller, or, for example, across multiple memory controllers having traffic slot ring circuits linked as shown in Figure 3. Detecting a low-activity state may include detecting that no read or write commands are issued to the memory during a first period, or that some commands below a specified threshold are issued during the first period.
[0036] In response to detecting the low-activity state, the process at block 506 limits the number of read and write commands issued by the arbiter during a second predetermined period. The second period is also configurable as a number of memory clock cycles, for example 32. Note that read and write commands may be assigned different weights when monitored. In this embodiment, limiting the number of commands is achieved by setting a ramp-up limit imposed during the time window by preventing read and write commands (specifically, CAS commands) from being issued in the arbiter or in the memory interface queue for the arbiter.
[0037] In block 508, the memory controller then exits the low-activity state and begins to dispatch memory commands. For example, this could be because the memory controller exits the idle state or simply because command traffic to the memory controller ceases for a period of time and then traffic resumes. As shown in block 510, the throttle circuit in each arbiter imposes a ramp-up limit to ensure that the arbiter does not immediately increase traffic to 100% of the available bandwidth. The ramp-up limit is imposed over a second window or specified period of time and is imposed on each arbiter involved in the process. When a window monitor circuit (e.g., 448A, 448B of FIG. 4, 602 of FIG. 6) determines that the ramp-up limit has been reached for the current window or time period, the traffic slot-ringing circuit in each arbiter slot-rings the traffic for that respective arbiter. For example, if the process is used with the memory controller 410 of FIG. 4, the two sub-channel arbiters collectively have a ramp-up limit. As another example, when the process is used with quadrant-level slot-ringing enabled for the group of memory controllers of FIG. 3, all three memory controllers 310, 320, 330 collectively have a ramp-up limit.
[0038] In block 512, after the second specified period, the throttle circuit increases the ramp-up limit for each arbiter. Then, blocks 510 and 512 are repeated until the limit reaches 100% of the available bandwidth.
[0039] In the case of sub-channel slotting, the process can be configured using configuration settings set in registers within the memory controller. Register settings are provided to enable sub-channel slotting. The register settings are provided to specify the number of clock cycles during which the process does not send CAS to both sub-channels simultaneously when a throttle condition is met, and to specify the size of the time window to be monitored. Another register setting is provided to specify the maximum number of clocks for which both channels are idle before starting slotting between sub-channels. Another register setting provides a bandwidth increase limit for each monitored time window and sets a "ramp rate" for each time window during slotting, for example, as a percentage of the total bandwidth.
[0040] Generally, the ramp-up limit sets an acceptable increase in read commands and write commands over a second predetermined period and is based on a first value held in a programmable control register. It is preferable that the length of the first period and the length of the second period are based on a second value held in a programmable control register indicating the number of clock cycles.
[0041] FIG. 6 is a block diagram of a traffic slotting circuit 600 according to some embodiments. The traffic slotting circuit 600 is suitable for use in the memory system of FIG. 3, the memory controller of FIG. 4, or other suitable systems that use multiple memory channels or sub-channels.
[0042] The traffic slot ring circuit 600 in this example is embodied within a memory controller designated as "UMC0", which is part of a group of memory controllers that includes two other memory controllers, "UMC1" and "UMC2". The traffic slot ring circuit 600 has a first input labeled "CASSENT UMC1", a second input labeled "CASSENT UMC2", and an output labeled "CASSENT UMC0", and includes a window monitoring circuit 602, a synchronization counter 604, a ramp counter 606, and slot ring logic that implements a slot ring process. Also, the traffic slot ring circuit 600 may include a synchronization input (not shown) for adjusting the synchronization counter 604 with the synchronization counters of other memory controllers.
[0043] The CASSENT UMC0 output is used to indicate that a column address strobe (CAS) has been dispatched by UMC0 and is supplied, for example, to the traffic slot ring circuits of UMC1 and UMC2 as shown in FIG. 3. CASSENT UMC1 and CASSENT UMC2 receive similar signals from UMC1 and UMC2. The window monitoring circuit 602 measures the number of read and write commands picked up by the arbiter of UMC0 and the additional arbiters of UMC1 and UMC2 over a first predetermined period. In this embodiment, the measurement is performed by counting, during the first period, the CAS commands dispatched at UMC0 and the CASSENT indicators received from UMC1 and UMC2.
[0044] The synchronous counter 604 generally operates as a "toggle" circuit that imposes a ramp-up limit by adjusting among a plurality of memory controllers. In this embodiment, the synchronous counter 604 cycles through counts of 0, 1, and 2 in each command cycle of the memory controller (8 clock cycles per command cycle), and directly indicates which of the memory controllers UMC0, UMC1, and UMC2 is permitted to send read and write commands in each command cycle under specific conditions, as will be further described below. The synchronous counter can receive a synchronization signal and ensure synchronization with the synchronous counters in UMC1 and UMC2.
[0045] The ramp counter 606 is used to track the ramp-up limit in this example. The ramp counter 606 is started when one of UMC1 or UMC2 sends a CAS command, as indicated by the rising edge of the CASSENT signal received from UMC1 or UMC2. In this embodiment, the ramp counter 606 counts down from a specified value and is used in two different modes to delay the transmission of the CAS command in UMC0, as will be further described with respect to FIG. 7.
[0046] In some embodiments, this counter and signaling configuration is used, but other embodiments, of course, can achieve similar functions using various counter and timer configurations.
[0047] FIG. 7 is a flowchart 700 of another process for throttling traffic according to some embodiments. The process is suitable for use with the memory system of FIG. 3 and the traffic throttling circuit of FIG. 6, and generally provides one exemplary method of imposing a ramp-up limit after the memory controller exits the low activity state (e.g., block 510 of FIG. 5). The process is executed at each memory controller within a group or quadrant of memory controllers. Although the process steps are shown in order, this order is not limiting, and the steps are generally executed in parallel by digital logic within the traffic throttling circuit 600.
[0048] The process begins at block 702, where it is determined whether quadrant-level throttling is enabled, which in this implementation is determined by a bit in the configuration register. Quadrant-level throttling provides one way to coordinate the ramp-up from the idle state among multiple memory controllers such as memory controllers 310, 320, 330, etc. of FIG. 3.
[0049] If quadrant-level throttling is used, process block 704 blocks the CAS command selected by the arbiter as the CASSENT signals of two memory controllers (UMCs) are asserted and the ramp counter is non-zero. As described above, the ramp counter 606 in this implementation counts down to 0 to provide a delay period for the throttling command.
[0050] At block 706, if only one CASSENT signal is asserted from another memory controller, the process allows the CAS to be sent from the current memory controller if the synchronous counter 604 indicates the current memory controller.
[0051] If quadrant-level slotting is not enabled, the processes of blocks 708 - 712 are used. In block 708, the process blocks the CAS commands being sent from the current memory controller if any other memory controller within the group asserts the CASSENT signal and the lamp counter 606 is non-zero. In block 710, once the lamp counter 606 reaches 0, the CAS is permitted to be sent from the current memory controller if the synchronous counter 604 indicates the current memory controller. In block 712, the process ends the slotting and resumes the normal operation of the current memory controller if two CASSENT signals are asserted simultaneously from other memory controllers.
[0052] This process uses synchronous and lamp counters, although other processes can of course use other means to limit the number of read and write commands issued by the arbiter.
[0053] Register settings are provided to configure quadrant-level slotting. To enable quadrant-level slotting, register settings are provided. The register settings are provided to configure the detection of idle states that require slotting and include the number of inactive clock cycles to account for the active idle state. The lamp rate setting provides the maximum increase for each monitored window period during quadrant-level slotting. The throttle length setting provides the number of clock cycles to delay the next CAS if the lamp rate is reached during the time window.
[0054] The memory system of FIG. 3, the dual-channel memory controller of FIG. 2, and the traffic slot ring circuit of FIG. 6, or any portion thereof, can be described or represented by a computer-accessible data structure in the form of a database or other data structure that can be read by a program and used, directly or indirectly, in manufacturing an integrated circuit. For example, this data structure can be a behavioral-level description or a register-transfer level (RTL) description of the hardware functionality in a high-level design language (HDL) such as Verilog or VHDL. The description can be read by a synthesis tool that can synthesize the description to generate a netlist including a list of gates from a synthesis library. The netlist includes a set of gates that also represent the functionality of the hardware including the integrated circuit. The netlist can then be placed and routed to generate a data set that describes the geometric shapes to be applied to a mask. The mask can then be used in various semiconductor manufacturing processes to fabricate the integrated circuit. Alternatively, the database on the computer-accessible storage medium can, if desired, be a netlist (with or without a synthesis library) or a data set, or a graphic data system (GDS) II data.
[0055] While specific embodiments have been described, various modifications to these embodiments will be apparent to those skilled in the art. For example, although a dual-channel memory controller has been used as an example, the techniques herein may be applied to three or more memory channels in order to combine their capacities in a manner transparent to the data fabric and the host data processing system. For example, three or four memory channels may be controlled using the techniques herein by providing individual command queues and memory channel control circuits for each channel, while providing a single interface, an address decoder, and a credit control circuit that issues request credits independent of the individual memory channels to the data fabric. Further, the internal architecture of the dual-channel memory controller 210 may vary in different embodiments. The dual-channel memory controller 210 may interface with other types of memory other than DDRx, such as high-bandwidth memory (HBM), RAMbus DRAM (RDRAM), etc. The illustrated embodiments show each rank of memory corresponding to an individual DIMM or SIMM, but in other embodiments, each module may support multiple ranks. Still other embodiments may include other types of DRAM modules or DRAMs not included in a particular module, such as DRAM attached to the host motherboard. Accordingly, the appended claims are intended to cover all modifications of the disclosed embodiments that fall within the scope of the disclosed embodiments.
Claims
1. A memory controller operable to select commands for dispatching to a memory, comprising: an arbiter; a traffic throttle circuit connected to the arbiter to cooperate with one or more additional arbiters to mitigate an increase in excessive power consumption; wherein the traffic throttle circuit comprises a monitoring circuit for measuring the number of read commands and write commands selected by the arbiter and the one or more additional arbiters over a first predetermined period; and a throttle circuit for limiting the number of read commands and write commands issued during a second predetermined period in response to a low activity state. A memory controller.
2. The throttle circuit comprises a programmable control register for holding a limit value indicating an acceptable increase amount of read commands and write commands over the second predetermined period. The memory controller according to Claim 1.
3. The programmable control register holds values indicating the first predetermined period and the second predetermined period as the number of clock cycles. The memory controller according to Claim 2.
4. The monitoring circuit includes one or more first inputs for receiving signals from the one or more additional arbiters respectively, the signals indicating that each of the additional arbiters has transmitted a column address strobe command associated with a read command or a write command to the memory. The memory controller according to Claim 1.
5. The monitoring circuit includes one or more second inputs for receiving second signals from the one or more additional arbiters respectively, the second signals indicating that each of the additional arbiters has selected a read command or a write command to be transmitted to the memory. The memory controller according to Claim 4.
6. At least one of the one or more additional arbiters is a sub-channel arbiter for dispatching commands to a sub-channel of a memory channel at least partially shared with the arbiter. The memory controller according to Claim 1.
7. At least one of the one or more additional arbiters is a channel arbiter for dispatching commands to a channel different from the channel used by the arbiter. The memory controller of claim 1.
8. The throttle circuit includes a toggle circuit operable to indicate to the arbiter whether the arbiter is permitted to send read and write commands following the low activity state. The toggle circuit is synchronized with a similar toggle circuit in the one or more additional arbiters. The memory controller of claim 1.
9. A method for mitigating an excessive increase in power consumption in an arbiter in cooperation with one or more additional arbiters, measuring the number of read and write commands selected by the arbiter and the one or more additional arbiters over a first predetermined period, limiting the number of read and write commands issued by the arbiter for transmission to the memory during a second predetermined period in response to a low activity state associated with the read and write commands. Method.
10. Determining an acceptable increase amount of read and write commands over the second predetermined period based on a first value held in a programmable control register. The method of claim 9.
11. Determining the first period and the second period based on a second value held in the programmable control register indicating the number of clock cycles. The method of claim 10.
12. Receiving a signal from each of the one or more additional arbiters, the signal indicating that each of the additional arbiters has sent a column address strobe command associated with a read or write command to the memory. The method of claim 9.
13. Receiving a second signal from each of the one or more additional arbiters, the second signal indicating that each of the additional arbiters has selected a read or write command to be sent to the memory. The method of claim 9.
14. At least one of the one or more additional arbiters is a sub-channel arbiter for dispatching commands to a sub-channel of a memory channel at least partially shared with the arbiter. The method of claim 9.
15. At least one of the one or more additional arbiters is a channel arbiter for dispatching commands to a channel different from the channel used by the arbiter. The method of claim 9.
16. Based on the value of the toggle circuit, indicating to the arbiter whether the arbiter is permitted to send read commands and write commands following the low activity state, synchronizing the toggle circuit with a similar toggle circuit in the one or more additional arbiters. The method of claim 9.
17. A processing system, a data processor, a plurality of memory channels coupling the data processor to a memory, for any one of the plurality of memory channels, a memory controller operable to select a command for dispatching to the memory, the memory controller an arbiter, for each of the plurality of memory channels, a traffic throttle circuit in cooperation with one or more additional arbiters for mitigating an excessive increase in power usage, the traffic throttle circuit a monitoring circuit for measuring the number of read commands and write commands selected by the arbiter and the one or more additional arbiters over a first predetermined period, a throttle circuit for limiting the number of read commands and write commands issued during a second predetermined period in response to a low activity state. A processing system.
18. The throttle circuit includes a programmable control register that holds a limit value indicating an acceptable increase amount of read commands and write commands over the second predetermined period. The processing system of claim 17.
19. The programmable control register holds a value indicating the first predetermined period and the second predetermined period as the number of clock cycles. The processing system of claim 18.
20. The monitoring circuit includes one or more first inputs for receiving signals from the one or more additional arbiters, the signals indicating that each of the additional arbiters has sent a column address strobe command associated with a read command or a write command to the memory. The processing system of claim 17.
Citation Information
Patent Citations
Multi-ported memory controller with ports associated with traffic classes
JP2012074042A
Command arbitration for high-speed memory interfaces
JP2019525271A
Credit Scheme for Multi-Queue Memory Controllers - Patent application
JP2024512623A
Proportional memory operation throttling
US20130054901A1