Channel and subchannel throttling for memory controllers

By introducing monitoring and throttling circuits into the memory controller to monitor the number of read and write commands and limit the command rate, the problem of current peaks caused by sudden changes in memory channel traffic is solved, and more stable power management is achieved.

CN119452336BActive Publication Date: 2026-03-24ADVANCED MICRO DEVICES INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-22
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

In a multi-memory controller system, sudden surges in traffic on memory channels can cause current spikes, resulting in unnecessary power consumption and voltage drops. These spikes can occur over a relatively short period of time, especially when multiple memory controllers are operating simultaneously.

Method used

A combination of monitoring and throttling circuits is used. The monitoring circuit measures the number of read and write commands selected by the arbitrator within a predetermined time period, while the throttling circuit limits the number of read and write commands in a low-activity state. The rate at which commands are issued is controlled by ramp-up limiting technology to avoid current peaks.

Benefits of technology

It effectively reduces excessive power consumption, stabilizes voltage, avoids current peaks, and improves the power management efficiency of the memory system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119452336B_ABST
    Figure CN119452336B_ABST
Patent Text Reader

Abstract

An arbiter is operable to pick commands from a command queue to dispatch to a memory. The arbiter includes a traffic throttling circuit to cooperate with one or more additional arbiters to mitigate excess power usage increases. The traffic throttling circuit includes a monitoring circuit and a throttling circuit. The monitoring circuit is to measure a number of read and write commands picked by the arbiter and the one or more additional arbiters within a first predetermined time period. The throttling circuit is to limit a number of read and write commands issued by the arbiter during a second predetermined time period in response to a low activity state.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] Computer systems typically use inexpensive and high-density dynamic random access memory (DRAM) chips as main memory. Most DRAM chips sold today are compatible with various double data rate (DDR) DRAM standards published by the Joint Electron Device Engineering Council (JEDEC). DDR DRAM uses a conventional DRAM memory cell array with high-speed access circuitry to achieve high transfer rates and improve utilization of the memory bus. DDR memory controllers can interface with multiple DDR channels in order to accommodate more DRAM modules and exchange data with memory faster than using a single channel. Furthermore, modern server systems typically include multiple memory controllers in a single data processor. For example, some modern server processors include eight or twelve memory controllers, each connected to a respective DDR channel.

[0002] Sometimes traffic on a memory channel is throttled or slowed down to conserve power. Such throttling occurs over a relatively long period of time relative to the memory clock speed and is implemented, for example, by adjusting the memory clock and / or the memory controller clock. Traffic on a memory channel is often "bursty," that is, it can have periods of idleness or low traffic followed by periods of high traffic. A memory controller can also be placed into a low power mode in which the memory controller does not issue commands to the memory and then begins issuing commands when the low power mode ends. This sudden increase in traffic can cause harmful peaks in the current consumed by the memory controller and memory channel circuitry, especially when multiple memory controllers are involved. These peaks can occur over a relatively short period of time compared to the memory clock speed. BRIEF DESCRIPTION OF DRAWINGS

[0003] Figure 1 A data processing system 100 according to the prior art is shown in block diagram form;

[0004] Figure 2 An APU suitable for use in the data processing system 100 of Figure 1 is shown in block diagram form.

[0005] Figure 3 A memory system including multiple memory controllers according to some embodiments is shown in block diagram form;

[0006] Figure 4 A portion of a data processing system including a dual-channel memory controller suitable for use in an APU similar to Figure 2 is shown in block diagram form.

[0007] Figure 5This is a flowchart of another process for throttling traffic at the memory controller, according to some implementation schemes;

[0008] Figure 6 A flow throttling circuit according to some embodiments is shown in block diagram form; and

[0009] Figure 7 This is a flowchart of another process for throttling traffic, based on some implementation schemes.

[0010] In the following description, the same reference numerals are used in different figures to indicate similar or identical items. Unless otherwise stated, the word “coupled” and its associated verb form include both direct connection and indirect electrical connection by means known in the art, and unless otherwise stated, any description of direct connection also implies alternative embodiments using suitable forms of indirect electrical connection. Detailed Implementation

[0011] The memory controller is operable to select commands from a command queue for dispatch to memory. The memory controller includes an arbitrator and a flow throttling circuit, which works in conjunction with one or more additional arbitrators to mitigate increased excessive power usage. The flow throttling circuit includes monitoring circuitry and a throttling circuit. The monitoring circuitry measures the number of read and write commands selected by the arbitrator and one or more additional arbitrators within a first predetermined time period. The throttling circuitry, in response to low activity associated with read and write commands, limits the number of read and write commands issued by the arbitrator during a second predetermined time period.

[0012] A method is provided to work with one or more additional arbitrators to mitigate increased excessive power usage at the arbitrators. The method includes measuring the number of read and write commands selected by the arbitrator and one or more additional arbitrators during a first predetermined time period. The method further includes limiting the number of read and write commands issued by the arbitrator during a second predetermined time period in response to a low activity state associated with the read and write commands.

[0013] A processing system includes a data processor and a plurality of memory channels connecting the data processor to memory. A memory controller is operable to select commands from a command queue for dispatch to memory for one of the memory channels. The memory controller includes an arbitrator and a flow throttling circuit for cooperating with one or more additional arbitrators for a respective memory channel among the memory channels. The flow throttling circuit includes a monitoring circuit and a throttling circuit. The monitoring circuit measures the number of read and write commands selected by the arbitrator and the one or more additional arbitrators during a first predetermined time period. The throttling circuit, connected to the monitoring circuit and the arbitrator, limits the number of read and write commands issued by the arbitrator during a second predetermined time period in response to a low activity state associated with the read and write commands.

[0014] Figure 1 A data processing system 100 according to the prior art is illustrated in block diagram form. The data processing system 100 typically includes a data processor 110 in the form of an accelerated processing unit (APU), a memory system 120, a peripheral component interconnect high-speed (PCIe) system 150, a universal serial bus (USB) system 160, and a disk drive 170. The data processor 110 operates as the central processing unit (CPU) of the data processing system 100 and provides a variety of buses and interfaces available in modern computer systems. These interfaces include two Double Data Rate (DDRx) memory channels, a PCIe root complex for connecting to a PCIe link, a USB controller for connecting to a USB network, and an interface to Serial Advanced Technology Attachment (SATA) mass storage devices.

[0015] Memory system 120 includes memory channel 130 and memory channel 140. Memory channel 130 includes a set of dual in-line memory modules (DIMMs) connected to DDRx bus 132, including representative DIMMs 134, 136, and 138, which in this example correspond to individual memory columns. Similarly, memory channel 140 includes a set of DIMMs connected to DDRx bus 142, including representative DIMMs 144, 146, and 148.

[0016] PCIe system 150 includes PCIe switch 152 (which is connected to the PCIe root complex in data processor 110), PCIe devices 154, 156, and 158. PCIe device 156 is then connected to system basic input / output system (BIOS) memory 157. System BIOS memory 157 can be any of various non-volatile memory types, such as read-only memory (ROM), electrically erasable programmable flash ROM (EEPROM), etc.

[0017] USB system 160 includes a USB hub 162 connected to a USB host in data processor 110, and representative USB devices 164, 166, and 168 each connected to the USB hub 162. USB devices 164, 166, and 168 may be devices such as keyboards, mice, flash EEPROM ports, etc.

[0018] Disk drive 170 is connected to data processor 110 via SATA bus and provides large-capacity storage for operating system, applications, application files, etc.

[0019] By providing memory channels 130 and 140, the data processing system 100 is suitable for use in modern computing applications. Each memory channel in memory channels 130 and 140 can be connected to state-of-the-art DDR memories (such as DDR4, LPDDR4, GDDR5) and high-bandwidth memory (HBM), and is adaptable to future memory technologies. These memories provide high bus bandwidth and high-speed operation. Simultaneously, they also provide low-power modes to save power for battery-powered applications such as laptops, and also provide built-in thermal monitoring.

[0020] Figure 2 The diagram illustrates the appropriate approach for use in conjunction with... Figure 1 The APU 200 used in a similar data processing system as System 100. The APU 200 typically includes a central processing unit (CPU) core complex 210, a graphics core 220, a set of display engines 230, a memory management hub 240, a data texture 250, a set of peripheral controllers 260, a set of peripheral bus controllers 270, a system management unit (SMU) 280, and a set of memory controllers 290.

[0021] CPU core complex 210 includes CPU core 212 and CPU core 214. In this example, CPU core complex 210 includes two CPU cores, but in other embodiments, CPU core complex 210 may include any number of CPU cores. Each core in CPU cores 212 and 214 is bidirectionally connected to a system management network (SMN) (which forms a control structure) and data structure 250, and is able to provide memory access requests to data texture 250. Each core in CPU cores 212 and 214 may be a monolithic core, or it may be a core complex of two or more monolithic cores sharing certain resources such as cache.

[0022] The graphics core 220 is a high-performance graphics processing unit (GPU) capable of performing graphics operations such as vertex processing, fragment processing, shading, and texture blending in a highly integrated and parallel manner. The graphics core 220 is bidirectionally connected to the SMN and data texture 250 and can provide memory access requests to the data structure 250. In this regard, the APU 200 can support a unified memory architecture in which the CPU core complex 210 and the graphics core 220 share the same memory space, or a memory architecture in which the CPU core complex 210 and the graphics core 220 share a portion of the memory space, while the graphics core 220 also uses a private graphics memory that the CPU core complex 210 cannot access.

[0023] Display engine 230 renders and rasterizes objects generated by graphics core 220 for display on a monitor. Graphics core 220 and display engine 230 are bidirectionally connected to a common memory management hub 240 for uniform translation into appropriate addresses in memory system 120, and memory management hub 240 is bidirectionally connected to data texture 250 for generating such memory accesses and receiving read data returned from memory system.

[0024] Data texture 250 includes crossbars for routing memory access requests and memory responses between any memory access agent and memory controller 290. The data texture also includes a system memory map defined by the BIOS for determining the destination of memory accesses based on system configuration, and buffers for each virtual connection.

[0025] Peripheral controller 260 includes a USB controller 262 and a SATA interface controller 264, each of which is bidirectionally connected to the system hub 266 and the SMN bus. These two controllers are merely examples of peripheral controllers that can be used with the APU 200.

[0026] Peripheral bus controller 270 includes a system controller or "southbridge" (SB) 272 and a PCIe controller 274, each of which is bidirectionally connected to input / output (I / O) hub 276 and the SMN bus. I / O hub 276 is also bidirectionally connected to system hub 266 and data texture 250. Thus, for example, the CPU core can program registers in USB controller 262, SATA interface controller 264, SB 272, or PCIe controller 274 via access routed through I / O hub 276 from data texture 250.

[0027] The SMU 280 is a local controller that controls the operation of resources on the APU 200 and synchronizes communication between these resources. The SMU 280 manages the power-on sequence of the various processors on the APU 200 and controls multiple off-chip devices via reset, enable, and other signals. The SMU 280 includes one or more clock sources (…). Figure 2 (Not shown in the diagram) such as phase-locked loops (PLLs) are used to provide clock signals for each component of the APU 200. The SMU 280 also manages the power of various processors and other functional blocks, and can receive measured power consumption values ​​from CPU cores 212 and 214 and graphics core 220 to determine appropriate power states.

[0028] In this specific implementation, a set of memory controllers 290 includes memory controllers 292, 294, 296, and 298. Each memory controller has an upstream bidirectional connection to the data texture 250 and a downstream connection to a memory channel for accessing memory such as DDR memory, as further described below.

[0029] Figure 3 A memory system 300, comprising multiple memory controllers, is illustrated in block diagram form according to some embodiments. The memory system 300 is typically embodied on an integrated circuit, such as an APU, GPU, or similar. Figure 2 The CPU system-on-a-chip (SoC) includes memory controllers 310, 320, and 330, and DDR PHY interfaces 340, 342, and 344 for interfacing with one or more external memory modules. The memory system 300 typically also includes a power distribution network (PDN) comprising a first power domain 302 labeled “VDD domain” (where circuitry is powered by voltage VDD) and a second power domain 304 labeled “VDDP domain” (where circuitry is powered by voltage VDDP). While three memory controllers are shown, in most implementations of modern server systems, the SoC may have more DDR channels with more memory controllers, such as twelve, sixteen, or more. In some implementations, such memory controllers are arranged in “quadrants” where a group of memory controllers coordinates throttling within that group. In this implementation, memory controllers 310, 320, and 330 constitute such a group.

[0030] Memory channel controllers 310, 320, and 330 are typically provided in the first power domain 302 and may be included in lower-level domains specific to each memory controller under the first power domain 302 in the PDN. Similarly, DDR PHY interfaces 340, 342, and 344 are typically provided in the second power domain 304 and may be included in lower-level domains specific to each DDR PHY interface under the second power domain 304 in the PDN. Typically, in operation, memory channel controllers 310, 320, and 330 handle various traffic loads exhibiting “burst” behavior, where the memory channel controller is idle and then quickly becomes busy fulfilling memory access requests. Such behavior tends to cause voltage drops in one or both of the VDDP domain 304 and VDD domain 302. Specifically, when all memory controllers in a quadrant send read (RD) and write (WR) commands at their full bandwidth capacity from idle approximately simultaneously, it often results in voltage drops on the order of nanoseconds (10 to 20 nanoseconds).

[0031] Each memory channel controller includes one or more corresponding arbitrators 312, 322, and 332, and corresponding flow throttling circuits 314, 324, and 334. Each flow throttling circuit is bidirectionally connected to the other flow throttling circuits for cooperation with the arbitrators of other memory controllers to mitigate increased excessive power usage. In this specific embodiment, each flow throttling circuit 314, 324, and 334 includes: a monitoring circuit for measuring the number of read and write commands selected by its corresponding arbitrator and the other arbitrators within a first predetermined time period; and a throttling circuit coupled to the monitoring circuit and the corresponding arbitrator. In response to a low activity state associated with read and write commands, each throttling circuit is used to limit the number of read and write commands issued by the arbitrators during a second predetermined time period, which may be repeated as further described below.

[0032] Figure 4 The diagram illustrates the components suitable for use in conjunction with... Figure 2This is a portion of the data processing system 400 using a dual-channel memory controller 410, similar to that used in APUs. The dual-channel memory controller 410 is shown connected to a data texture 250, to which it can communicate with multiple memory agents present in the data processing system 400. The dual-channel memory controller 410 can control, for example, two sub-channels defined in the DDR5 specification for use with DDR5 DRAMs, or those sub-channels defined in the High Bandwidth Memory 2 (HBM2) and HBM3 standards. In some implementations, instead of two sub-channels, the dual-channel memory controller 410 can replace two separate memory channel controllers and control two DDRx channels together in a manner transparent to the various memory address agents in the data texture 250 and APU 200, allowing a single memory controller interface 412 to be used for sending memory access commands and receiving results. The dual-channel memory controller 410 typically includes two instances of an interface 412, an address decoder 422, and memory channel control circuitry 423A and 423B, each assigned to a different memory channel. Each instance of the memory channel control circuits 423A and 423B includes memory interface queues 414A and 414B, command queues 420A and 420B, content addressable memory (CAM) 424A and 424B, timing blocks 434A and 434B, page tables 436A and 436B, arbitrators 438A and 438B, ECC generation blocks 444A and 444B, data buffers 446A and 446B, and flow throttling circuits 432A and 432B, the flow throttling circuits including window monitoring circuits 448A and 448B.

[0033] In other embodiments, for each memory channel or subchannel used, only command queues 430A, 430B, arbitrators 438A, 438B, flow throttling circuits 432A, 432B, and memory interface queues 414A, 414B are replicated; the remaining circuitry shown is adapted for use with both channels. Furthermore, while the depicted dual-channel memory controller 410 includes two instances of arbitrators 438A, 438B, command queues 420A, 420B, and memory interface queues 414A, 414B for controlling two memory channels or subchannels, other embodiments may include more instances, such as three, four, or more instances, for communicating with DRAM on three or four channels or subchannels according to the credit management techniques described herein. Moreover, while a dual-channel memory controller is used in this embodiment, other embodiments may apply the throttling techniques described herein to a single-channel memory controller.

[0034] Interface 412 has a first bidirectional connection to data texture 250 via a communication bus and a second bidirectional connection to credit control circuitry 421. In this embodiment, interface 412 uses a Scalable Data Port (SDP) link to establish several channels for communication with data texture 250, but other interface link standards are also applicable. For example, in another embodiment, the communication bus is compatible with Advanced Scalable Interface Version 4 (referred to as AXI4) as specified by ARM Holdings, PLC of Cambridge, England, but in other embodiments it may be other types of interfaces. Interface 412 translates memory access requests from a first clock domain called the “FCLK” (or “MEMCLK”) domain to a second clock domain called the “UCLK” domain within the dual-channel memory controller 410. Similarly, memory interface queue 414 provides memory access from the UCLK domain to the “DFICLK” domain associated with the DFI interface.

[0035] Address decoder 422 has a bidirectional link to interface 412, a first output connected to a first command queue 420A (labeled "Command Queue 0"), and a second output connected to a second command queue 420B (labeled "Command Queue 1"). Address decoder 422 decodes the address of a memory access request received on data texture 250 through interface 412. The memory access request includes an access address in the physical address space represented in a normalized format. Based on the access address, address decoder 422 selects a memory channel and an associated command queue in command queues 420A and 420B to process the request. The selected channel for each request is identified to credit control circuitry 421, enabling a credit issuance decision. Address decoder 422 converts the normalized address into a format that can be used to address the actual memory devices in memory system 130 and efficiently schedule related accesses. This format includes region identifiers that associate the memory access request with specific memory column, row address, column address, bank address, and bank group. At startup, the system BIOS queries the memory devices in memory system 130 to determine their size and configuration, and programs a set of configuration registers associated with address decoder 422. Address decoder 422 uses the configuration stored in the configuration registers to translate normalized addresses into appropriate formats. For each memory access request selected by address decoder 422, each memory access request is loaded into command queue 420A or 420B.

[0036] Each command queue 420A, 420B is a queue of memory access requests received from various memory access engines in the APU 200, such as CPU cores 212 and 214 and graphics core 220. Each command queue 420A, 420B is bidirectionally connected to a corresponding arbiter 438A, 438B for selecting memory access requests to be issued through the associated memory channel from the command queues 420A, 420B. Each command queue 420A, 420B stores the address field decoded by the address decoder 422 and other address information that allows the corresponding arbiters 438A, 438B to efficiently select memory accesses, including access type and quality of service (QoS) identifiers. Each CAM 424A, 424B includes information on implementing ordering rules such as write-after-write (WAW) and read-after-write (RAW) ordering rules.

[0037] Arbitrators 438A and 438B are bidirectionally connected to their respective command queues 420A and 420B to select memory access requests to be completed with appropriate commands, and bidirectionally connected to their respective flow throttling circuits 432A and 432B to receive throttling signals. Arbitrators 438A and 438B typically improve the utilization of the memory bus of their respective memory channels through intelligent access scheduling to improve the efficiency of those memory channels. Each arbitrator 438A and 438B uses its respective timing blocks 434A and 434B to implement correct timing relationships by determining, based on DRAM timing parameters, whether certain accesses in their respective command queues 420A and 420B are eligible to be issued. Each page table 436A and 436B maintains status information about the active pages in each bank and column of their respective memory channels for their respective arbitrators 438A and 438B, and is bidirectionally connected to their respective replay queues 430A and 430B. Each arbiter 438A, 438B uses decoded address information, timing qualification information indicated by timing blocks 434A, 434B, and active page information indicated by page tables 436A, 436B to efficiently schedule memory accesses while adhering to other standards such as Quality of Service (QoS) requirements. In some implementations, arbiters 438A, 438B determine whether to allow the release of the selected command based on signals from flow throttling circuits 432A, 432B.

[0038] Each flow throttling circuit 432A, 432B is bidirectionally connected to the corresponding arbitrators 438A, 438B, and bidirectionally connected to one or more other flow throttling circuits that can be in other memory controllers. As shown, flow throttling circuits 432A, 432B are bidirectionally connected to each other to provide a signal marked CASSENT, thereby indicating to each other when a column address strobe (CAS) command has been sent to the memory channel. See below for reference. Figure 5 to Figure 7Further described, the CASSENT signal is used in the throttling process. Each flow throttling circuit includes window monitoring circuits 448A, 448B, which measure the number of read and write commands selected by their respective arbitrators 438A, 438B and one or more additional arbitrators within a first predetermined time period. Typically, as further described below, each flow throttling circuit 432A, 432B, in response to a low activity state associated with read and write commands, limits the number of read and write commands issued by the arbitrators 438A, 438B during a second predetermined time period. Each flow throttling circuit 432A, 432B can also send an "idle" signal to indicate that an idle condition has been reached at the arbitrator that sent the flow throttling circuit, i.e., a specified length of time for which no CAS signal has been dispatched. In response to this idle signal, the flow throttling circuits 432A, 432B currently in throttling mode adjust their own throttling rate to allow more CAS commands to be sent.

[0039] Each Error Correction Code (ECC) generation block 444A, 444B determines the ECC of the write data to be sent to memory. In response to receiving a write memory access request from interface 412, ECC generation blocks 444A, 444B calculate the ECC based on the write data. Data buffers 446A, 446B store the write data and ECC of the received memory access request. When the corresponding arbitrators 438A, 438B select the corresponding write access to assign to the memory channel, data buffers 446A, 446B output the combined write data / ECC to the corresponding memory interface queues 414A, 414B.

[0040] Figure 5 This is a flowchart 500 of a process for throttling traffic at a memory controller, according to some implementation schemes. This process is suitable for use with... Figure 3 memory system, Figure 4 The memory controller may use multiple memory channels or sub-channels and include, for example, memory controllers. Figure 3 , Figure 4 or Figure 6 The flow throttling circuit shown can be used in conjunction with other suitable systems.

[0041] The process begins at box 502, where the flow throttling circuitry monitors multiple arbiters or sub-channel arbiters for multiple memory access commands (specifically read and write) selected for dispatch to memory. During this monitoring, at box 504, the flow throttling circuitry detects low activity of data commands over a specified prior time window. The time window is measured as a configurable number of clock cycles of the memory clock, such as 64 cycles. For example, monitoring can be performed for two sub-channels of a single memory controller, or across multiple sub-channels with different memory controllers. Figure 3The flow throttling circuit linked to the diagram is monitored by multiple memory controllers. Detecting a low-activity state may include detecting that no read or write commands were issued to the memory during a first time period, or that the number of commands issued during the first time period is less than a specified threshold.

[0042] The process at box 506, in response to the detection of a low-activity state, limits the number of read and write commands issued by the arbitrator during a second predetermined time period. This second time period can also be configured as the number of memory clock cycles, for example, 32. It should be noted that read and write commands can be assigned different weights during monitoring. In this specific implementation, limiting the number of commands is achieved by setting a ramp-up limit, which is implemented during the time window by preventing read and write commands (specifically, CAS commands) from being issued at the arbitrator or the arbitrator's memory interface queue.

[0043] At box 508, the memory controller then leaves the low-activity state and begins dispatching memory commands. This could be because the memory controller leaves the idle state, or simply because there was no command traffic to the memory controller for a period of time, and then traffic appears. As shown in box 510, the throttling circuitry at each arbiter implements ramp-up limiting to ensure that the arbiter does not immediately increase traffic to 100% of the available bandwidth. Ramp-up limiting is implemented in a second window or a specified time period, and at each arbiter involved in this process. When the window monitoring circuitry (e.g., Figure 4 448A and 448B in the middle, Figure 6 (602) When it is determined that the ramp-up limit has been reached for the current window or time period, the flow throttling circuit at each arbitrator throttles the flow to that corresponding arbitrator. For example, when the process and Figure 4 When used together with the memory controller 410, the two sub-channel arbiters share a ramp-up limit. As another example, when used for... Figure 3 When the process is used with quadrant-level throttling enabled in the memory controller group, all three memory controllers 310, 320 and 330 have ramp-up limits together.

[0044] At box 512, after a second specified time period, the throttling circuit increases the ramp-up limit at each arbiter. Boxes 510 and 512 are then repeated until the limit reaches 100% of the available bandwidth.

[0045] For subchannel throttling, the procedure can be configured using configuration settings set in the memory controller registers. Register settings are provided to enable subchannel throttling. One register setting specifies the number of clock cycles during which the procedure prevents two subchannels from sending CAS simultaneously when throttling conditions are met, thus specifying the size of the monitored time window. Another register setting specifies the maximum number of clock cycles that the two channels can remain idle before initiating throttling between subchannels. Yet another register setting provides a bandwidth increase limit for each monitored time window, setting a "ramp rate" for each time window during throttling, for example, as a percentage of the total bandwidth.

[0046] Typically, the ramp-up limit is set by the allowable increase in read and write commands within a second predetermined time period, and is based on a first value stored in a programmable control register. The lengths of the first and second time periods are preferably based on a second value stored in the programmable control register indicating the number of clock cycles.

[0047] Figure 6 A block diagram of a flow throttling circuit 600 according to some embodiments is shown. The flow throttling circuit 600 is suitable for use in... Figure 3 memory system, Figure 4 It is used in memory controllers or other suitable systems that use multiple memory channels or sub-channels.

[0048] In this example, the flow throttling circuit 600 is implemented in a memory controller designated "UMC0," which is part of a set of memory controllers including two other memory controllers, "UMC1" and "UMC2." The flow throttling circuit 600 has a first input labeled "CASSENT UMC1," a second input labeled "CASSENT UMC2," and an output labeled "CASSENT UMC0," and includes window monitoring circuitry 602, a synchronization counter 604, and a ramp counter 606, as well as throttling logic for implementing the throttling process. The flow throttling circuit 600 may also include a synchronization input (not shown) for coordinating the synchronization counter 604 with the synchronization counters of other memory controllers.

[0049] For example, such as Figure 3As shown, the CASSENT UMC0 output indicates that the column address strobe (CAS) has been dispatched by UMC0 and fed to the flow throttling circuits of UMC1 and UMC2. CASSENT UMC1 and CASSENT UMC2 receive similar signals from UMC1 and UMC2. Window monitoring circuit 602 measures the number of read and write commands selected by the arbitrator of UMC0 and the additional arbitrators UMC1 and UMC2 within a first predetermined time period. In this embodiment, the measurement is performed by counting the CAS commands dispatched at UMC0 during the first time period and the CASSENT indicators received from UMC1 and UMC2.

[0050] Synchronous counter 604 typically acts as a "switchback" circuit to implement ramp-up limiting by coordinating among multiple memory controllers. In this specific implementation, synchronous counter 604 cycles through 0, 1, and 2 in each command cycle of the memory controller (each command cycle is 8 clock cycles), and directly indicates which of the memory controllers UMC0, UMC1, and UMC2 is authorized to send read and write commands in each command cycle under certain conditions, as further described below. The synchronous counter can receive a synchronization signal to ensure that it is synchronized with the synchronous counters at UMC1 and UMC2.

[0051] In this specific implementation, ramp counter 606 is used to track ramp rise limits. Ramp counter 606 starts when either UMC1 or UMC2 sends a CAS command, as indicated by the rising edge of the CASSENT signal received from UMC1 or UMC2. In this specific implementation, ramp counter 606 counts down from a specified value and is used in two different modes to delay the transmission of the CAS command at UMC0, as referenced... Figure 7 Further description.

[0052] While this counter and signaling arrangement is used in some implementations, other implementations can certainly achieve similar functionality using various counter and timer arrangements.

[0053] Figure 7 This is flowchart 700 of another process for throttling flow according to some implementation schemes. This process is suitable for use with… Figure 3 memory system and Figure 6 They are used together with flow throttling circuits and are typically provided as an exemplary method for implementing ramp-up limits after the memory controller leaves a low-activity state (e.g., Figure 5(See box 510). The process is performed at each memory controller in the memory controller group or quadrant. Although the process steps are depicted sequentially, this order is not restrictive, and these steps are typically performed in parallel by digital logic in the flow throttling circuit 600.

[0054] The process begins at box 702, where it determines whether quadrant-level throttling is enabled, which in this specific implementation is determined by bits in the configuration register. Quadrant-level throttling provides... Figure 3 A method for coordinating the ramp-up from the idle state among multiple memory controllers of memory controllers 310, 320 and 330.

[0055] If quadrant-level throttling is used, process block 704 prevents the CAS command selected by the arbitrator if the CASSENT signals of both memory controllers (UMCs) are asserted and the ramp counter is non-zero. As described above, in this embodiment, ramp counter 606 decrements to zero, providing a delay period for the throttling command.

[0056] At block 706, if only one CASSENT signal from another memory controller is asserted, the procedure allows CAS to be sent from the current memory controller if the synchronization counter 604 indicates the current memory controller.

[0057] If quadrant-level throttling is not enabled, the procedures at blocks 708-712 are used. At block 708, if any other memory controller in the group has asserted the CASSENT signal and ramp counter 606 is non-zero, the procedure prevents CAS commands from being sent from the current memory controller. At block 710, once ramp counter 606 reaches zero, CAS commands are allowed from the current memory controller if synchronization counter 604 indicates to the current memory controller. At block 712, if two CASSENT signals from other memory controllers are asserted simultaneously, the procedure ends throttling and restores normal operation of the current memory controller.

[0058] Although this process uses a synchronization counter and a ramp counter, other processes can certainly use other schemes to limit the number of read and write commands issued by the arbitrator.

[0059] Register settings are provided to configure quadrant-level throttling. Register settings are provided to enable quadrant-level throttling. Register settings are provided to configure an idle condition that detects when throttling is required, containing the number of inactive clock cycles to account for when the idle condition is activated. During quadrant-level throttling, the ramp rate setting provides the maximum increment for each monitored window period. When the ramp rate is reached during the time window, the throttling length setting provides the number of clock cycles to delay the next CAS.

[0060] Figure 3 memory system, Figure 2 Dual-channel memory controller and Figure 6 The flow throttling circuit or any part thereof can be described or represented by a computer-accessible data structure in the form of a database or other data structure that can be read by a program and is used directly or indirectly for manufacturing integrated circuits. For example, the data structure can be a behavioral-level description or register-transfer-level (RTL) description of hardware functionality in a high-level design language (HDL) such as Verilog or VHDL. The description can be read by a synthesis tool that can synthesize the description to produce a netlist including a list of gates from a synthesis library. The netlist includes gate sets that also represent the functionality of the hardware comprising the integrated circuit. The netlist can then be placed and routed to produce a dataset describing the geometry to be applied to a mask. The mask can then be used in various semiconductor manufacturing steps to produce the integrated circuit. Alternatively, the database on a computer-accessible storage medium can be a netlist (with or without a synthesis library), a dataset (as needed), or Graphical Data System (GDS) II data.

[0061] While specific embodiments have been described, various modifications to these embodiments will be apparent to those skilled in the art. For example, although a dual-channel memory controller is used as an example, the techniques described herein can also be applied to more than two memory channels to combine their capacity in a manner transparent to the data texture and host data processing system. For example, the techniques described herein can be used to control three or four memory channels by providing independent command queues and memory channel control circuitry for each channel, while providing a single interface, address decoder, and credit control circuitry that issues credit requests to the data texture, independent of each individual memory channel. Furthermore, the internal architecture of the dual-channel memory controller 210 can vary in different embodiments. The dual-channel memory controller 210 can interface to other types of memory besides DDRx, such as high-bandwidth memory (HBM), RAMbus DRAM (RDRAM), etc. While the illustrated embodiments show each memory storage column corresponding to a single DIMM or SIMM, in other embodiments, each module may support multiple storage columns. Still other embodiments may include other types of DRAM modules or DRAM not included in a particular module, such as DRAM mounted to a host motherboard. Therefore, the appended claims are intended to cover all modifications of the disclosed embodiments that fall within the scope of the disclosed embodiments.

Claims

1. A memory controller operable to select commands for dispatching to memory, the memory controller comprising: Arbitrator; A flow throttling circuit, connected to the arbitrator, for cooperating with one or more additional arbitrators to mitigate increased excessive power usage, the flow throttling circuit comprising: A monitoring circuit is used to measure the number of read and write commands selected by the arbitrator and the one or more additional arbitrators within a first predetermined time period; and A throttling circuit, in response to the number of commands sent to the memory during the first predetermined time period being less than a specified threshold, limits the number of read and write commands issued by the arbitrator during a second predetermined time period.

2. The memory controller according to claim 1, wherein: The throttling circuit further includes a programmable control register that stores a limit value indicating the allowable increase in read and write commands within the second predetermined time period.

3. The memory controller according to claim 2, wherein: The programmable control register further stores values ​​that indicate the first predetermined time period and the second predetermined time period as clock cycle numbers.

4. The memory controller according to claim 1, wherein: The monitoring circuit includes one or more first inputs, each of which receives a signal from one or more additional arbitrators to indicate that the respective additional arbitrator has sent a column address strobe command associated with a read or write command to the memory.

5. The memory controller according to claim 4, wherein: The monitoring circuit includes one or more second inputs, each receiving a second signal from one or more additional arbitrators to indicate that the respective additional arbitrator has selected a read or write command to be sent to the memory.

6. The memory controller according to claim 1, wherein: At least one of the one or more additional arbitrators is a subchannel arbitrator, which is used to dispatch commands to a subchannel of a memory channel that is at least partially shared with the arbitrator.

7. The memory controller according to claim 1, wherein: At least one of the one or more additional arbitrators is a channel arbitrator, which is used to dispatch commands to a channel different from the channel used by the arbitrator.

8. The memory controller according to claim 1, wherein: The throttling circuit includes a back-and-forth switching circuit that is operable to indicate to the arbitrator whether it is authorized to send read and write commands after the number of commands sent to the memory during the first predetermined time period falls below the specified threshold. The back-and-forth switching circuit is synchronized with a similar back-and-forth switching circuit in one or more additional arbitrators.

9. A method for mitigating increased excessive power usage at an arbitrator in cooperation with one or more additional arbitrators, the method comprising: Measure the number of read and write commands selected by the arbitrator and the one or more additional arbitrators within a first predetermined time period; as well as In response to the number of read and write commands sent to the memory during the first predetermined time period being less than a specified threshold, the number of read and write commands issued by the arbitrator for transmission to the memory during the second predetermined time period is limited.

10. The method according to claim 9, wherein the method further comprises: The permissible increase in read and write commands during the second predetermined time period is determined based on a first value stored in the programmable control register.

11. The method according to claim 10, further comprising: The first predetermined time period and the second predetermined time period are determined based on a second value indicating the number of clock cycles stored in the programmable control register.

12. The method according to claim 9, wherein the method further comprises: Signals are received from the one or more additional arbitrators to indicate that the respective additional arbitrator has sent a column address strobe command associated with a read or write command to the memory.

13. The method according to claim 12, wherein the method further comprises: A second signal is received from each of the one or more additional arbitrators to indicate that the respective additional arbitrator has selected a read or write command to be sent to the memory.

14. The method according to claim 9, wherein: At least one of the one or more additional arbitrators is a subchannel arbitrator, which is used to dispatch commands to a subchannel on a memory channel that is at least partially shared with the arbitrator.

15. The method according to claim 9, wherein: At least one of the one or more additional arbitrators is a channel arbitrator, which is used to dispatch commands to a channel different from the channel used by the arbitrator.

Citation Information

Patent Citations

  • Multi-ported memory controller with ports associated with traffic classes

    JP2012074042A

  • Command arbitration for high-speed memory interfaces

    JP2019525271A