Memory context restoration and boot time reduction for a system on chip through reduced double data rate memory training

CN114868111BActive Publication Date: 2026-09-22ADVANCED MICRO DEVICES INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080090909.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-30
Filing Date
2020-11-25
Publication Date
2026-09-22
Estimated Expiration
2040-11-25

Smart Images

  • Figure CN114868111B_ABST
    Figure CN114868111B_ABST
Patent Text Reader

Abstract

Methods for reducing boot time of a system on a chip (SOC) by reducing double data rate (DDR) memory training and for memory context recovery. Dynamic random access memory (DRAM) controller and DDR physical interface (PHY) settings are stored into non-volatile memory and the DRAM controller and the DDR PHY are powered down. Upon system recovery, the basic input / output system recovers the DRAM controller and the DDR PHY settings from non-volatile memory and finalizes the DRAM controller and the DDR PHY settings to operate with the SOC. Reducing the boot time of the SOC by reducing DDR training includes setting a DRAM to a self-refresh mode and programming a self-refresh state machine memory operation (MOP) array to exit the self-refresh mode and update any DRAM device states for a target power management state. The DRAM device is reset and the self-refresh state machine MOP array reinitializes the DRAM device states for the target power management state.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims the benefit of U.S. nonprovisional patent application number 16 / 730,086, filed December 30, 2019, the contents of which are incorporated herein by reference. Background Technology

[0003] Dynamic Random Access Memory (DRAM) is a commonly used type of memory in computer systems. DRAM is volatile memory, requiring proper initialization and periodic calibration to maintain performance. Attached Figure Description

[0004] A more detailed understanding can be obtained from the following description, given by way of example in conjunction with the accompanying drawings:

[0005] Figure 1 This is a block diagram of an example apparatus that can implement one or more features of this disclosure;

[0006] Figure 2 yes Figure 1 A block diagram of the device, showing further details;

[0007] Figure 3 A data processing system according to some implementation schemes is illustrated in block diagram form;

[0008] Figure 4 The diagram illustrates the applicable... Figure 3 The acceleration processing unit (APU) of the data processing system;

[0009] Figure 5 The diagram illustrates, in block form, an application of some implementation schemes. Figure 4 The memory controller and associated physical interface (PHY) of the APU;

[0010] Figure 6 The diagram illustrates, in block form, an application of some implementation schemes. Figure 4 Another memory controller and associated PHY of the APU;

[0011] Figure 7 A memory controller according to some embodiments is shown in block diagram form;

[0012] Figure 8 The diagram illustrates the corresponding implementation schemes according to some embodiments. Figure 3 The data processing system is part of the data processing system;

[0013] Figure 9 The diagram illustrates the corresponding implementation schemes according to some embodiments. Figure 7The memory channel controller portion of the memory channel controller; and

[0014] Figure 10 This paper demonstrates a method for reducing system-on-chip (SoC) boot time by reducing double data rate (DDR) training; and

[0015] Figure 11 A method for memory context recovery by reducing DDR training is shown. Detailed Implementation

[0016] This teaching content provides a method to reduce system boot time by canceling or reducing the Double Data Rate (DDR) training step during subsequent reboots. A hardware-based mechanism is used to quickly reinitialize the Dynamic Random Access Memory (DRAM) device using settings from the previous boot. This technique allows for flexible use of Dual In-line Memory Modules (DIMMs), which can be replaced at the factory or by the end customer in the field, while still maintaining fast subsequent boots to improve user experience. For example, for advanced DDR training steps, it may take 1 to 2 seconds during the first boot to optimize the timing, voltage, and decision feedback equalizer (DFE) / feedforward equalizer (FFE) of the DDR channels. However, for a given processor / platform (motherboard / module / DRAM) combination, once these values ​​are known, subsequent training can be canceled (e.g., for DDR4, or reduced in the case of LPDDR4 systems). This allows the system to skip or reduce DDR training steps, including loading training firmware code and running multiple verbose training firmware steps.

[0017] This tutorial utilizes an initialization process that saves and restores DRAM controller and DDR physical interface (PHY) configuration settings from non-volatile memory. During system recovery, DRAM contents and settings can be retained, or the Basic Input / Output System (BIOS) can optionally reset and reinitialize the DRAM and its contents.

[0018] Methods for reducing SoC boot time by reducing DDR training and for memory context recovery by reducing DDR training are disclosed. These methods include storing DRAM controller and DDR PHY settings in non-volatile memory and powering down the DRAM controller and DDR PHY. Upon system recovery, the BIOS restores the DRAM controller and DDR PHY settings from the non-volatile memory and finalizes the DRAM controller and DDR PHY settings for performing task-mode operation with the SoC. Methods for reducing SoC boot time by reducing DDR training also include setting the DRAM to self-refresh mode and programming the self-refresh state machine memory operation (MOP) array to exit self-refresh and update any DRAM device states for a target power management state. Methods for memory context recovery by reducing DDR training also include resetting the DRAM devices and programming the self-refresh state machine MOP array to reinitialize the DRAM device states for a target power management state.

[0019] While this disclosure includes a discussion of DRAM memory and DRAM controller as specific embodiments, those skilled in the art will recognize that other types of memory may be utilized in the present embodiments. Therefore, DRAM includes any form of memory, and these memory types are alternatives to the DRAM discussed herein. A DRAM controller is understood to be a memory controller that controls the corresponding memory in use, even though the examples herein pertain to a DRAM controller.

[0020] Figure 1 This is a block diagram of an example device 100 that can implement one or more features of the present disclosure. Device 100 may include, for example, a computer, a gaming device, a handheld device, a set-top box, a television, a mobile phone, or a tablet computer. Device 100 includes a processor 102, a memory 104, a storage device 106, one or more input devices 108, and one or more output devices 110. Device 100 may also optionally include an input driver 112 and an output driver 114. It should be understood that device 100 may include... Figure 1 Additional components not shown.

[0021] In various alternatives, processor 102 includes a central processing unit (CPU), a graphics processing unit (GPU), a CPU and GPU located on the same die, or one or more processor cores, where each processor core may be a CPU or a GPU. In various alternatives, memory 104 is located on the same die as processor 102 or is located separately from processor 102. Memory 104 includes volatile or non-volatile memory, such as random access memory (RAM), dynamic RAM, or cache.

[0022] Storage device 106 includes fixed or removable storage devices, such as hard disk drives, solid-state drives, optical disk drives, or flash drives. Input device 108 includes, but is not limited to: keyboards, keypads, touchscreens, touchpads, detectors, microphones, accelerometers, gyroscopes, biometric scanners, or network connections (e.g., wireless LAN cards for transmitting and / or receiving wireless IEEE 802 signals). Output device 110 includes, but is not limited to: displays, speakers, printers, haptic feedback devices, one or more lights, antennas, or network connections (e.g., wireless LAN cards for transmitting and / or receiving wireless IEEE 802 signals).

[0023] Input driver 112 communicates with processor 102 and input device 108, and permits processor 102 to receive input from input device 108. Output driver 114 communicates with processor 102 and output device 110, and permits processor 102 to send output to output device 110. It should be noted that input driver 112 and output driver 114 are optional components, and device 100 will operate in the same manner in the absence of input driver 112 and output driver 114. Output driver 114 includes an accelerated processing unit (“APD”) 116 coupled to display device 118. The APD receives computation commands and graphics rendering commands from processor 102, processes those commands, and provides pixel output to display device 118 for display. As described in further detail below, APD 116 includes one or more parallel processing units to perform computations according to the Single Instruction Multiple Data (“SIMD”) paradigm. Therefore, although various functions are described herein as being performed by or in combination with APD 116, in various alternatives, the functions described as being performed by APD 116 may additionally or alternatively be performed by other computing devices with similar capabilities, which are not driven by a host processor (e.g., processor 102) and provide graphics output to display device 118. For example, it is conceivable that any processing system performing processing tasks according to the SIMD paradigm can perform the functions described herein. Alternatively, it is conceivable that a computing system not performing processing tasks according to the SIMD paradigm can perform the functions described herein.

[0024] Figure 2This is a block diagram of device 100, illustrating additional details related to performing processing tasks on APD 116. Processor 102 maintains one or more control logic modules in system memory 104 for execution by processor 102. The control logic modules include operating system 120, kernel-mode driver 122, and application program 126. These control logic modules control various features of the operation of processor 102 and APD 116. For example, operating system 120 communicates directly with the hardware and provides an interface to the hardware for other software executing on processor 102. Kernel-mode driver 122 controls the operation of APD 116 by providing, for example, an application programming interface (API) to software executing on processor 102 (e.g., application program 126) to access various functions of APD 116. Kernel-mode driver 122 also includes a just-in-time (JIT) compiler that compiles programs for execution by processing units of APD 116, such as SIMD unit 138, which is discussed in further detail below.

[0025] APD 116 executes commands and procedures related to selected functions, such as graphics and non-graphics operations suitable for parallel processing. APD 116 can be used to perform graphics pipeline operations based on commands received from processor 102, such as pixel manipulation, geometric calculations, and rendering images to display device 118. APD 116 can also perform computational processing operations not directly related to graphics operations based on commands received from processor 102, such as operations related to video, physics simulations, computational fluid dynamics, or other tasks.

[0026] APD 116 includes a computation unit 132 comprising one or more SIMD units 138 that perform operations in parallel upon request from processor 102 according to a SIMD paradigm. A SIMD paradigm is one in which multiple processing elements share a single program control flow unit and program counter, and thus execute the same program, but can execute that program using different data. In one example, each SIMD unit 138 includes sixteen channels, where each channel executes the same instruction simultaneously with other channels in the SIMD unit 138, but can execute that instruction using different data. If not all channels need to execute a given instruction, an assertion can be used to shut down a channel. Assertions can also be used to execute programs with divergent control flow. More specifically, for programs with conditional branches or other instructions, where control flow is based on computation performed by a single channel, assertions corresponding to channels with currently unexecuted control flow paths and the serial execution of different control flow paths allow for arbitrary control flow.

[0027] The basic unit of execution in computing unit 132 is a work item. Each work item represents a single instantiation of a program that will be executed in parallel on a specific channel. Work items can be executed simultaneously as “wavefronts” on a single SIMD processing unit 138. One or more wavefronts are included in a “workgroup,” which comprises a set of work items designated to execute the same program. A workgroup can be executed by executing each of the wavefronts that make up the workgroup. Alternatively, wavefronts are executed sequentially on a single SIMD unit 138, or partially or completely in parallel on different SIMD units 138. A wavefront can be considered as the largest set of work items that can be executed simultaneously on a single SIMD unit 138. Therefore, if a command received from processor 102 indicates that a particular program will be parallelized to the extent that the program cannot be executed simultaneously on a single SIMD unit 138, the program is decomposed into wavefronts that are parallelized on two or more SIMD units 138 or serialized on the same SIMD unit 138 (or parallelized and serialized as needed). Scheduler 136 performs various wavefront-related operations on different computing units 132 and SIMD units 138.

[0028] The parallelism provided by the computing unit 132 is suitable for graphics-related operations, such as pixel value calculation, vertex transformation, and other graphics operations. Therefore, in some instances, the graphics pipeline 134, which receives graphics processing commands from the processor 102, provides computing tasks to the computing unit 132 for parallel execution.

[0029] The computing unit 132 is also used to perform computational tasks that are unrelated to graphics or are not part of the "normal" operation of the graphics pipeline 134 (e.g., performing custom operations to supplement the processing performed for the operation of the graphics pipeline 134). The application program 126 or other software executing on the processor 102 transfers programs defining such computational tasks to the APD 116 for execution.

[0030] As described below, a memory controller includes a controller and a MOP array. The controller has inputs for receiving power state change request signals and outputs for providing memory operations. The MOP array includes multiple entries, each including multiple encoded fields. In response to activation of the power state change request signal, the controller accesses the MOP array to obtain at least one entry and issues at least one memory operation indicated by that entry. For example, the memory controller may have a portion of the MOP array describing a specific memory operation for implementing the power state change request. For example, DDR4 DRAM and LPDDR4 DRAM implement different state machines and different low-power modes, and require different sequences to transition from an active state to a low-power state. In one case, the memory controller may use the MOP array to define commands to be written to the DDR register DIMM and the load reduction DIMM using register control words (RCW) or buffer control words (BCW).

[0031] In another form, such a memory controller may be included in the processor of a processing system that includes a processor and memory modules. The processor may also include a physical interface (PHY) coupled between the memory controller and the memory system.

[0032] In another form, a method for controlling the power state of a memory system is disclosed. A power state change signal is received. A memory operation block (MOP) array is accessed in response to a power state change request signal. Entries in the MOP array are decoded into at least one memory operation. Each memory operation thus decoded is output. Decoding and output are repeated for consecutive entries in the MOP array until a predetermined termination condition is met. For example, the predetermined termination condition could be an empty entry in the MOP array. The received power state change request signal can be a change from an active state to a low-power state, such as a precharge power-off, self-refresh power-off, or idle power-off, or it can be a change from one operating frequency to another while in an active state. The BIOS can also program the MOP array in response to detected characteristics of the memory system.

[0033] Figure 3A data processing system 300 according to some embodiments is illustrated in block diagram form. The data processing system 300 generally includes a data processor 310 (in the form of an Accelerated Processing Unit (APU)), a memory system 320, a Peripheral Component Rapid Interconnect (PCIe) system 350, a Universal Serial Bus (USB) system 360, and a disk drive 370. The data processor 310 operates as the central processing unit (CPU) of the data processing system 300 and provides various buses and interfaces useful in modern computer systems. These interfaces include two Double Data Rate (DDRx) memory channels, a PCIe root complex for connecting PCIe links, a USB controller for connecting USB networks, and an interface to Serial Advanced Technology Attached (SATA) mass storage devices.

[0034] Memory system 320 includes memory channel 330 and memory channel 340. Memory channel 330 includes a set of dual in-line memory modules (DIMMs) connected to DDRx bus 332, including representative DIMMs 334, 336, and 338 corresponding to individual classes in this example. Similarly, memory channel 340 includes a set of DIMMs connected to DDRx bus 342, including representative DIMMs 344, 346, and 348.

[0035] PCIe system 350 includes PCIe switch 352, PCIe device 354, PCIe device 356, and PCIe device 358 connected to the PCIe root complex in data processor 310. PCIe device 356 is then connected to system basic input / output system (BIOS) memory 357. System BIOS memory 357 can be any of a variety of non-volatile memory types, such as read-only memory (ROM), electrically erasable programmable flash ROM (EEPROM), etc.

[0036] USB system 360 includes a USB hub 362 connected to a USB host device in data processor 310, and representative USB devices 364, 366, and 368 each connected to the USB hub 362. USB devices 364, 366, and 368 may be devices such as keyboards, mice, flash EEPROM ports, etc.

[0037] Disk drive 370 is connected to data processor 310 via SATA bus and provides large-capacity storage for operating system, applications, application files, etc.

[0038] By providing memory channels 330 and 340, the data processing system 300 is suitable for modern computing applications. Each of memory channels 330 and 340 can connect to state-of-the-art DDR memories, such as DDR version 4 (DDR4), low-power DDR4 (LPDDR4), graphics DDR version 5 (gDDR5), and high-bandwidth memory (HBM), and is compatible with future memory technologies. These memories offer high bus bandwidth and high-speed operation. They also provide low-power modes to save power for battery-powered applications such as laptops, and offer built-in thermal monitoring. As will be described in more detail below, the data processor 310 includes a memory controller capable of throttling power in certain situations to avoid overheating and reduce the chance of thermal overload.

[0039] Figure 4 The diagram illustrates the applicable... Figure 3 The data processing system 300 has an APU 400. The APU 400 generally includes a central processing unit (CPU) core complex 410, a graphics core 420, a set of display engines 430, a memory management hub 440, a data braid 450, a set of peripheral controllers 460, a set of peripheral bus controllers 470, a system management unit (SMU) 480, and a set of memory controllers 490.

[0040] CPU core complex 410 includes CPU core 412 and CPU core 414. In this example, CPU core complex 410 includes two CPU cores, but in other embodiments, CPU core complex 410 may include any number of CPU cores. Each of CPU cores 412 and 414 is bidirectionally connected to a system management network (SMN) (which forms a control weave) and a data weave 450, and is capable of providing memory access requests to the data weave 450. Each of CPU cores 412 and 414 may be a single core, or may further be a core complex of two or more single cores sharing certain resources such as cache.

[0041] The graphics core 420 is a high-performance graphics processing unit (GPU) capable of performing graphics operations, such as vertex processing, fragment processing, shading, and texture blending, in a highly integrated and parallel manner. The graphics core 420 is bidirectionally connected to the SMN and data weave 450 and can provide memory access requests to the data weave 450. In this respect, the APU 400 can support a unified memory architecture where the CPU core complex 410 and the graphics core 420 share the same memory space, or a memory architecture where the CPU core complex 410 and the graphics core 420 share a portion of the memory space, while the graphics core 420 also uses private graphics memory that the CPU core complex 410 cannot access.

[0042] Display engine 430 renders and rasterizes objects generated by graphics core 420 for display on a monitor. Graphics core 420 and display engine 430 are bidirectionally connected to a common memory management hub 440 for uniformly translating them into appropriate addresses in memory system 320, and memory management hub 440 is bidirectionally connected to data braiding 450 for generating such memory accesses and receiving read data returned from memory system.

[0043] Data weaving 450 includes a crossbar switch for routing memory access requests and responses between any memory access agent and memory controller 490. The data weaving also includes a BIOS-defined system memory map for determining the destination of memory accesses based on system configuration, as well as buffers for each virtual connection.

[0044] Peripheral controller 460 includes a USB controller 462 and a SATA interface controller 464, each of which is bidirectionally connected to system hub 466 and the SMN bus. These two controllers are merely examples of peripheral controllers that can be used with the APU 400.

[0045] The peripheral bus controller 470 includes a system controller or southbridge (SB) 472 and a PCIe controller 474, each of which is bidirectionally connected to an input / output (I / O) hub 476 and the SMN bus. The I / O hub 476 is also bidirectionally connected to a system hub 466 and a data braid 450. Therefore, for example, the CPU core can program registers in the USB controller 462, SATA interface controller 464, SB 472, or PCIe controller 474 via access routed through the data braid 450 to the I / O hub 476.

[0046] The SMU 480 is a local controller that controls the operation of resources on the APU 400 and synchronizes communication between them. The SMU 480 manages the power-on sequence of various processors on the APU 400 and controls multiple off-chip devices via reset signals, enable signals, and other signals. The SMU 480 includes... Figure 4 One or more clock sources (such as phase-locked loops (PLLs) not shown) are used to provide clock signals to each component in the APU 400. The SMU 480 also manages the power of various processors and other functional blocks and can receive measured power consumption values ​​from CPU cores 412 and 414 and graphics core 420 to determine appropriate power states.

[0047] The APU 400 also implements various system monitoring and power-saving functions. Specifically, one system monitoring function is thermal monitoring. For example, if the APU 400 gets hot, the SMU 480 can reduce the frequency and voltage of CPU cores 412 and 414 and / or graphics core 420. If the APU 400 becomes too hot, it can be completely shut down. Thermal events can also be received by the SMU 480 from external sensors via the SMN bus, and in response, the SMU 480 can reduce the clock frequency and / or supply voltage.

[0048] Figure 5 The diagram illustrates, in block form, an application of some implementation schemes. Figure 4 The APU400 includes a memory controller 500 and an associated physical interface (PHY) 530. The memory controller 500 includes a memory channel 510 and a power engine 520. The memory channel 510 includes a host interface 512, a memory channel controller 514, and a physical interface 516. The host interface 512 bidirectionally connects the memory channel controller 514 to a data braid 450 via an Extensible Data Port (SDP). The physical interface 516 bidirectionally connects the memory channel controller 514 to the PHY 530 via a bus conforming to the DDR-PHY Interface Specification (DFI). The power engine 520 bidirectionally connects to the SMU 480 via the SMN bus, bidirectionally connects to the PHY 530 via the Advanced Peripheral Bus (APB), and also bidirectionally connects to the memory channel controller 514. The PHY 530 has interfaces such as... Figure 3 The memory channel 330 or memory channel 340 is bidirectionally connected. The memory controller 500 is an instantiation of a memory controller using a single memory channel using a single memory channel controller 514, and has a power engine 520 that controls the operation of the memory channel controller 514 in a manner that will be further described below.

[0049] Figure 6 The diagram illustrates, in block form, an application of some implementation schemes. Figure 4The APU400 includes another memory controller 600 and associated PHYs 640 and 650. Memory controller 600 includes memory channels 610 and 620 and a power engine 630. Memory channel 610 includes a host interface 612, a memory channel controller 614, and a physical interface 616. Host interface 612 connects memory channel controller 614 bidirectionally to data braid 450 via an SDP. Physical interface 616 connects memory channel controller 614 bidirectionally to PHY 640 and is DFI compliant. Memory channel 620 includes a host interface 622, a memory channel controller 624, and a physical interface 626. Host interface 622 connects memory channel controller 624 bidirectionally to data braid 450 via another SDP. Physical interface 626 connects memory channel controller 624 bidirectionally to PHY 650 and is DFI compliant. The power engine 630 is bidirectionally connected to the SMU 480 via the SMN bus, bidirectionally connected to the PHYs 640 and 650 via the APB bus, and also bidirectionally connected to the memory channel controllers 614 and 624. The PHY 640 has connections to, for example... Figure 3 The memory channel 330 has bidirectional connections to the memory channel. The PHY 650 has connections to, for example... Figure 3 The memory channel 340 has a bidirectional connection to the memory channel. The memory controller 600 is an instantiation of a memory controller with two memory channel controllers, and uses a shared power engine 630 to control the operation of both memory channel controller 614 and memory channel controller 624 in a manner that will be further described below.

[0050] Figure 7 A memory controller 700 according to some embodiments is shown in block diagram form. The memory controller 700 generally includes a memory channel controller 710 and a power controller 750. The memory channel controller 710 generally includes an interface 712, a queue 714, a command queue 720, an address generator 722, a content-addressable memory (CAM) 724, a replay queue 730, a refresh logic block 732, a timing block 734, a page table 736, an arbitrator 738, an error correction code (ECC) check block 742, an ECC generation block 744, and a write data buffer (WDB) 746.

[0051] Interface 712 has a first bidirectional connection to data braid 450 via an external bus and has an output. In memory controller 700, this external bus is compatible with Advanced Scalable Interface Version 4 (referred to as "AXI4") as specified by ARM Holdings, PLC of Cambridge, England, but may be other types of interfaces in other embodiments. Interface 712 translates memory access requests from a first clock domain called the FCLK (or MEMCLK) domain to a second clock domain called the UCLK domain within memory controller 700. Similarly, queue 714 provides memory access from the UCLK domain to the DFICLK domain associated with the DFI interface.

[0052] Address generator 722 decodes the addresses of memory access requests received from data weaving 450 via the AXI4 bus. The memory access request includes an access address in the physical address space represented in a normalized format. Address generator 722 translates the normalized address into a format that can be used to address the actual memory devices in memory system 320 and efficiently schedule the associated access. This format includes a region identifier that associates the memory access request with a specific level, row address, column address, bank address, and bank group. At startup, the system BIOS queries the memory devices in memory system 320 to determine their size and configuration and programs a set of configuration registers associated with address generator 722. Address generator 722 uses the configuration stored in the configuration registers to translate the normalized address into the appropriate format. Command queue 720 is a queue of memory access requests received from memory access agents (such as CPU cores 412 and 414 and graphics core 420) in data processing system 300. Command queue 720 stores the address field decoded by address generator 722, as well as other address information that allows arbitrator 738 to effectively select memory access, including access type and quality of service (QoS) identifiers. CAM 724 includes information on enforcing ordering rules, such as write-after-write (WAW) and read-after-write (RAW) ordering rules.

[0053] Replay queue 730 is a temporary queue used to store memory accesses awaiting responses, selected by arbitrator 738, such as address and command parity responses, write cyclic redundancy check (CRC) responses for DDR4 DRAM, or write and read CRC responses for gDDR5 DRAM. Replay queue 730 accesses ECC check block 742 to determine whether the returned ECC is correct or indicates an error. Replay queue 730 allows replay accesses in the event of a parity or CRC error in one of these cycles.

[0054] The refresh logic 732 includes a state machine for generating various power-down, refresh, and termination resistor (ZQ) calibration cycles separately from normal read and write memory access requests received from the memory access agent. For example, if the memory level is in a pre-charge power-down state, it must be periodically woken up to run a refresh cycle. The refresh logic 732 periodically generates refresh commands to prevent data errors caused by charge leakage from the storage capacitors of the memory cells in the DRAM chip. Furthermore, the refresh logic 732 periodically calibrates ZQ to prevent on-die termination resistor mismatch due to thermal variations in the system.

[0055] Arbitrator 738 is bidirectionally connected to command queue 720 and is the core of memory channel controller 710. This arbitrator improves efficiency by intelligently scheduling accesses to increase memory bus utilization. Arbitrator 738 uses timing block 734 to enforce correct timing relationships by determining whether certain accesses in command queue 720 are eligible to be issued based on DRAM timing parameters. For example, each DRAM has a minimum specified time between activation commands, called "tRC". Timing block 734 maintains a set of counters that determine eligibility based on this and other timing parameters specified in the JEDEC specification, and this timing block is bidirectionally connected to replay queue 730. Page table 736 maintains status information for active pages in each bank and level of the memory channel for arbitrator 738 and is bidirectionally connected to replay queue 730.

[0056] In response to a write memory access request received from interface 712, ECC generation block 744 calculates the ECC based on the write data. DB 746 stores the write data and ECC of the received memory access request. When arbitrator 738 selects the corresponding write access for dispatch to the memory channel, the DB outputs the combined write data / ECC to queue 714.

[0057] The power controller 750 typically includes an interface 752 to an Advanced Scalable Interface version one (AXI), an APB interface 754, and a power engine 760. Interface 752 has a first bidirectional connection to the SMN, which includes a method for receiving signals from the SMN. Figure 7The input and output of the event signal labeled “EVENT_n” are shown separately. The APB interface 754 has inputs connected to the output of interface 752, and outputs for connecting to the PHY via APB. The power engine 760 has inputs connected to the output of interface 752, and outputs connected to the input of queue 714. The power engine 760 includes a set of configuration registers 762, a microcontroller (μC) 764, a self-refresh controller (SLFREF / PE) 766, and a reliable read / write timing engine (RRW / TE) 768. The configuration registers 762 are programmed via the AXI bus and store configuration information to control the operation of various blocks in the memory controller 700. Therefore, the configuration registers 762 have inputs connected to the PHY via APB. Figure 7 The outputs of these blocks are not shown in detail. The self-refresh controller 766 is an engine that allows manual refresh generation in addition to automatic refresh generation by the refresh logic 732. The reliable read / write timing engine 768 provides a continuous stream of memory accesses to memory or I / O devices for purposes such as DDR interface maximum read latency (MRL) training and loopback testing.

[0058] The memory channel controller 710 includes circuitry that allows it to select memory accesses for assignment to associated memory channels. To make the desired arbitration decision, the address generator 722 decodes address information into pre-decoded information, including the memory system level, row address, column address, bank address, and bank group, and the command queue 720 stores the pre-decoded information. The configuration register 762 stores configuration information to determine how the address generator 722 decodes the received address information. The arbitrator 738 uses the decoded address information, timing qualification information indicated by the timing block 734, and active page information indicated by the page table 736 to efficiently schedule memory accesses while adhering to other criteria such as QoS requirements. For example, the arbitrator 738 implements a preference for accessing open pages to avoid the overhead of precharge and activation commands required to change memory pages, and hides overhead accesses to one bank by interleaving overhead accesses to one bank with read and write accesses to another bank. Specifically, during normal operation, the arbitrator 738 typically keeps pages open in different banks until they need to be precharged before selecting a different page.

[0059] Figure 8 The diagram illustrates the corresponding implementation schemes according to some embodiments. Figure 3 The data processing system 300 is a part of the data processing system 800. The data processing system 800 generally includes a memory controller labeled "MC" 810, a PHY 820, and a memory module 830.

[0060] The memory controller 810 receives memory access requests from the processor's memory access agent (such as CPU core 412 or graphics core 420) and provides it with a memory response. The memory controller 810 corresponds to... Figure 4 Any one of the memory controllers in memory controller 490. Memory controller 810 outputs memory access to PHY 820 and receives responses from it via a DFI-compatible bus.

[0061] The PHY 820 is connected to the memory controller 810 via the DFI bus. The PHY performs physical signaling in response to a received memory access by providing a set of command and address outputs labeled “C / A” and a set of 72 bidirectional data signals labeled “DQ” (including 64 bits of data and 8 bits of ECC).

[0062] Memory module 830 can support any of several memory types and speed levels. In the illustrated embodiment, memory module 830 is a DDR4 registered DIMM (RDIMM) comprising a set of memory chips 840 labeled "DDR4", a register clock driver 850 labeled "RCD", and a set of buffers 860 labeled "B". Memory chip 840 comprises a set of M-bit × N memory chips. To support 72 data signals (64 bits of data plus 8 bits of ECC), M*N = 72. For example, if each memory chip is 4 × 4 (N = 4), then memory module 830 comprises 18 DDR4 memory chips. Alternatively, if each memory chip is 8 × 8 (N = 8), then memory module 830 comprises 9 DDR4 memory chips. Each buffer in buffer 860 is associated with an N × N memory chip and is used to latch the corresponding N bits of data. Figure 8 In the example shown, memory module 830 includes DDR4 memory, and the C / A signals include those described in the DDR4 specification. The DDR4 specification specifies a “fly-through” architecture, in which the same C / A signals received and latched by RCD 850 are redriven left and right to each memory chip in memory chip 840. However, the data signal DQ is only provided to the corresponding buffer and memory.

[0063] The memory module 830 operates according to control information programmed into the register control word (RCW) for the RCD 850 and into the buffer control word (BCW) for the buffer 860. Therefore, when the memory controller 810 places the memory module 830 in a low-power state, it also changes the settings in the RCW and BCW in a manner that will be described more fully below.

[0064] Although the data processing system 800 uses registered, buffered DDR4 DRAM DIMMs as memory modules 830, the memory controller 810 and PHY 820 can also connect to several different types of memory modules. Specifically, the memory controller 810 and PHY 820 can support several different types of memory (e.g., DDR, flash memory, PCM, etc.), several different register conditions (unused, RCD, flash controller, etc.), and several different buffer conditions (unused, data buffer only, etc.), enabling the memory controller 810 to support multiple combinations of memory types, register conditions, and buffer conditions. To support these combinations, the memory controller 810 implements an architecture that allows for unique plans for entering and exiting low-power modes, which the system BIOS can program for specific memory system characteristics. These features will now be described.

[0065] Figure 9 The diagram illustrates the corresponding implementation schemes according to some embodiments. Figure 7 The memory channel controller 750 is a portion of the memory channel controller 900. The memory channel controller 900 includes, as described above... Figure 7 The diagram shows the UMCSMN 752 and self-refresh controller 766, as well as the memory operation (MOP) array 710. The UMCSMN 752, as described above, has a first port for connecting to the SMN and, as shown in related details here, an input for receiving a power state change request signal labeled "Power Request" from the data weave 450, and an output for providing a power state change confirmation signal labeled "Power Confirmation" to the data weave 450. The UMCSMN 752 also has a second port with a first output for providing a memory power state change request signal labeled "M_PSTATE REQ" and a second output for providing data stored in the MOP array 710. The self-refresh controller 766 has an input connected to the first output of the second port of the UMCSMN 752, a bidirectional port, and an output connected to the BEQ 714 for providing decoded MOPs to the BEQ 714. The MOP array 910 has an input to the second output of the second port connected to the UMCSMN 752 and a bidirectional connection to the self-refresh controller 766, and is divided into a first part 912 for storing commands (i.e., MOPs) and a second part 914 for storing data.

[0066] In one example, at startup, the system BIOS, stored in system BIOS memory 357, queries memory system 320 to determine the type and capabilities of the installed memory. This is typically achieved by reading registers in the Serial Presence Detection (SPD) memory present on each DIMM in the system. For example, the PHY may support any of DDR3, DDR4, Low Power DDR4 (LPDDR4), and Graphics DDR Version 5 (gDDR5) memory. In response to detecting the type and capabilities of the memory installed in memory system 320, the system BIOS populates MOP array 910 with a command sequence that initiates entry into and exit from low-power modes supported by the specific type of memory.

[0067] In the illustrated embodiment, the memory channel controller 750 supports various device low-power states defined according to the model described in the Advanced Configuration and Power Interface (ACPI) specification. According to the ACPI specification, the operating state of a device (such as memory controller 700) is referred to as the D0 state or "fully on" state. Other states are low-power states and include D1, D2, and D3 states, where the D3 state is the "off" state. Memory controller 700 is capable of placing memory system 320 in a low-power state corresponding to the D state of memory controller 700, and of performing frequency and / or voltage changes in the D0 state. Upon receiving a power request, UMCSMN 752 provides an M_PSTATE REQ signal to self-refresh controller 766 to indicate which power state is requested. Self-refresh controller 766 accesses MOP array 910 in response to executing a sequence of MOPs that places the RCW and BCW of the memory chip and DIMM (if supported) in the appropriate state of the requested D state. The self-refresh controller 766 outputs the index to the MOP array 910, and in response, the MOP array 910 returns an encoded command (MOP).

[0068] The memory channel controller 750 is implemented with a relatively small circuit area, supporting multiple memory types with different characteristics, by including a MOP array 910 to store programmable commands from the self-refresh controller 766. Furthermore, this provides an upward-compatible architecture that allows for memory state changes for memory types and characteristics not yet specified but potentially to be specified in the future. Therefore, the memory channel controller 750 is also modular and avoids the need for costly redesigns in the future.

[0069] The interaction between these memory controller device power states (D states) and DRAM operation will now be described. The D0 state is the operating state of the memory controller 700. In the D0 state, the memory controller 700 supports four programmable power states (P states), each with a different MEMCLK frequency and associated timing parameters. The memory controller 700 maintains a set of registers for each P state, which stores the timing parameters of that P state and defines the context. The memory controller 700 puts the DRAM into self-refresh mode to change the P state / context. The MOP array 910 includes a set of commands that are used in conjunction with frequency changes in the D0 state to support correct sequencing.

[0070] The D1 state is known as the stop clock state and is used for memory state change requests. When the memory controller 700 is in the D1 state, entry and exit latency are minimal unless the PHY 820 needs to be retrained. As a result of the D1 power state change request, the memory controller 700 typically does not flush any arbitration queue entries. However, the memory controller 700 pre-flushes all writes in the command queue 720, without typically performing a normal pending flush. The memory controller 700 places all memory chips in the system into a precharge-down or self-refresh state.

[0071] The D2 state is referred to as the standby state and corresponds to system C1 / C2 / C3 and the stop clock / intermittent state. This state is a lower power state for the operation of the memory controller 700. In the D2 state, the memory controller 700 uses local clock gating and optional power gating to further reduce power. The memory controller 700 refreshes writes and reads from command queue 720. In the D2 state, the memory controller 700 also places all memories in the system into a precharge-down state while enabling automatic self-refresh. However, because D2 is a deeper power state, the memory controller performs all pending (“under-refreshed”) refreshes using automatic self-refresh before entering the precharge-down state.

[0072] The D3 state is referred to as the suspended state. The memory controller 700 supports two D3 states. The first D3 state is used for the system S3 state. The memory controller 700 places the DRAM and PHY in a minimum power state in anticipation of entering the system S3 state. The memory controller 700 typically refreshes writes from the command queue 720 and performs a suspended refresh cycle. The second D3 state is used for asynchronous DRAM refresh (ADR-style self-refresh). ADR is a feature in servers used to refresh suspended write data to non-volatile memory during power failures or system crashes. The DRAM and PHY are again placed in a precharge-down state, and automatic self-refresh is enabled.

[0073] As used herein, the power request signal indicates a change from one power state to another. The available power states differ for different memory types. Furthermore, as used herein, a "low power state" refers to a state that saves power compared to another state. For example, DDR4 SDRAM supports two low power states, called self-refresh and precharge power-off. However, LPDDR4 supports three low power states: active power-off, self-refresh power-off, and idle power-off. The conditions for entering and exiting these states differ and are specified in the state diagram of the corresponding published JEDEC standard, and a "low power state" includes any of these states.

[0074] The MOP array 910 supports a command format that allows for efficient encoding of commands to support all these power state changes. The MOP array 910 uses two arrays, called "SrEnterMOP" and "SrExitMOP," for each of the four P-state contexts. SrEnterMOP is processed before a self-refresh upon entering a P-state request. SrExitMOP is processed after a self-refresh upon exiting a P-state request. The MOP array specifies a sequential list of commands for the Mode Register (MR), MR with per-DRAM accessibility (PDA), Register Control Word (RCW), or Buffer Control Word (BCW). Upon receiving a power state change request, the self-refresh controller 766 accesses the commands of the selected context in the MOP array 910 to determine the sequence and timing of the MOPs issued to the memory system.

[0075] The MOPs in section 912 include fields representing one or more corresponding D states within section 912. Therefore, the self-refresh controller 766 scans the MOP array 912 starting from a first position to find commands applicable to a specific context and ignores MOPs that are not applicable to the current context. The MOP array 912 also includes counter values ​​to determine the correct timing between MOPs when appropriate, thereby satisfying the dynamic timing parameters of the memory chip. After the start of the command sequence, the self-refresh controller 766 continues scanning the MOP array 912 and executing valid commands until it encounters an empty entry, indicating the end of the power state change sequence.

[0076] Figure 10 A method 1000 for reducing the boot time of a System-on-a-Chip (SOC) during memory context recovery is shown, which reduces the boot time of the SOC by reducing DDR training. Method 1000 includes, at step 1010, storing DRAM controller and DDR PHY settings (including any values ​​of the processor / platform / DRAM combination) in a non-volatile location prior to recovery.

[0077] At step 1020, in DDR4 mode, for S3, the DRAM controller then sets the DRAM to self-refresh mode. Then at step 1025, the DRAM controller and DDR PHY are powered down to save total system power.

[0078] At step 1030, during system recovery, the BIOS restores the DRAM controller and DDR PHY settings from non-volatile memory. For S3, at step 1035, the self-refresh state machine MOP array (small code of the optimized state machine) is programmed to exit self-refresh and update the state of any DRAM device for the target power management state (memory P state).

[0079] At step 1040, the DRAM controller and DDR PHY settings are finalized to enable task mode operation with the SOC.

[0080] Figure 11 A method 1100 for performing memory context recovery during memory context recovery is illustrated, which reduces the boot time of the SoC by reducing DDR training. Method 1100 includes, at step 1110, storing DRAM controller and DDR PHY settings (including any values ​​of the processor / platform / DRAM combination) in a non-volatile memory location prior to recovery.

[0081] At step 1115, in DDR4 mode, the DRAM controller sets the DRAM to self-refresh mode. Then, at step 1120, the DRAM controller and DDR PHY are powered down to save total system power or the entire system may have been de-energized.

[0082] At step 1130, following step 1030 of method 1000, the BIOS restores the DRAM controller and DDR PHY settings from non-volatile memory during system recovery, and / or optionally resets the DRAM device at step 1135. This reset includes, at step 1140, programming the self-refresh state machine MOP array to reinitialize the DRAM device (according to the JEDEC specification sequence) for a target power management state (memory P state). At step 1145, the DRAM controller and DDR PHY settings are finalized for task-mode operation with the SOC.

[0083] Although methods 1000 and 1100 are described using separate figures, each part of methods 1000 and 1100 is interchangeable or can be used to supplement the steps described in methods 1000 and 1100. In methods 1000 and 1100, a software-mode register access mechanism can be used to finalize the DRAM settings. Although DRAM is used in this description for clarity, the described methods are also applicable to other associated components on RDIMM or LRDIMM modules, such as RCDs or DBs.

[0084] It should be understood that many variations are possible based on the disclosure herein. Although features and elements are described above in specific combinations, each feature or element may be used alone without other features and elements, or in various combinations with or without other features and elements.

[0085] The various functional units shown in the figures and / or described herein (including, but not limited to, processor 102, input driver 112, input device 108, output driver 114, output device 110, accelerated processing device 116, scheduler 136, graphics processing pipeline 134, computing unit 132, SIMD unit 138, and APU 310) may be implemented as a general-purpose computer, processor, or processor core, or implemented as a program, software, or firmware, stored on a non-transitory computer-readable medium or another medium, executable by a general-purpose computer, processor, or processor core. The methods provided can be implemented in a general-purpose computer, processor, or processor core. Suitable processors include, for example, general-purpose processors, special-purpose processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, application-specific integrated circuits (ASICs), field-programmable gate array (FPGA) circuits, any other type of integrated circuit (IC), and / or state machines. Such processors can be manufactured by configuring a manufacturing process using the results of processed hardware description language (HDL) instructions and other intermediate data, including netlists (such instructions can be stored on a computer-readable medium). The result of this processing can be a mask, which is then used in a semiconductor manufacturing process to manufacture a processor that implements the features of this disclosure.

[0086] The methods or flowcharts provided herein can be implemented in a computer program, software, or firmware incorporated into a non-transitory computer-readable storage medium for execution by a general-purpose computer or processor. Examples of non-transitory computer-readable storage media include read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (such as internal hard disks and removable disks), magneto-optical media, and optical media such as CD-ROMs and digital versatile optical discs (DVDs).

Claims

1. A method for reducing boot time of a system-on-chip (SoC) by reducing training of double data rate (DDR) memory, the method comprising: The dynamic random access memory (DRAM) controller and DDR physical interface (PHY) settings are stored in non-volatile memory for use during boot. Set the DRAM to self-refresh mode; Power off the DRAM controller and DDR PHY; During system recovery to boot training, the system basic input / output system BIOS is used to restore the DRAM controller and DDR PHY settings from the non-volatile memory to avoid memory retraining; The self-refresh state machine memory operation MOP array is programmed to cause the DRAM to exit self-refresh mode and update the state of any DRAM device for the target power management state, thereby reducing boot time. as well as Configure the DRAM controller and DDR PHY settings to operate in conjunction with the SOC for use at boot time.

2. The method of claim 1, wherein setting the DRAM to self-refresh mode includes retaining the memory contents.

3. The method of claim 2, wherein retaining the memory contents includes the DRAM entering a low-power mode.

4. The method of claim 1, wherein power-off of the DRAM controller and the DDR PHY enables rapid booting from the system state.

5. The method of claim 1, wherein a warm reset is achieved by powering off the DRAM controller and the DDR PHY.

6. The method of claim 1, wherein power-off of the DRAM controller and the DDR PHY achieves a reset of the DRAM controller and the DDR PHY.

7. The method of claim 1, wherein configuring the DRAM controller and DDR PHY settings to operate with the SOC includes configuring any controller settings and PHY settings by programming the configuration.

8. The method of claim 7, wherein the configuration settings include updating the MOP array and timing for optimal P state switching.

9. The method of claim 7, wherein the final determination of settings includes an initialization sequence.

10. A method for recovering the memory context of a system-on-a-chip (SoC) by reducing double data rate (DDR) memory training, the method comprising: The DRAM controller and DDR physical interface PHY settings are stored in non-volatile memory for use during boot. Power off the DRAM controller and DDR PHY; During system recovery to boot training, the system basic input / output system BIOS is used to restore the DRAM controller and DDR PHY settings from the non-volatile memory to avoid memory retraining; Reset DRAM; The self-refresh state machine memory operation MOP array is programmed to reinitialize the DRAM device state for the target power management state, thereby reducing boot time; as well as Configure the DRAM controller and DDR PHY settings to operate in conjunction with the SOC for use at boot time.

11. The method of claim 10, wherein powering off the DRAM controller does not preserve the memory state.

12. The method of claim 10, wherein power-off of the DRAM controller and the DDR PHY provides fast boot from system state.

13. The method of claim 10, wherein powering down the DRAM controller and the DDR PHY provides at least one of a warm reset, a cold reset, and a complete new power cycle starting from a mechanical shutdown.

14. The method of claim 10, wherein power-off of the DRAM controller and the DDR PHY provides a reset of the DRAM controller and the DDR PHY.

15. The method of claim 10, wherein configuring the DRAM controller and DDR PHY settings to operate with the SOC includes configuring any controller settings and PHY settings by programming the configuration.

16. The method of claim 15, wherein the configuration settings include updating the MOP array and timing for optimal P state switching.

17. The method of claim 15, wherein the final determination of settings includes an initialization sequence.

18. A system for reducing the boot time of a system-on-chip (SoC) by reducing training with double data rate (DDR) memory, the system comprising: The dynamic random access memory (DRAM) controller and DDR physical interface (PHY) have their settings stored in non-volatile memory for use during boot. Multiple DRAMs, wherein the multiple DRAMs are configured in self-refresh mode, The DRAM controller and the DDR PHY are powered off. During system recovery to boot training, the system basic input / output system BIOS restores the DRAM controller and DDR PHY settings from the non-volatile memory to avoid memory retraining; as well as A self-refreshing state machine memory operates a MOP array, which is programmed to cause the plurality of DRAMs to exit self-refresh mode and update the state of any DRAM device for a target power management state, thereby reducing boot time. The DRAM controller and DDR PHY settings are configured to operate in conjunction with the SOC for use during boot.

19. The system of claim 18, wherein setting the DRAM to self-refresh mode includes retaining the memory contents.

20. The system of claim 19, wherein retaining memory contents includes the DRAM entering a low-power mode.

Citation Information

Patent Citations

  • Fast exit from self-refresh state of a memory device

    CN102214152A

  • Method and system for protecting DRAM stored data of embedded system software

    CN105608023A