Methods and systems for transmitting commands to dynamic random access memory
By using a single clock signal and a synchronization signal in DRAM to receive and transmit commands and data, the problem of multi-clock signal deviation is solved, reducing the complexity and power consumption of DRAM and improving system performance.
Patent Information
- Application Number
- CN202210094417.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-11-10
- Filing Date
- 2022-01-26
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-01-26
AI Technical Summary
The prior art requires internal synchronization and training circuits to resolve deviations between multiple clock signals when transmitting commands and data to DRAM, resulting in increased internal circuit complexity of DRAM, increasing power consumption, and requiring additional receivers and I/O pins, limiting the use of other signals.
By receiving the synchronization signal on the input pin of the memory device, synchronizing the memory device to the clock edge of the single clock signal relative to the synchronization signal, different parts of the receiving command are on different clock edges, avoiding the use of internal synchronization and training circuits.
The implementation of commands and data receiving at different transmission rates through a single clock signal, reducing the internal circuit complexity and power consumption of DRAM, releasing the I/O pins for other signals, and improving system performance.
Smart Images

Figure CN114840454B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of priority to U.S. Provisional Patent Application Serial No. 63 / 144,971, filed on February 2, 2021, entitled “TECHNIQUES FOR TRANSFERRING COMMANDS TO A DRAM.” This application further claims the benefit of priority to U.S. Provisional Patent Application Serial No. 63 / 152,814, filed on February 23, 2021, entitled “DATA SCRAMBLING ON A MEMORY INTERFACE.” This application further claims the benefit of priority to U.S. Provisional Patent Application Serial No. 63 / 152,817, filed on February 23, 2021, entitled “DRAM COMMAND INTERFACE TRAINING.” This application further claims the benefit of priority to U.S. Provisional Patent Application Serial No. 63 / 179,954, entitled “DRAM WRITE TRAINING,” filed on April 26, 2021. The subject matter of these related applications is hereby incorporated by reference herein. background Technical Field
[0004] Various embodiments relate generally to computer memory devices and, more particularly, to techniques for transferring commands to dynamic random access memory. Background Art
[0006] A computer system typically includes one or more processing units, such as a central processing unit (CPU) and / or a graphics processing unit (GPU), and one or more memory systems, among other things. One type of memory system is called system memory, which is accessible to both the CPU and the GPU. Another type of memory system is graphics memory, which is typically only accessible to the GPU. These memory systems include multiple memory devices. An example memory device employed in system memory and / or graphics memory is synchronous dynamic random access memory (SDRAM, or more concisely, DRAM).
[0007] Conventionally, high-speed DRAM memory devices use multiple interfaces. These interfaces use multiple independent clock signals to transmit commands and data to or from DRAM. A low-speed clock signal is used to transmit commands to DRAM via a command interface. Such commands include commands to initiate write operations, commands to initiate read operations, etc. After the instruction is transmitted to the DRAM, a second high-speed clock signal is used to transfer data to the DRAM and transfer data from the DRAM via a data interface. In some cases, commands and data can be overwritten. For example, a command for a first DRAM operation can be transferred to the DRAM via a low-speed clock signal. Subsequently, while a command for a second DRAM operation is transmitted via a low-speed clock signal, data for a first DRAM operation can be transferred to the DRAM via a high-speed clock signal. Then, while a command for a third DRAM operation is transmitted via a low-speed clock signal, data for a second DRAM operation can be transferred to the DRAM via a high-speed clock signal, etc.
[0008] When the command interface and the data interface use different clock signals, the high-speed clock signal and the low-speed clock signal need to be synchronized with each other at the clock signal source generator. This clock signal is called a source synchronous clock signal. The high-speed clock signal and the low-speed clock signal are transmitted to the DRAM from the clock signal source generator via independent signal paths. These signal paths may have different lengths, resulting in different delay times between the clock signal source generator and the DRAM. In addition, the signal path can travel through different intermediate devices, which may have different internal delays and are subject to changes in internal delays. These changes are due to process changes during manufacturing and local changes caused by changes in operating temperature, power supply voltage, etc.
[0009] As a result, even if the two clock signals are synchronized at the source, it is not assumed that the two clock signals are synchronized when the clock signals arrive at the DRAM. To address this phenomenon, the DRAM includes synchronization and training circuits that determine the deviation between the two clock signals. This synchronization and training circuit allows the DRAM to properly manage internal timing so that commands and data are correctly transmitted to and from the DRAM.
[0010] One disadvantage of this technique is that the synchronization and training circuits increase the complexity of the internal circuitry of the DRAM, consume the surface area of the DRAM die, and increase power consumption. Another disadvantage of this technique is that two receivers are required and two input / output (I / O) pins of each DRAM memory device are consumed to receive the two clock signals. As a result, the additional receiver and I / O pins for receiving the second clock signal are not available to accommodate other signals, such as additional command bits, data bits, or control signals. In addition, some DRAM modules include multiple DRAM devices. In addition, each clock signal can be a differential signal, requiring two I / O pins for each clock signal. In one example, a DRAM module with four DRAM devices and differential clock signals will require 8 I / O pins for the data clock signal and 8 additional I / O pins for the command clock signal.
[0011] Another disadvantage of this technique is that the overhead of performing this synchronization and training takes a finite amount of time. In addition, this synchronization and training is performed each time the DRAM memory device exits a low power state, such as a power-off state or a self-refresh state. Therefore, the latency for DRAM memory devices with multiple clock inputs to exit a low power state is relatively high. This relatively high latency to exit from a low power state reduces the performance of systems that employ these types of DRAM memory devices. Alternatively, systems that employ these types of DRAM memory devices may choose not to utilize these low power states with longer exit latencies. As a result, such a system may have higher performance, but may not be able to obtain the benefits of low power states, such as a power-off state, a self-refresh state, etc.
[0012] As noted above, there is a need in the art for more efficient techniques for transferring commands and data to and from memory devices. Summary of the invention
[0013] Various embodiments of the present disclosure set forth a computer-implemented method for transmitting a command to a memory device. The method includes: receiving a synchronization signal on an input pin of the memory device, wherein the synchronization signal specifies a starting point of a first command. The synchronization signal may be in the form of a signal (e.g., a pulse signal) received on any one or more input / output pins (e.g., a command input / output pin) of the memory device. Additionally or alternatively, the synchronization signal may be any signal and / or other indication of a phase of an input clock signal used by the memory device to identify a starting point for a command.
[0014] The method also includes: synchronizing the memory device to a first clock edge of the clock signal input relative to the synchronization signal. The method also includes: receiving a first portion of a first command at the first clock edge. The method also includes: receiving a second portion of the first command at a second clock edge of the clock signal input subsequent to the first clock edge.
[0015] Other embodiments include, but are not limited to, a system implementing one or more aspects of the disclosed technology, and one or more computer-readable media including instructions for performing one or more aspects of the disclosed technology, and a method for performing one or more aspects of the disclosed technology.
[0016] At least one technical advantage of the disclosed technology over the prior art is that, using the disclosed technology, commands and data are received by the memory device via a single clock signal at different transfer rates. As a result, the memory device does not require internal synchronization and training circuits to resolve possible deviations between multiple clock signals. Another advantage of the disclosed technology is that only one receiver and I / O pin are required to receive the clock signal, rather than two receivers and I / O pins. As a result, the complexity, surface area, and power consumption of the internal circuitry of the DRAM die can be reduced relative to methods involving multiple clock signals. In addition, the I / O pin previously used to receive the second clock signal can be used for another function, such as an additional command bit, data bit, or control signal. These advantages represent one or more technical improvements over the prior art methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to understand the features of the various embodiments described above in more detail, the above briefly summarized inventive concept can be described in more detail by referring to various embodiments (some of which are illustrated in the accompanying drawings). However, it should be noted that the attached drawings only illustrate typical embodiments of the inventive concept and should not be considered to limit the scope in any way, and there are other equally effective embodiments.
[0018] Figure 1 is a block diagram of a computer system configured to implement one or more aspects of various embodiments;
[0019] Figure 2 According to various embodiments, it is included in Figure 1 A block diagram of a clock architecture for a memory device in a system memory and / or a parallel processing memory of a computer system;
[0020] Figure 3 is a method for including Figure 1 A more detailed block diagram of a command address clock architecture of a memory device in a system memory and / or parallel processing memory of a computer system;
[0021] Figure 4 It is a diagram showing that initialization according to various embodiments includes Figure 1 A timing diagram for receiving commands from a memory device in a system memory and / or a parallel processing memory of a computer system;
[0022] Figure 5 is a diagram illustrating the transmission of consecutive commands to a Figure 1 timing diagrams of memory devices in a system memory and / or a parallel processing memory of a computer system; and
[0023] Figure 6 is a method for transmitting commands to a Figure 1 A flow chart of method steps for storing memory devices in a system memory and / or parallel processing memory of a computer system. DETAILED DESCRIPTION
[0024] In the following description, numerous specific details are set forth to provide a more thorough understanding of various embodiments. However, it will be apparent to one skilled in the art that these inventive concepts may be practiced without one or more of these specific details.
[0025] System Overview
[0026] Figure 1 1 is a block diagram of a computer system 100 configured to implement one or more aspects of various embodiments. As shown, computer system 100 includes, but is not limited to, a central processing unit (CPU) 102 and a system memory 104, which is coupled to a parallel processing subsystem 112 via a memory bridge 105 and a communication path 113. Memory bridge 105 is coupled to system memory 104 via a system memory controller 130. Memory bridge 105 is also coupled to an I / O (input / output) bridge 107 via a communication path 106, and I / O bridge 107 is coupled to a switch 116. Parallel processing subsystem 112 is coupled to parallel processing memory 134 via a parallel processing subsystem (PPS) memory controller 132.
[0027] In operation, I / O bridge 107 is configured to receive user input information from input device 108 (such as a keyboard or mouse) and forward the input information to CPU 102 for processing via communication path 106 and memory bridge 105. Switch 116 is configured to provide connections between I / O bridge 107 and other components of computer system 100 (such as network adapter 118 and various add-in cards 120 and 121).
[0028] As also shown, the I / O bridge 107 is coupled to a system disk 114, which can be configured to store content, applications, and data for use by the CPU 102 and the parallel processing subsystem 112. In general, the system disk 114 provides non-volatile storage for applications and data, and can include a fixed or removable hard drive, a flash memory device, and a CD-ROM (Compact Disc Read Only Memory), DVD-ROM (Digital Versatile Disk-ROM), Blu-ray, HD-DVD (High Definition DVD), or other magnetic, optical, or solid-state storage device. Finally, although not explicitly shown, other components (such as a universal serial bus or other port connection, an optical disk drive, a digital versatile disk drive, a film recording device, etc.) can also be connected to the I / O bridge 107.
[0029] In various embodiments, memory bridge 105 may be a north bridge chip, and I / O bridge 107 may be a south bridge chip. In addition, communication paths 106 and 113 and other communication paths may be implemented within computer system 100 using any technically suitable protocol, including but not limited to AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol known in the art.
[0030] In some embodiments, parallel processing subsystem 112 includes a graphics subsystem that transmits pixels to display device 110, which can be any conventional cathode ray tube, liquid crystal display, light emitting diode display, etc. In such embodiments, parallel processing subsystem 112 incorporates circuits optimized for graphics and video processing, including, for example, video output circuits. Such circuits can be combined across one or more parallel processing units (PPUs) included in parallel processing subsystem 112. In some embodiments, each PPU includes a graphics processing unit (GPU), which can be configured to implement a graphics rendering pipeline for performing various operations related to generating pixel data based on graphics data provided by CPU 102 and / or system memory 104. Each PPU can be implemented using one or more integrated circuit devices, such as a programmable processor, an application specific integrated circuit (ASIC), or a memory device, or in any other technically feasible manner.
[0031] In some embodiments, parallel processing subsystem 112 incorporates circuits optimized for general purpose and / or computational processing. Again, such circuits may be combined across one or more PPUs included within parallel processing subsystem 112, which are configured to perform such general purpose and / or computational operations. In yet another embodiment, one or more PPUs included within parallel processing subsystem 112 may be configured to perform graphics processing, general purpose processing, and computational processing operations. System memory 104 includes at least one device driver 103, which is configured to manage processing operations of one or more PPUs within parallel processing subsystem 112.
[0032] In various embodiments, the parallel processing subsystem 112 may be coupled to Figure 1 For example, parallel processing subsystem 112 may be integrated with CPU 102 and other connected circuits on a single chip to form a system on a chip (SoC).
[0033] In operation, CPU 102 is the main processor of computer system 100, controlling and coordinating the operation of other system components. In particular, CPU 102 issues commands that control the operation of the PPUs within parallel processing subsystem 112. In some embodiments, CPU 102 writes a command stream for the PPUs within parallel processing subsystem 112 to a data structure ( ) that may be located in system memory 104, PP memory 134, or another storage location accessible to both CPU 102 and the PPUs. Figure 1 The push buffer is a data structure that is not explicitly shown in the figure. A pointer to the data structure is written to the push buffer to start processing the command stream in the data structure. The PPU reads the command stream from the push buffer and then executes the commands asynchronously with respect to the operation of the CPU 102. In an embodiment where multiple push buffers are generated, the application can specify an execution priority for each push buffer via the device driver 103 to control the scheduling of different push buffers.
[0034] Each PPU includes an I / O (input / output) unit that communicates with the rest of the computer system 100 via a communication path 113 and a memory bridge 105. The I / O unit generates packets (or other signals) for transmission on the communication path 113, and also receives all incoming packets (or other signals) from the communication path 113, directing the incoming packets to the appropriate components of the PPU. The connection of the PPU to the rest of the computer system 100 can vary. In some embodiments, the parallel processing subsystem 112 including at least one PPU is implemented as an add-in card that can be inserted into an expansion slot of the computer system 100. In other embodiments, the PPU can be integrated on a single chip with a bus bridge (such as, memory bridge 105 or I / O bridge 107). Likewise, in other embodiments, some or all elements of the PPU can be included in a single integrated circuit or chip system (SoC) with the CPU 102.
[0035] The CPU 102 and the PPU within the parallel processing subsystem 112 access the system memory via the system memory controller 130. The system memory controller 130 sends signals to the memory devices included in the system memory 104 to start the memory devices, send commands to the memory devices, write data to the memory devices, read data from the memory devices, etc. An example memory device employed in the system memory 104 is a double data rate SDRAM (DDR SDRAM, or more simply, DDR). DDR memory devices perform memory write and read operations at twice the data rate of previous generation single data rate (SDR) memory devices.
[0036] In addition, the PPU and / or other components within the parallel processing subsystem 112 access the PP memory 134 via a parallel processing subsystem (PPS) memory controller 132. The PPS memory controller 132 sends signals to memory devices included in the PP memory 134 to start the memory devices, send commands to the memory devices, write data to the memory devices, read data from the memory devices, etc. An example memory device used in the PP memory 134 is a synchronous graphics random access memory (SGRAM), which is a special form of SDRAM used for computer graphics applications. A specific type of SGRAM is a graphics double data rate SGRAM (GDDR SDRAM, or more simply, GDDR). Compared to DDR memory devices, GDDR memory devices are configured with a wider data bus to transfer more data bits during each memory read and write operation. By using double data rate technology and a wider data bus, GDDR memory devices are able to achieve the high data transfer rates typically required by the PPU.
[0037] It should be understood that the system shown herein is illustrative and that variations and modifications are possible. The connection topology, including the number and arrangement of bridges, the number of CPUs 102, and the number of parallel processing subsystems 112, may be modified as desired. For example, in some embodiments, system memory 104 may be directly connected to CPU 102 rather than being connected through memory bridge 105, and other devices will communicate with system memory 104 via memory bridge 105 and CPU 102. In other alternative topologies, parallel processing subsystem 112 may be connected to I / O bridge 107 or directly to CPU 102 rather than being connected to memory bridge 105. In other embodiments, I / O bridge 107 and memory bridge 105 may be integrated into a single chip rather than existing as one or more discrete devices. Finally, in some embodiments, there may be no memory bridge 105. Figure 1 For example, the switch 116 may be eliminated, and the network adapter 118 and the add-in cards 120 , 121 would be connected directly to the I / O bridge 107 .
[0038] It should be understood that the core architecture described herein is illustrative and that variations and modifications are possible. Within the scope of the disclosed embodiments, Figure 1 The computer system 100 may include any number of CPUs 102, parallel processing subsystems 112, or memory systems such as system memory 104 and parallel processing memory 134, etc. In addition, as used herein, references to shared memory may include any one or more technically feasible memories, including but not limited to local memory shared by one or more PPUs within a parallel processing subsystem 112, memory shared between multiple parallel processing subsystems 112, cache memory, parallel processing memory 134 and / or system memory 104. Please also note that, as used herein, references to cache memory may include any one or more technically feasible memories, including but not limited to L1 cache, L1.5 cache, and L2 cache. In summary, a person of ordinary skill in the art will understand that Figure 1 The architecture described in this document in no way limits the scope of the various embodiments of the present disclosure.
[0039] Transfers commands and data to and from DRAM via a single clock signal
[0040] Various embodiments include an improved DRAM that uses a single clock to transmit commands and data to and from the DRAM. The single command / data clock in the DRAM can be selected to run at a speed similar to or higher than the high-speed clock of a conventional multiple clock signal high-speed DRAM. Using the disclosed technology, the bits of the command are serialized by the memory controller and sent to the DRAM through a small number of connections to the DRAM command (CA) I / O pins. In some examples, the bits of the command are sent to a single DRAM CA I / O pin through a single connection using the single data / command clock of the DRAM. In order to initialize the DRAM to receive one or more commands, the memory controller sends a synchronization command to the DRAM. The synchronization command establishes a clock edge corresponding to the start of each command, called a command start point. The synchronization command can be in the form of a synchronization signal applied to one or more I / O pins of the DRAM.
[0041] Thereafter, the memory controller sends subsequent instructions to the DRAM according to a predetermined instruction length. The predetermined command length is based on the number of clock cycles required to transmit each command to the DRAM. In other words, the time period between the first command start point and the second consecutive command start point is based on a command length that specifies the total number of parts of the command transmitted in consecutive clock cycles. Adjacent command start points are separated from each other by a predetermined command length. In some examples, the memory controller sends commands to the DRAM via five I / O pins labeled CA[4:0]. The memory controller sends each command within four clock cycles of a high-speed clock signal, with one quarter of the command sent per clock cycle. Therefore, a complete command includes up to 24 bits. In this way, the DRAM avoids the need for a second, lower-speed clock signal to transmit commands to the DRAM.
[0042] Figure 2 is a method for including Figure 1 A block diagram of a clock architecture 200 for memory devices in the system memory 104 and / or the parallel processing memory 134 of the computer system 100 is shown.
[0043] As shown, the clock architecture 200 for the memory device includes a single clock signal WCK 202 that synchronizes various commands transmitted to the memory device. In particular, the WCK 202 clock signal is received by the memory device from the memory controller via a WCK receiver 220 and then sent to various synchronization registers to capture commands and data transmitted to and from the memory device. In this regard, the synchronization register 240 captures data presented on the command (CA) pin 204 via a receiver 222 at the clock edge of the WCK 202 clock signal. After the synchronization register 240 is synchronized, the synchronized CA bit is stored in the command DRAM core 260.
[0044] Similarly, a single clock signal WCK 202 synchronizes the various data transferred to the memory devices. In this regard, the synchronization register 242 captures the main data and extended data (DQ / DQX) bits 206 via the receiver 224 at the clock edge of the WCK 202 clock signal. After the synchronization register 242 is synchronized, the synchronized DQ / DQX bits 206 are stored in the data DRAM core 262. Similarly, the synchronization register 246 captures the error detection and correction data (EDC) bits 208 via the receiver 228 at the clock edge of the WCK 202 clock signal. After the synchronization register 246 is synchronized, the synchronized EDC bits 208 are stored in the data DRAM core 262.
[0045] The single clock signal WCK 202 of the clock architecture 200 for memory devices also synchronizes various data transferred from the memory devices to other devices. In this regard, the synchronization register 244 captures the main data and extended data (DQ / DQX) read from the data DRAM core 262 at the clock edge of the WCK 202 clock signal. After the synchronization register 244 is synchronized, the synchronized DQ / DQX bits 206 are transmitted to the other device via the transmitter 226. Similarly, the synchronization register 248 captures the error detection and correction data (EDC) bits 208 read from the data DRAM core 262 bits at the clock edge of the WCK 202 clock signal. After the synchronization register 248 is synchronized, the synchronized EDC bits 208 are transmitted to the other device via the transmitter 230.
[0046] During a read operation of the DQ / DQX bit 206 and / or the EDC bit 208, the memory device may send a read clock (RCK) signal 210 that is synchronized with the DQ / DQX bit 206 and / or the EDC bit 208 sent by the memory device. In this case, the synchronization register 250 synchronizes the read clock (RCK) generated by the read clock (RCK) generator 264 to be synchronized with the WCK 202. The transmitter 232 sends the synchronized RCK signal 210 to the memory controller. As a result, the RCK signal 210 is synchronized with the DQ / DQX bit 206 synchronized by the synchronization register 244 and / or with the EDC bit 208 synchronized by the synchronization register 248.
[0047] Figure 3 is a method for including Figure 1A more detailed block diagram of a command address clock architecture 300 for memory devices in the system memory 104 and / or parallel processing memory 134 of a computer system 100 is shown. As shown, the command address clock architecture 300 includes an unsynchronized state detection logic 306. The unsynchronized state detection logic 306 detects whether the command pin (CA) interface is synchronous or unsynchronized based on various conditions. In some examples, the unsynchronized state detection logic 306 includes an asynchronous logic circuit that does not receive a clock signal. Additionally or alternatively, the unsynchronized state detection logic 306 includes a synchronous logic circuit that receives a clock signal, such as a version of the WCK 202 clock signal. The unsynchronized state detection logic 306 detects when the memory device attempts to exit from a low power, reset, or CA training state. In response, the unsynchronized state detection logic 306 enables the command start point detection logic 308. After receiving a synchronization command or a command start point command, the memory device synchronizes the synchronization command decoding 314 and / or the clock logic 312 based on the phase of the WCK 202 that received the synchronization command. This condition completes the synchronization process of the CA interface, and the memory device is now ready to accept regular synchronization commands from the memory controller.
[0048] In some examples, the unsynchronized state detection logic 306 detects that the CA interface is unsynchronized. The unsynchronized state detection logic 306 detects this state when the memory device is initially powered on, such as by completely powering off and powering on VPP, VDD, VDDQ, etc. In some examples, the unsynchronized state detection logic 306 detects an assertion followed by a deassertion of the reset (RST) input signal 302. When the unsynchronized state detection logic 306 detects these conditions, the unsynchronized state detection logic 306 determines that the CA interface is unsynchronized. In addition, the memory controller initiates a CA training process to train the unsynchronized CA interface, as described herein. In general, the unsynchronized state detection logic 306 does not determine when the CA training process is required. Instead, the memory controller determines when the CA training process is required. After the CA training process is completed, the unsynchronized state detection logic 306 sends a signal to the command start point detection logic 308 to indicate that the CA interface is now synchronized.
[0049] In some examples, the unsynchronized state detection logic 306 detects that the memory device is recovering from a low power state, such as a power-off state, a self-refresh state, and / or the like, without experiencing a complete power-off and power-on of reset 302 or VPP, VDD, and / or VDDQ. Generally speaking, when the memory device is in a low power state, the memory device powers off one or more receivers that receive external inputs and enters an asynchronous state. In this case, the CA interface may lose synchronization with the memory controller. When the memory device exits from a low power state, a power-off state, a self-refresh state, and / or the like, the CA training process is optional. The memory controller can reestablish synchronization via an asynchronous process without asserting reset 302 or a complete power-off and power-on of VPP, VDD, and / or VDDQ. Through this asynchronous process, the memory device can remove power from the receivers and transmitters of all other I / O pins, including WCK 202, except for the receivers of one or more I / O pins of the memory device involved in the asynchronous process. When resuming from a power-off state or a self-refresh state, the memory controller application and the out-of-sync state detection logic 306 search for a specific value on one or more I / O pins of the memory device with an active receiver. For example, the memory device may maintain a CA 204 command during a power-off or self-refresh state that a receiver on one of the I / O pins is active.
[0050] When recovering from a power-off state or a self-refresh state, the memory controller can apply and the unsynchronized state detection logic 306 can detect a low value on the CA 204 command I / O pin within four consecutive clock cycles of WCK 202. In response, the memory device begins a synchronization phase and waits to receive a synchronization command from the memory controller to establish a new first command starting point. The synchronization command can be in the form of a synchronization signal applied to one or more I / O pins of the memory device. Advantageously, the asynchronous process allows the memory controller to re-establish synchronization with the CA interface without incurring the delay and penalty of performing another CA training process and / or other signal training processes. In contrast, when recovering from a low power state (e.g., a power-off state, a self-refresh state, and / or the like), the memory device quickly resumes synchronous operation with the memory controller. After the asynchronous process is completed, the unsynchronized state detection logic 306 sends a signal to the command starting point detection logic 308 to indicate that the CA interface is now synchronized.
[0051] When the CA interface is not synchronized, the command start point detection logic 308 receives notification from the unsynchronized state detection logic 306. The command start point detection logic 308 receives notification when the memory device exits from a self-refresh state, a power-off state, a CA training operation, a reset, etc. In response, the command start point detection logic 308 begins detecting a specific command start point command received via the CA 204 command I / O pin. After the command start point detection logic 308 receives the command start point, and the command start point is aligned with the memory controller, the command start point detection logic 308 determines that the CA interface is synchronized. The command start point detection logic 308 sends a signal to the command start point generation logic 310 to begin the process of generating a command start point, as described herein.
[0052] The command start point generation logic 310 generates a signal called a command start point, which indicates the start of each command received via the CA204 command I / O pin. The command start point generation logic 310 implements the capture of synchronous multi-cycle commands. The command start point generation logic 310 generates a command start point through various techniques. In some examples, the command start point generation logic 310 includes a counter-based logic that counts "n" phases or cycles of WCK 202, where n is the number of partial command words in each complete command word. The command start point generation logic 310 generates a command start point every n cycles. Additionally or alternatively, the command start point generation logic 310 may include other counter-based logic, clock divider circuits, clock detection logic, etc. In some examples, each command may include four partial command words (n=4), and then the command start point generation logic 310 generates a signal when the first partial command word is present on the CA204 command I / O pin. When the second, third and fourth partial command words are present on the CA 204 command I / O pin, the command start point generation logic 310 does not generate a signal. When the first partial command word of the subsequent command is present on the CA 204 command I / O pin, the command start point generation logic 310 generates a signal again. The command start point generation logic 310 sends the generated command start point to the clock logic 312 and the synchronous command decoding logic 314.
[0053] The clock logic 312 receives the WCK clock signal 202 via the receiver 220 and also receives the command start point from the command start point generation logic 310. In some examples, the clock logic 312 generates a synchronized and divided phase of the WCK 202 to send to the synchronization register 240 so that the synchronization register 240 accurately captures the partial command word received via the CA 204 command I / O pin.
[0054] In various examples, the clock logic 312 may or may not use the command start point indication received from the command start point generation logic 310. In some examples, the memory device captures the state of the CA204 command I / O pin at certain rising and / or falling edges of WCK 202. In such an example, the clock logic 312 does not need to use the command start point to determine when to sample the CA 204 command I / O pin. Instead, only the command deserialization logic and / or the synchronous command decoding logic 314 determine the command start point. The command start point can be determined via a counter that is initially synchronized using the command start point. Once synchronized, the counter will run freely and keep in sync with the memory controller. Additionally or alternatively, the clock logic 312 receives a single command start point to set the phase of the divided clock signal. The clock logic 312 synchronizes the internal clock divider with the single command start point. From then on, the clock logic 312 generates a divided clock signal that continues to keep in sync with the original command start point.
[0055] The synchronous command decoding logic 314 receives a signal from the command starting point generation logic 310 to identify the starting point of each command received via the CA 204 command I / O pin. After the command starting point detection is completed, the synchronous command decoding logic 314 is enabled, indicating that the CA interface has been synchronized. After the CA interface is synchronized, the synchronous command decoding logic 314 can decode synchronous commands received via the CA 204 command I / O pin, including read commands, write commands, activation commands, etc. Additionally or alternatively, after the CA interface is synchronized, the synchronous command decoding logic 314 can decode asynchronous commands received via the CA 204 command I / O pin, including commands without a command starting point. The synchronous command decoding logic 314 sends the decoded commands to the command DRAM core 260.
[0056] Figure 4 is a diagram illustrating that initialization according to various embodiments includes Figure 1 4. A timing diagram 400 of a computer system's system memory 104 and / or a memory device in a parallel processing memory 134 to receive a command.
[0057] The memory device uses a single clock signal scheme to capture both commands and data. The rate of the clock signal is determined by the transfer rate of the highest speed interface of the memory device. Typically, the data interface transfers data at a higher rate than the command interface. However, in some embodiments, the command interface can transfer data at a higher rate than the data interface. The rate of the clock signal rate is set to the transfer rate of the highest speed interface (such as, the data interface). The clock signal is used to transfer data to and from the memory device, typically at a rate of one data transfer per clock cycle.
[0058] The clock signal is also used to transmit commands to the memory device at a lower transmission rate. More specifically, the command is transmitted to the memory device within a plurality of clock cycles of the high-speed clock signal, for example, within four clock cycles. The high-speed clock signal is labeled as WCK 406, which illustrates Figure 2 The command interface includes any number of I / O pins used to transmit commands to the memory device, including Figure 2 CA I / O pin 204. In some embodiments, the command interface includes five I / O pins, labeled CA[4:0], shown as CA[4:1] 408 command I / O pins and CA[0] 410 command I / O pins.
[0059] In some embodiments, each command is transmitted within four clock cycles of WCK 406. References to 0, 1, 2, and 3 represent four phases of command word 412. Within four cycles of WCK 406, a complete command word 412 is transmitted to the memory device in a series of consecutive clock cycles 0, 1, 2, and 3. Therefore, a complete command includes up to 4 clock cycles x 6 bits per clock cycle = 24 bits. Each complete command word 412 represents a command to be executed by the memory device, such as a write operation, a read operation, an activate operation, etc.
[0060] In order to synchronize the transmission of commands to the memory devices, a memory controller (such as the system memory controller 130 or the parallel processing subsystem (PPS) memory controller 132) sends a synchronization (sync) command 418 to the memory devices before transmitting the commands to the memory devices. As shown, the synchronization command 418 is in the form of a synchronization pulse signal received on the CA[0]410 command I / O pin of the memory device. Additionally or alternatively, the synchronization command 418 can be in the form of a synchronization pulse signal received on any other technically feasible input / output pin of the memory device (e.g., one of the CA[4:1]408 command I / O pins). Additionally or alternatively, the synchronization command 418 can be in the form of a synchronization signal received on any technically feasible combination of input / output pins of the memory device, such as two or more CA[4:1]408 and / or CA[0]410 command I / O pins. Additionally or alternatively, the synchronization command 418 can be any signal and / or other indication used by the memory device to identify the phase of the WCK 406 that sets the command start point 414.
[0061] As shown, after receiving the synchronization command 418, the memory device receives a first command start point 414 indicating phase 0 of a first command from the memory controller at four phases of the WCK 406. Additionally or alternatively, after receiving the synchronization command 418, the memory device may receive the first command start point 414 at any technically feasible number of phases of the WCK 406, such as a multiple of four phases, a non-multiple of four phases, and / or less than four phases.
[0062] The synchronization command 418 indicates a valid command starting point 414 for transmitting a command, i.e., which clock edge corresponds to the first part of a multi-cycle command. At some point, the memory device loses synchronization and does not know which clock cycles are valid command starting points 414. For example, the memory device loses synchronization when it is powered on, when it recovers from a reset, when it recovers from a low power state (such as a power-off state or a self-refresh state, etc.). In this case, the memory controller sends a synchronization command 418 to the memory device, which enforces a new command starting point 414 and synchronizes the memory device with the memory controller. Once synchronized, the memory device can begin to accept commands from the memory controller.
[0063] More specifically, the memory device may be powered on when VPP, VDD, and VDDQ 402 are applied to the memory device, where VPP is a pump voltage, VDD is a main power supply voltage, and VDDQ is an I / O voltage. The memory controller applies a low voltage to the reset 404 input of the memory device to place the memory device in a reset state. Subsequently, the memory controller applies a high voltage to the reset 404 input of the memory device to bring the memory device out of a reset state. Prior to applying the high voltage to the reset 404 input, the memory controller may apply a fixed bit pattern to the CA[4:1] 408 and CA[0] 410 command I / O pins of the memory device. The fixed bit pattern is referred to herein as a “strap”. The memory device samples the state of the strap at the rising edge of reset 404 to determine the value of the fixed bit pattern. Based on the fixed bit pattern, the memory device may undergo certain boot procedures, such as an optional command pin (CA) training 416 process, to command the memory device to determine the deviation between the WCK 406 and the CA[4:1] 408 and CA[0] 410 command I / O pins. The memory controller completes the boot procedure via an asynchronous communication sequence with the memory device, such as the optional CA training 416 process. The optional CA training 416 process determines the optimal skew of the CA[4:1] 408 and CA[0] 410 command I / O pins relative to the WCK 406 to ensure that the CA[4:1] 408 and CA[0] 410 command I / O pins meet the setup and hold time requirements. The optional CA training 416 process further detects and corrects any multi-cycle skew between any two or more command I / O pins to ensure that all command I / O pins capture the command bits of the same command word 412 at the same rising or falling edge of the WCK 406.
[0064] After completing the optional CA training 416 process, the memory device is in a state where it can receive commands synchronously with respect to the rising and / or falling edges of WCK 406. Alternatively, if the memory controller and memory device do not perform the optional CA training 416 process, the memory device is ready to receive commands synchronously at any time after the rising edge of reset 404. In either case, the memory controller receives the command on one of the command I / O pins ( Figure 4410 command I / O pin). The memory controller applies phase 0 of the first command word 412 to the CA[4:1]408 and CA[0]410 command I / O pins. The memory controller applies phase 0 of the first command word 412 so that it is valid at the fourth rising edge of WCK 406 after the trailing edge of the synchronization command 418. The memory controller applies phase 1, 2, and 3 of the first command word 412 so that it is valid at consecutive rising edges of WCK 406. The memory device samples the four phases of the first command word 412 on CA[4:1] 408 and CA[0] 410 on these same four rising edges of WCK 406. The first rising edge of WCK 406 after phase 3 of the first command word 412 indicates a second command start point 414. The memory controller applies and the memory device transmits the four phases 0, 1, 2, 3 of the second command word 412 on four consecutive rising edges of WCK 406 starting from the second command start point 414. The first rising edge of WCK 406 after phase 3 of the second command word 412 indicates a third command start point 414, and so on.
[0065] In some embodiments, the memory device can be recovered from a power-off state, a self-refresh state, and / or the like without undergoing a complete power-off and power-on of reset 404 or VPP, VDD, VDDQ 402. In this case, the memory device may lose synchronization with the memory controller. In this case, the memory controller can reestablish synchronization via an asynchronous process without asserting a complete power-off and power-on of reset 404 or VPP, VDD, VDDQ 402. Through the asynchronous process, the memory device can remove power from the receivers and transmitters of all other I / O pins, including WCK 406, except for the receivers of one or more I / O pins of the memory device involved in the asynchronous process. When recovering from a power-off state or a self-refresh state, the memory controller applies and the memory device searches for a specific value on one or more I / O pins of the memory device with an active receiver. For example, the memory device can keep the receiver of the CA[0]410 command I / O pin active during a power-off or self-refresh state. When recovering from a power-off or self-refresh state, the memory controller may apply and the memory device may detect a low value on the CA[0]410 command I / O pin within four consecutive clock cycles of WCK 406. In response, the memory device begins a synchronization phase and waits to receive a synchronization command 418 from the memory controller to establish a new first command start point 414. The synchronization command 418 may be in the form of a synchronization signal applied to one or more I / O pins of the memory device. Advantageously, this asynchronous process allows the memory controller to reestablish synchronization with the memory device without incurring the delays and penalties of performing another optional CA training 416 process and / or other signal training processes. In contrast, when recovering from a low power state (e.g., a power-off state, a self-refresh state, and / or the like), the memory device quickly resumes synchronous operation with the memory controller.
[0066] Figure 5 is a diagram illustrating the transmission of successive commands to a Figure 1 A timing diagram 500 of memory devices in the system memory 104 and / or the parallel processing memory 134 of a computer system.
[0067] As shown, the high speed clock signal is a single clock signal for both commands and data, labeled WCK 406, and illustrates Figure 2 The command interface includes any number of I / O pins used to transmit commands to the memory device, including Figure 2 CA I / O pin 204. In some embodiments, the command interface includes five I / O pins, labeled CA[4:0] 502, and is Figure 4The CA[4:1] 408 and CA[0] 410 command I / O pins are shown separately from the same command I / O pins. In some embodiments, the command bits CA[4:0] 502 may be encoded via a non-return-to-zero (NRZ) data signaling mode.
[0068] Figure 5 4 shows five command start points 414, each of which coincides with a rising edge of WCK 406, which coincides with phase 0 of a four-phase command. Three consecutive phases 1, 2, 3 of the command coincide with three consecutive rising edges of WCK 406. The clock rising edge of WCK 406 following phase 3 of the command is followed by a command start point 414 for phase 0 of the following command.
[0069] The data transmitted to and from the memory device may include main data bits (DQ), extended data bits (DQX), and error detection bits (EDC). The error detection bits are used to detect and / or correct bit errors in the main data bits and / or extended data bits via any technically feasible error detection and correction code (e.g., cyclic redundancy check (CRC) code).
[0070] The memory device can adopt a variety of data signaling modes based on different data transmission modes. For example, the DQ and EDC data bits can adopt a redundant data strobe (RDQS) data transmission mode, as shown in the DQ / EDC 504 timing diagram. In this case, the DQ and EDC data bits can be encoded via an NRZ data signaling mode. In the RDQS data transmission mode, at each rising edge and each falling edge of WCK406, the data is sent to and from the memory device as a one-bit symbol captured at twice the command phase rate. Therefore, each DQ and EDC symbol includes one bit of data. Additionally or alternatively, the data sent to and from the memory device can adopt a data transmission mode that transmits symbols including two or more bits of data. In one example, the DQ, DQX, and EDC data bits can be encoded with symbols carrying more than one bit of data via a high-speed multi-level mode. One such data transmission mode is a 4-level pulse amplitude modulation (PAM4) data transmission mode using two-bit symbols, as shown in the DQ / DQX / EDC 506 timing diagram. In PAM4 mode, data is sent to and from the memory device as two-bit symbols captured at twice the rate of the command phase at each rising edge and each falling edge of WCK 406. The PAM4 data transmission mode allows each data I / O pin to carry two bits of data captured at each rising edge and each falling edge of WCK 406. Therefore, in PAM4 data transmission mode, the data transmission rate is four times the command transmission rate. Whether the memory device operates in RDQS mode, PAM4 mode, or any other data transmission mode, the same clock signal WCK 406 captures both command bits and data bits.
[0071] It should be understood that the system shown herein is illustrative and variations and modifications are possible. A single command word may include multiple groups of four phases. In some examples, a single command word may include multiples of four phases, such as eight phases, twelve phases, etc. In such examples, each command is sent via a CA[4:0]I / O pin through a plurality of four-phase commands. For a single command including eight phases, the command is sent as two consecutive four-phase commands. When the memory controller sends the first four-phase command to the memory device, the memory device recognizes that the command is an eight-phase command. The memory device receives the first four phases of the command starting from a certain command starting point 414, and receives the last four phases of the command starting from the next consecutive command starting point 414. Similarly, for a single command including twelve phases, the command is sent as three consecutive four-phase commands. When the memory controller sends the first four-phase command to the memory device, the memory device recognizes that the command is a twelve-phase command. The memory device receives the first four phases of a command starting from a certain command starting point 414, and receives the second four phases and the third four phases of a command starting from the next two consecutive command starting points 414, and so on.
[0072] In another example, a command transmitted by a memory controller to a memory device is described as up to 24 command bits, which are sent as four phases of five bits. However, within the scope of the disclosed embodiments, the number of phases may be more than four phases or less than four phases. Furthermore, within the scope of the disclosed embodiments, the number of command bits may be more than five bits or less than five bits. In yet another example, the signals disclosed herein are described in terms of rising and / or falling edges, high or low levels, etc. However, rising and falling edges may be interchangeable, high and low levels may be interchangeable, and any other technically feasible changes may be made with respect to signal edges and levels within the scope of the disclosed embodiments.
[0073] Figure 6 is a method for transmitting commands to a Figure 1 Flowchart of method steps for storing memory devices in system memory 104 and / or parallel processing memory 134 of a computer system. Figure 1-4 Although the method steps are described with reference to a system, one of ordinary skill in the art will understand that any system configured to perform the method steps in any order is within the scope of the present disclosure.
[0074] As shown, method 600 begins at step 602, where the memory device receives a synchronization command 418 on an input of the memory device. In order to synchronize the transmission of commands to the memory device, a memory controller such as system memory controller 130 or parallel processing subsystem (PPS) memory controller 132 sends a synchronization command 418 to the memory device before transmitting the command to the memory device. The synchronization command can be in the form of a synchronization signal applied to one or more I / O pins of the DRAM. The synchronization command 418 indicates a valid command starting point 414 for transmitting the command, i.e., which clock edge corresponds to the first part of the multi-cycle command. At some point, the memory device loses synchronization and does not know which clock cycles are valid command starting points 414. For example, the memory device loses synchronization when it is powered on, when it is restored from a reset, when it is restored from a low power state (such as a power-off state or a self-refresh state, etc.). In this case, the memory controller sends a synchronization command 418 to the memory device, which enforces a new command starting point 414 and synchronizes the memory device with the memory controller. Once synchronized, the memory device can begin to accept commands from the memory controller.
[0075] More specifically, the memory device may be powered on when VPP, VDD, and VDDQ 402 are applied to the memory device, where VPP is a pump voltage, VDD is a main power supply voltage, and VDDQ is an I / O voltage. The memory controller applies a low voltage to the reset 404 input of the memory device to place the memory device in a reset state. Subsequently, the memory controller applies a high voltage to the reset 404 input of the memory device to take the memory device out of the reset state.
[0076] At step 604, the memory device synchronizes to the clock edge based on the synchronization command 418. When the memory device receives the synchronization command 418, the memory device counts the number of rising edges or falling edges of the high-speed clock WCK 406 starting from the leading edge or trailing edge of the synchronization command 418. The high-speed clock WCK 406 is the same clock used by the memory device to receive and send data. In some examples, the memory device counts four rising edges of WCK 406 after the falling edge of the synchronization command 418 to determine the first command starting point 414.
[0077] At step 606, the memory device receives the first portion of the command, Phase 0, on the WCK 406 clock edge determined at step 604. The memory controller in turn applies Phase 0 of the first command word 412 to the CA[4:1] 408 and CA[0] 410 command I / O pins. The memory controller applies Phase 0 of the first command word 412 so that it is valid on the fourth rising edge of WCK 406 following the falling edge of the Sync Command 418.
[0078] At step 608, the memory device receives additional portions of the command, phases 1, 2, and 3, on successive WCK 406 clock edges following the clock edge determined at step 604. The memory controller applies phases 1, 2, and 3 of the first command word 412 to be valid on successive rising edges of WCK 406. The memory device samples the four phases of the first command word 412 on CA[4:1] 408 and CA[0] 410 on these same four rising edges of WCK 406.
[0079] At step 610, the memory device receives the portion of the additional command on successive WCK 406 clock edges after the clock edge of phase 3 of the first command. The first rising edge of WCK 406 after phase 3 of the first command word 412 indicates the second command starting point 414. The memory controller applies and the memory device transmits the four phases 0, 1, 2, 3 of the second command word 412 on four consecutive rising edges of WCK 406 starting from the second command starting point 414. The first rising edge of WCK 406 after phase 3 of the second command word 412 indicates the third command starting point 414, and so on.
[0080] The method 600 then terminates. Alternatively, the method 600 proceeds to step 610 to transmit additional commands to the memory device. Thus, by repeatedly transmitting commands to the memory device in the manner described, commands and data may be transmitted to and from the memory device via a single high-speed clock signal. If the memory device subsequently loses synchronization, such as when powered on, when recovering from a reset, when recovering from a low-power state (e.g., a powered-off state or a self-refresh state, etc.), the method 600 proceeds to step 602 to begin synchronization again.
[0081] In summary, various embodiments include an improved DRAM that uses a single clock to transmit commands and data to and from the DRAM. The single command / data clock in the DRAM can be selected to run at a speed similar to or higher than the high-speed clock of a conventional multiple clock signal high-speed DRAM. Using the disclosed technology, the bits of the command are serialized by the memory controller and sent to the DRAM through a small number of connections to the DRAM command (CA) I / O pins. In some examples, the bits of the command are sent to a single DRAM CA I / O pin through a single connection using the single data / command clock of the DRAM. In order to initialize the DRAM to receive one or more commands, the memory controller sends a synchronization command to the DRAM. The synchronization command establishes a clock edge corresponding to the start of each command, called a command start point. The synchronization command can be in the form of a synchronization signal applied to one or more I / O pins of the DRAM.
[0082] Thereafter, the memory controller sends subsequent commands to the DRAM according to a predetermined command length. The predetermined command length is based on the number of clock cycles required to transmit each command to the DRAM. Adjacent command start points are separated from each other by a predetermined command length. In some examples, the memory controller sends commands to the DRAM via five I / O pins labeled CA[4:0]. The memory controller sends each command within four clock cycles of a high-speed clock signal, with one quarter of the command sent per clock cycle. Therefore, a complete command includes up to 24 bits. In this way, the DRAM avoids the need for a second, lower-speed clock signal to transmit commands to the DRAM.
[0083] At least one technical advantage of the disclosed technology over the prior art is that, using the disclosed technology, commands and data are received by the memory device via a single clock signal at different transfer rates. As a result, the memory device does not require internal synchronization and training circuits to resolve possible deviations between multiple clock signals. Another advantage of the disclosed technology is that only one receiver and I / O pin are required to receive the clock signal, rather than two receivers and I / O pins. As a result, the complexity, surface area, and power consumption of the internal circuitry of the DRAM die can be reduced relative to methods involving multiple clock signals. In addition, the I / O pin previously used to receive the second clock signal can be used for another function, such as an additional command bit, data bit, or control signal. These advantages represent one or more technical improvements over the prior art methods.
[0084] Any and all combinations of any claim elements recited in any claim and / or any elements described in this application, in any manner, are within the intended scope of this disclosure and protection.
[0085] The description of various embodiments has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments.
[0086] Aspects of the present embodiment may be embodied as systems, methods, or computer program products. Therefore, aspects of the present disclosure may take the form of complete hardware embodiments, complete software embodiments (including firmware, resident software, microcode, etc.), or embodiments combining software and hardware aspects, which may be collectively referred to herein as "modules" or "systems." In addition, aspects of the present disclosure may take the form of a computer program product contained in one or more computer-readable media, which has a computer-readable program code contained thereon.
[0087] Any combination of one or more computer-readable media may be used. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. For example, a computer-readable storage medium includes, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination thereof. More specific examples (non-exhaustive) of computer-readable storage media may include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the context of this document, a computer-readable storage medium may be any tangible medium that may include or store a program for use by or in conjunction with an instruction execution system, device, or apparatus.
[0088] Aspects of the present disclosure are described with reference to flowcharts and / or block diagrams of methods, devices (systems) and computer program products according to embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram, as well as the combination of boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to generate a machine so that when the instructions are executed by the processor of the computer or other programmable data processing device, the functions / actions specified in one or more boxes of the flowchart and / or block diagram can be implemented. Such processors can be, but are not limited to, general-purpose processors, special-purpose processors, application-specific processors, or field programmable gate arrays.
[0089] The flowchart and block diagram in the figure show the possible architecture, function and operation of the system, method and computer program product according to each embodiment of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, segment or part of the code, and the code includes one or more executable instructions for implementing one or more specified logical functions. It should also be noted that in some alternative embodiments, the functions indicated in the box may not occur in the order indicated in the figure. For example, the two boxes shown in succession can actually be executed roughly at the same time, or these boxes can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented by a system based on special-purpose hardware, which performs a specified function or action or a combination of special-purpose hardware and computer instructions based on the system based on special-purpose hardware.
[0090] While the foregoing is directed to embodiments of the present disclosure, other and further embodiments of the disclosure may be devised without departing from the basic scope thereof, the scope of which is determined by the appended claims.
Claims
1. A computer-implemented method for transmitting a command to a memory device, the method comprising: receiving a synchronization signal at an input pin of the memory device, wherein the synchronization signal indicates a starting point of a first command; synchronizing the memory device to a first clock edge of a clock signal input relative to the synchronization signal; receiving a first portion of the first command at the first clock edge; as well as A second portion of the first command is received at a second clock edge of the clock signal input subsequent to the first clock edge.
2. The computer-implemented method of claim 1 , further comprising: Based on the synchronization signal, establishing a first command starting point at the first clock edge; as well as A second command start point is established at a third clock edge of the clock signal input following the second clock edge.
3. The computer-implemented method of claim 2, wherein a time period between the first command starting point and the second command starting point is based on a command length that specifies a total number of parts of the first command including the first part and the second part.
4. The computer implemented method of claim 3, wherein the command length comprises four clock cycles of the clock signal input.
5. The computer implemented method of claim 1, wherein the first portion of the first command and the second portion of the first command are received via a plurality of command input pins.
6. The computer implemented method of claim 5, wherein the synchronization signal is received via a first command input pin included in the plurality of command input pins.
7. The computer-implemented method of claim 1 , further comprising: One or more data bits associated with the first command are received at a third clock edge of the clock signal input following the second clock edge.
8. The computer-implemented method of claim 7, further comprising: A first portion of a second command is received at the third clock edge.
9. The computer-implemented method of claim 1, wherein the first clock edge comprises a fourth rising clock edge of the clock signal input after a falling edge of the synchronization signal.
10. The computer implemented method of claim 1, wherein the synchronization signal is received after recovery from at least one of a power down state, a reset state, or a self-refresh state.
11. A computer-implemented system comprising: Memory controller; and A memory device coupled to the memory controller, which: receiving a synchronization signal at an input pin of the memory device, wherein the synchronization signal indicates a starting point of a first command; synchronizing the memory device to a first clock edge of a clock signal input relative to the synchronization signal; receiving a first portion of the first command at the first clock edge; and A second portion of the first command is received at a second clock edge of the clock signal input subsequent to the first clock edge.
12. The system of claim 11, wherein the memory device further: establishing a first command starting point at the first clock edge based on the synchronization signal; and A second command start point is established at a third clock edge of the clock signal input following the second clock edge.
13. The system of claim 12, wherein a time period between the first command start point and the second command start point is based on a command length indicating a total number of parts of the first command including the first part and the second part.
14. The system of claim 13, wherein the command length comprises four clock cycles of the clock signal input.
15. The system of claim 11, wherein the first portion of the first command and the second portion of the first command are received via a plurality of command input pins.
16. The system of claim 15, wherein the synchronization signal is received via a first command input pin included in the plurality of command input pins.
17. The system of claim 11, wherein the memory device further receives one or more data bits associated with the first command at a third clock edge of the clock signal input subsequent to the second clock edge.
18. The system of claim 17, wherein the memory device further receives the first portion of the second command at the third clock edge.
19. The system of claim 11, wherein the first clock edge comprises a fourth rising clock edge of the clock signal input after a falling edge of the synchronization signal.
20. The system of claim 11, wherein the synchronization signal is received after recovery from at least one of a power-off state, a reset state, or a self-refresh state.
Citation Information
Patent Citations
Method and apparatus for decoding commands
US20170110173A1