High-Performance Non-Volatile Memory Module
By adopting a point-to-point system architecture in the memory system, and using buffers and control logic to achieve interleaved access to non-volatile memory modules and DRAM memory modules, the problems of module capacity reduction and signal rate reduction in the prior art are solved, and the capacity and performance of the memory system are maximized.
Patent Information
- Application Number
- CN202111250123.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2015-09-22
- Filing Date
- 2016-03-11
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2036-03-11
AI Technical Summary
While the existing memory systems increase the signaling rate, the module capacity continues to decline, resulting in a decrease in signal integrity and a decrease in signal rate, limiting the expansion capability of the memory system and may be limited to single-row devices on a single module in the future.
Using a point-to-point system architecture, by introducing buffers and control logic into the memory module, non-volatile memory modules and DRAM memory modules are allowed to be interleaved, achieving high aggregated data transmission bandwidth, and supporting a variety of configurable transmission solutions.
Through the point-to-point system architecture, the capacity and performance of the memory system are maximized, the problems of reduced signal integrity and signal rate are avoided, and the coordinated work of multiple modules is supported, which improves the system's expansion capabilities.
Smart Images

Figure CN113806243B_ABST
Abstract
Description
[0001] This application is a divisional application of the application with the application date of March 11, 2016, application number 201680010277.X, and invention title "High-Performance Non-Volatile Memory Module". Technical Field
[0002] Disclosed herein are memory modules, memory controllers, memory devices, and related methods. Background Art
[0003] Successive generations of dynamic random access memory components (DRAMs) have emerged on the market, having continuously shrinking lithographic feature sizes. As a result, the storage capacity of devices from each generation has increased. In addition, due to improvements in transistor performance, the interface signaling rate of each generation has also increased.
[0004] Unfortunately, one metric in memory system design that has not shown significant improvement is the module capacity of a standard memory channel. As the signaling rate has increased, this capacity has been continuously decreasing. Part of the reason is the link topology used in standard memory systems. When more modules are added to the system, signal integrity decreases, and the signal rate must necessarily decrease. When operating at the maximum signaling rate, today's typical memory systems are limited to one or two modules.
[0005] Unless improvements are made, at the maximum signaling rate, future memory systems may be limited to a single rank of devices (or a stack of single rank of devices) on a single module. Brief Description of the Drawings
[0006] Embodiments of the present disclosure are shown in the drawings in an illustrative rather than restrictive manner, and like reference numerals in the drawings denote like elements, and in the drawings:
[0007] Figure 1 An embodiment of a memory system employing a memory controller, a non-volatile memory module, and a DRAM memory module is shown.
[0008] Figure 2 Shown is Figure 1 An embodiment of the non-volatile memory module shown
[0009] Figure 3 Shown is Figure 2 An embodiment of the steering logic employed in the non-volatile memory module buffer circuit of
[0010] Figure 4 Shown is Figure 1 An embodiment of the DRAM memory module of
[0011] Figure 5A A flowchart of an embodiment of a method for reading data from a non-volatile memory module is shown from the perspective of a memory controller.
[0012] Figure 5B A flowchart of an embodiment of a method for writing data to a non-volatile memory module is shown from the perspective of a memory controller.
[0013] Figure 6A An embodiment of a flowchart showing a read data transfer from a non-volatile memory module is shown from the perspective of the non-volatile memory module.
[0014] Figure 6B An embodiment of a flowchart showing a write data transfer to a non-volatile memory module is shown from the perspective of the non-volatile memory module.
[0015] Figure 7A An embodiment of a timing diagram related to the read data transfer of Figure 6A is shown.
[0016] Figure 7B An embodiment of a timing diagram related to the write data transfer of Figure 6B is shown.
[0017] Figure 8 A block diagram showing a read data transfer involving a non-volatile memory module and a DRAM memory module is shown, where each module is allocated half of the system bandwidth.
[0018] Figure 9 An embodiment of a timing diagram related to the read data transfer of Figure 8 is shown.
[0019] Figure 10 A block diagram showing a read data transfer involving a non-volatile memory module and a DRAM memory module similar to Figure 8 is shown, where each module is allocated half of the system bandwidth.
[0020] Figure 11 An embodiment of a timing diagram related to the read data transfer of Figure 10 is shown.
[0021] Figure 12 A block diagram showing a read data transfer similar to Figure 8 and Figure 10 is shown, but where the entire system bandwidth is allocated to the DRAM memory module.
[0022] Figure 13 An embodiment of a timing diagram related to the read data transfer of Figure 12 is shown.
[0023] Figure 14 A block diagram showing a read operation from a DRAM module and a read operation from a non-volatile memory module, where the non-volatile memory module directly transfers write data to the DRAM module during a separate write operation.
[0024] Figure 15 Shows the related to Figure 14 The timing diagram of data transfer.
[0025] Figure 16 A block diagram showing a non-volatile memory module and a buffered DRAM memory module, and where the entire system bandwidth is allocated to the non-volatile memory module.
[0026] Figure 17 Shows the related to Figure 16 The timing diagram of read data transfer.
[0027] Figure 18 Shows an alternative system arrangement having two buffered DRAM modules and a non-volatile memory module. Detailed Description
[0028] Memory modules, memory controllers, devices, and related methods are disclosed. In one embodiment, a memory module is disclosed that includes a pin interface for coupling to a bus. The bus couples the module to a memory controller. The module includes at least two non-volatile memory devices and a buffer disposed between the pin interface and the at least two non-volatile memory devices. The buffer receives non-volatile memory access commands interleaved with DRAM memory module access commands from the memory controller. This allows for a point-to-point system architecture that can use non-volatile memory modules and DRAM memory modules together to maximize capacity and performance.
[0029] Referring to Figure 1 , one embodiment of a memory system generally designated 100 employs a plurality of memory modules 102 and 104 coupled to a memory control circuitry 110 via point-to-point signaling links 106 and 108. Modules 102 and 104 can be of the same or different types, such as DRAM memory modules or non-volatile memory modules. The architecture described herein allows for the mixing of different module types in a point-to-point topology to maximize memory capacity and performance.
[0030] Continuing to refer to Figure 1, a particular embodiment of the memory control circuit device 110 may include, for example, a discrete memory controller separate from the requester integrated circuit (IC), or any IC that controls memory devices such as DRAM and non-volatile memory, and may be any type of system-on-chip (SoC). One embodiment of the memory control circuit device 110 uses the interface 112 to send signals to and receive signals from the memory modules 102 and 104. The write data signals sent through the interface may be protected by error detection and correction (EDC) bits encoded by the write error detection and correction (EDC) encoder 114. The write EDC encoder 114 generates error information associated with the write data symbols, such as EDC parity bits. The error coding may be generated according to one of a plurality of acceptable EDC algorithms, and acceptable EDC algorithms include, for example, a simple one-bit Hamming code, a more complex high-speed BCH (Bose, Ray-Chaudhuri, and Hocquenghem) code. A particular error code applicable to the embodiments described herein is the 64 / 72 error detection and correction code. Other EDC codes such as Reed-Solomon codes, turbo codes, cyclic redundancy codes (CRC), and low-density parity-check (LDPC) codes are also acceptable. The memory control circuit device 110 includes a read EDC decoder 116 for decoding the error information associated with the input read data symbols from the memory modules 102 and 104. The three-level cache 118 connects the memory control circuit device to the host processing resources (not shown).
[0031] Figure 2 A particular embodiment of a non-volatile memory module generally designated 200 is shown, which may be adapted to be included in Figure 1 system 100. The non-volatile memory module 200 includes a substrate 202, and the substrate mounts multiple sets of components, for example, at 204 (shown in dashed lines), to achieve the desired module bandwidth in a point-to-point memory system with similar or different memory modules. A more detailed view of one of the groups of components is shown at 206, and it should be understood that each group has the same structure. With this in mind, each group includes data buffer components DB 208i (nine groups are shown here, and "i" ranges from 1 to 9), and the data buffer components are connected to the memory control circuit device 110 via the first primary DQ nibble group DQu ( Figure 1 ). The D-group buffer components are also connected to the primary nibble group DQt shared with another memory module. For one embodiment, each data nibble group includes four data DQ links and differential strobe DQS links (not shown). The secondary data DQ nibble group DQn couples each data buffer component 208i to a group of non-volatile memory devices 210. Although Figure 2 the non-volatile memory module 200 is shown using nine data buffer components DB 2081 -DB 208 9 (to accommodate data transfers protected by error code used by the DRAM memory module as well), however, the buffer components may alternatively be incorporated into a smaller number of broader components (such as three components, each having, for example, six primary nibble interfaces).
[0032] Further reference Figure 2 , for a specific example, a set of non-volatile memory devices 210 includes a stack of four non-volatile memory dies. Each stack may contain eight non-volatile memory components. The interfaces of each non-volatile memory component may be connected in parallel using through-silicon vias or any other connection method. Other stack configurations are also possible. An example of a stack of devices is shown in an enlarged view Figure 2-2 showing the stacked components 212 within a single encapsulation 214. For some configurations, the opposite side of the module substrate 202 may mount the memory components such as at 216.
[0033] Continuing reference Figure 2 , the non-volatile memory module 200 includes a control / address (CA) buffer component RCD that drives an intermediate CAi link connected to each data buffer component such that at 218, each data buffer component drives a secondary CAn link to each non-volatile memory stack. In an alternative embodiment, the CA buffer may directly drive the secondary CAn links to each non-volatile memory stack.
[0034] In an alternative embodiment, the non-volatile memory module 200 may further include a DRAM component. The data buffer DB and CA buffer RCD components on the module will allow operation of the DRAM component as described above (as on a traditional DRAM DIMM module) or operation of the NVM component.
[0035] Figure 3 Shows further details of a specific embodiment of a data buffer component suitable for inclusion in a Figure 1 non-volatile memory module. Generally, the data buffer includes control logic 300 that manages non-volatile memory components connected to secondary data DQn and control / address CAn links. The control logic 300 may manage concurrent transactions to more than one non-volatile memory component. This concurrency allows the module to achieve a high aggregate data transfer bandwidth.
[0036] Further reference Figure 3, the data buffer component includes two primary nibble interfaces DQa and DQb, and each primary nibble interface has independent receiver and transmitter logic circuits 302 and 304 coupled to each nibble interface. The first independent logic circuit 302 for the first primary nibble DQa includes a receiving amplifier 30 that feeds a sampler 308. The output of the sampler is then routed to a transmit multiplexer 312 and a secondary interface logic circuit multiplexer 310 in the second primary nibble logic circuit 304. The first independent logic circuit 302 also includes a transmit logic path using a transmit multiplexer 314. The transmit multiplexer 314 selects between a first input connected to the output of an SRAM memory 330 associated with a second logic circuit 332 for a secondary nibble interface DQn and a second input coupled to the sampler 316 from the second independent logic circuit 304. The output of the multiplexer 314 feeds a phase and period adjustment circuit 318 and 320, which are coupled to a transmit amplifier 322.
[0037] Continuing reference Figure 3 , the second independent logic circuit 304 associated with the nibble interface DQb is similar to the first independent logic circuit 302. A receiving amplifier 324 feeds a sampler 316, and the sampler 316 feeds its sampled output to one input of the transmit multiplexer 314 of the first logic circuit and an input of the transmit multiplexer 310 of the secondary logic circuit 332. The second logic circuit also includes a transmit logic path using a transmit multiplexer 312. The transmit multiplexer 312 selects between a first input connected to the output of an SRAM memory 330 associated with the secondary logic circuit 332 and a second input coupled to the sampler 308 from the first independent logic circuit 302. The output of the multiplexer 312 feeds a phase and period adjustment circuit 326 and 328, which are coupled to a transmit amplifier 329.
[0038] Further reference Figure 3 , the secondary logic circuit 332 includes a read path using an amplifier 334 that feeds a sampler 336. The output of the sampler is fed to an SRAM memory 330 that serves as a temporary memory for the read data. For a particular embodiment, the read SRAM 330 can be organized into data rows or blocks of 2KB, with each data column including 64 bits, as shown at 338. The bank / row / column address tags at 340 are for columns. With this arrangement, the read data can be received as block data in the SRAM from a non-volatile device (usually via a non-volatile memory), aggregated in the SRAM, and then retrieved as column data from the read SRAM 330 and fed to either or both of the transmit multiplexers 314 and 312.
[0039] Continuing reference Figure 3, the secondary logic circuit 332 further includes a write path that includes a transmit multiplexer 310 that selects between the outputs of sampler 308 (associated with DQa) or sampler 316 (associated with DQb). Then at 344, the multiplexer output is fed to a temporary write SRAM storage device 342 with a corresponding write index. The write SRAM storage device 342 is organized in a manner similar to the read SRAM memory 330, except that write data is received as column data within the SRAM, aggregated in the SRAM, and then distributed as block data to the non-volatile memory. The output of the write SRAM is fed to corresponding phase and cycle adjustment circuits 346 and 348 and then driven towards the non-volatile memory via a transmit amplifier 350.
[0040] The SRAM storage device allows concurrent transactions to occur such that access to two different non-volatile memory components connected to the same data buffer can overlap. The storage device also allows parallel transactions to synchronize across all data buffer components. This feature is provided because non-volatile memory modules can use the same error detection and correction code (EDC) as that used by DRAM memory modules. Thus, the access granularity is a multiple of nine rather than a power of two. This allows the 9 / 8 transfers and storage device overhead required by standard syndromes (such as ECC, Chipkill, etc.). Thus, non-volatile memory transactions will involve transfers between nine non-volatile memory components and nine data buffer components on the module. This transfer will have a "block" granularity - typically 2KB / 4KB / 8KB per non-volatile memory component. Since nine DB components operate in parallel, the overall transfer granularity will be 18KB / 36KB / 72KB. The requester in the controller will see block sizes of 16KB / 32KB / 64KB because EDC syndromes will be generated and checked in the controller interface. This block size can be comparable to the row size of a DRAM module (with 18 DRAMs operating in parallel).
[0041] Once the non-volatile memory data block is transferred to the temporary SRAM memory in the DB component, it can be accessed in column blocks (the same column access granularity as a DRAM module). Once the block read has moved the block data from nine non-volatile memory components to the SRAM memory of nine DB components, the controller can perform column read accesses. These column read accesses can transfer all or just part of the block data from the SRAM memory to the controller. Assuming a 64B column block and non-volatile memory data block sizes of 2KB / 4KB / 8KB, typically 512 / 1024 / 2048 column accesses are required to transfer the data block between the SRAM memory and the controller.
[0042] If the controller wants to perform a column write access, it typically transfers all the block data from the controller to the SRAM memory (one column block at a time) before performing a block write access to transfer the block from the SRAM memory to the nine non-volatile memory components. If the controller only wants to write a part of the block, the block needs to be first read from the nine non-volatile memory components into the SRAM memory, a column write access is performed on the SRAM memory, and then a block write access is performed to transfer the modified block from the SRAM memory to the nine non-volatile memory components. This is also referred to as a read-modify-write transaction.
[0043] In some cases, it is desirable to direct the data received at one primary DQ (e.g., DQb) to another primary DQ (e.g., DQa), thus bypassing the secondary interface circuit 332. This can be achieved by using control signals applied to the appropriate multiplexers (here, by enabling the transmit multiplexer 314 to pass the output of the sampler 316 and disabling the multiplexer 310 of the secondary interface, the data can be transferred from DQb to DQa).
[0044] The buffer logic circuit also provides a pipeline latency for column data access that matches the pipeline latency associated with the DRAM module. Status bits generated in the logic generate a status return signal for one of the following conditions: (1) enabling parallel access to the non-volatile memory; (2) accommodating variable non-volatile memory access; and (3) accommodating a larger non-volatile memory access granularity.
[0045] The receive-transmit path of the data buffer component also provides a function to change the mode (adjust the phase relationship between the timing signals) with respect to the time domain. Most DB components operate in the clock domain created by a CLK link (not shown) associated with the CA bus. A small part of the interface operates within the domain of the DQS timing signal (not shown) received for the DQa interface. The buffer includes cross-domain logic to perform a domain crossing between the two time domains.
[0046] Figure 4 A particular embodiment of a DRAM memory module, generally designated 400, is shown, which is applicable to Figure 1A point-to-point memory system is provided such that it can be combined with another DRAM memory module or a non-volatile memory module 200 such as the above. The DRAM memory module 400 can be of the registered dual inline memory module (RDIMM) type and includes a substrate 402 on which multiple component groups are mounted (in a dashed manner) at, for example, 404 to achieve the desired module bandwidth for point-to-point in a point-to-point memory system with similar or different memory modules. A more detailed view of one of the component groups is shown at 406, and it should be understood that each group has the same structure. With this in mind, each group is connected to the memory control circuitry 110 ( Figure 1 ) via a primary DQ nibble group (e.g., DQv). Another primary nibble group DQt allows the module to be connected to the memory control circuitry (for a single module configuration) or to another module as a shared data path. For one embodiment, each data nibble group includes four data DQ links and differential strobe DQS links (not shown).
[0047] Further referring to Figure 4 , for a particular instance, each group of devices includes four stacks of DRAM memory modules 408, 410, 412, and 414. Each stack can include eight DRAM memory components. An example of a stacked device group is shown in an enlarged view Figure 4-4 showing the stacked components 416 within a single encapsulation 418. For some configurations, memory components can be mounted on opposite sides of the module substrate 402, for example, at 420. The interfaces of each DRAM memory component can be connected in parallel using through-silicon vias or any other connection method. Other stacked configurations are also possible.
[0048] Continuing to refer to Figure 4 , for one embodiment, the four stacks 408 - 414 of DRAM devices can be interconnected in a ring configuration such that the first DRAM stack 408 is directly connected to the DQv nibble. The second stack 410 is coupled to the first stack 408 via a nibble of the path at 411. The second stack 410 is associated with the third stack 412 via a nibble of the path at 413, while the fourth stack 414 is coupled to the fourth stack 414 via a nibble of the path at 415. It is also directly connected to the DQt nibble.
[0049] Continuing to refer to Figure 4 , the DRAM memory module 400 includes a control / address (CA) buffer component RCD that drives the CAI connections CAya and CAyb connected to a pair of DRAM memory stacks. For this configuration, a given pair of DRAM stacks, such as at 408 and 410, can be accessed independently of the pair of stacks at 412 and 414.
[0050] The operation of the various system components described above will begin with a discussion of the interaction between the memory control circuitry 110 and the non-volatile memory module 200. Then the complete system will be used to illustrate the various configurable operating environments, including the memory control circuitry, the non-volatile memory module, and the DRAM module.
[0051] As described above, aspects of the circuits described herein enable non-volatile and DRAM memory modules to be used in a point-to-point topology to advantageously expand system storage capacity while maintaining performance. To support including the non-volatile memory module in the system, for a read operation, the memory control circuitry typically operates according to Figure 5A the steps shown. At 502, a read access command is sent along the primary CA bus to the non-volatile memory module. As will be explained below in the context of various system examples, commands to the non-volatile memory module can be interleaved with commands to the DRAM memory module. At 504, after sending the command, the memory control circuitry waits for an indication or signal from the non-volatile memory module that the requested read data is ready to be transferred from the module to the memory control circuitry. As will be more fully explained below, this "wait" is the result of the non-volatile memory module buffer accumulating block read data from the non-volatile device to the SRAM read buffer 330 ( Figure 3 ). Specific embodiments of how the indication or signal is implemented will be described below. Then the block read data is read out as column read data to the memory control circuitry 110. Then at 506, the read data is received by the memory control circuitry 110 from the non-volatile memory module 200 along one of the primary DQ nibbles as column read data.
[0052] Now referring to Figure 5B , from the perspective of the memory control circuitry 110, at 508, a write operation is performed in a manner similar to a read transaction, where the memory control circuitry 110 issues a write access command to the non-volatile memory 200. Then at 510, the column write data is transferred to the non-volatile memory module. As will be more fully explained below, the column write data accumulates in the SRAM write buffer 342 ( Figure 3 ) until it is ready to be transferred to the non-volatile device along the secondary DQ path. When the block data accumulation is complete, the non-volatile memory module buffer sends an indication of write transfer completion to the memory control circuitry at 512.
[0053] Now referring to Figure 6A, from the perspective of the non-volatile memory module 200, at 602, a read transaction begins with the receipt of a read access command from the memory control circuitry 110. As described below, the received command may be interleaved with commands distributed to the DRAM memory modules. Data in the form of a read data block is accessed from the non-volatile memory device 210 at 604 and aggregated as a data block in the SRAM read data buffer 330 at 606. Once the block read is completed at 608, signals such as status bits may be sent along the status line to the memory control circuitry 110 at 610, indicating to the memory control circuitry that the block read is complete. Then at 612, the data is transferred out of the SRAM read buffer 330 and transmitted as column data to the memory control circuitry 110 along a point-to-point nibble.
[0054] Now referring Figure 6B , from the perspective of the non-volatile memory module 200, at 614, a write operation is performed in a manner similar to the read transaction, where a write access command is received by the non-volatile memory. At 616, the SRAM write buffer 342 on the non-volatile memory module 200 then receives column write data from the memory control circuitry 110. At 618, the column data is aggregated in the SRAM write buffer 342. At 620, once the column write data is fully aggregated and organized into write block data for transfer to the non-volatile memory device 210, a status bit is generated by the buffer logic at 622 and sent to the memory control circuitry 110 along the status link.
[0055] Figure 7A A timing diagram is shown, showing various timings for the sequence of the read operation for the non-volatile memory module 200 and the timing for the above-mentioned status bit “S”. The waveform CK represents a timing reference of 3.2 GHz, which corresponds to a primary DQ signaling rate of 6.4 Gb / s for the transfer operation. At 702, the non-volatile memory module receives an activate command along the primary CA path CAx, and at 704, re-transmits the activate command along the secondary CA path CAxa. The non-volatile memory device 210 then transfers the block read data to the SRAM read buffer 330. After an internal transfer to the SRAM is completed after a time interval tR of approximately 25 microseconds, the status bit “S” is sent to the memory control circuitry at 706. In response to receiving the status bit, the memory control circuitry starts distributing grouped read commands “R” starting from 708 and re-transmits at 710. Upon receiving the read command, the data buffer circuitry reads the data read out from the SRAM as column data for transfer to the memory control circuitry starting at 712.
[0056] Figure 7BShows the timing for a grouped write operation and the corresponding timing for the status bit “S”. At 714, the non-volatile memory module receives an activation command along the primary CA path CAx and, at 716, re-transmits it along the secondary CA path CAxa. Then starting at 718, a string of write commands is transmitted by the memory control circuitry and received by the module. As described above, starting at 720, column write data accumulates in the SRAM write buffer and is transmitted as block write data to the non-volatile memory. Once the block write is complete, at 722, the status bit is generated by the data buffer and sent along the status link to the memory control circuitry. The receipt of the status bit notifies the controller that the write operation is complete.
[0057] At the system level where multiple modules interact with the memory control circuitry 110, various configurable transfer schemes are possible depending on the application. Various examples illustrating the schemes are given below. Generally, the module configurations allow for increasing the capacity of the memory system without sacrificing performance and for using non-volatile memory modules with DRAM modules. These configurations also allow for distributing the total system bandwidth among the modules in a balanced or unbalanced manner.
[0058] Now referring to Figure 8 , a partial system view of a memory system, generally designated 800, is shown as being consistent with the above-described structure. The partial system view includes the memory control circuitry 802 and some portions of the non-volatile memory module 804 and a portion of the DRAM module 806. The corresponding module portions can be considered the corresponding “slices” or replicas of the circuitry corresponding to the non-volatile module nibble pair forming components 206 ( Figure 2 ) and the DRAM module nibble pair forming components 406 ( Figure 4 ). For clarity, the like components of each module are labeled in a manner consistent with the labels of Figure 2 and Figure 4 . For one particular embodiment, the full system employs nine “slices” of the circuitry to perform memory transfers.
[0059] Further referring to Figure 8, the memory control circuit device 802 includes a first data nibble interface circuit DQv, which is connected to the corresponding nibble interface on the DRAM module 806 in a point-to-point relationship along the data path 808. The second data nibble interface DQu is connected to the corresponding nibble interface on the non-volatile memory module 804 in a point-to-point relationship along the data path 810. Although not shown, for some embodiments, a source synchronous timing signal path for clock or strobe signals accompanying data may also be coupled between the memory control circuit device 802 and the modules 804 and 806 in a point-to-point relationship near each data path. The corresponding CA interface circuits CAx and Cay connect the memory control circuit device 802 to each of the module RCD buffers 812 and 814 via point-to-point paths 816 and 818.
[0060] Figure 9 An embodiment of a timing diagram corresponding to the Figure 8 system is shown, showing various interleaved commands and resulting data transfers related to two concurrent read transactions, where half of the system bandwidth is allocated to the DRAM module 806 and half to the DRAM module 804. The waveform CK represents a timing reference of 3.2 GHz, corresponding to a primary DQ signaling rate of 6.4 Gb / s for transmission operations. As the primary DQ rate changes, the relative signal rate of the bus will increase or decrease. Each interleaved read transaction includes an activate command shown at 902 and 904, a read command shown at 906 and 908, and read data shown at 910 and 912.
[0061] Further reference is made to Figure 9, the first read transaction begins with an activation command "A" distributed by the controller on the CAy bus at 904. For one embodiment, the bus has a point-to-point topology and a signaling rate of 1.6 Gb / s, which is one-fourth of the signaling rate of the point-to-point DQ bus. The RCD component on the DRAM module 806 receives the activation command "A" and re-transmits the command information as an "ACT" command on the secondary CA bus CAya at 905 (in this instance, the CA bus CAyb is not used because only the upper DRAM components 408 and 410 are accessed for the read transaction). The secondary CA bus CAya operates at 0.8 Gb / s, which is half the speed of the primary CA bus Cay and one-eighth the speed of the primary DQ bus DQv. One reason for the reduced rate is that the secondary CA bus CAya is a multi-drop bus that connects to approximately one-fourth of the DRAM stacks on the module. After delaying an appropriate amount to compensate for buffer latency, the memory control circuitry 802 then distributes a read command "R" along the CAya bus at 906, which is re-transmitted as an "RD" by the CA buffer component RCD at 914. The read data is then accessed from the DRAM components 408 and 410 and transmitted to the memory control circuitry 802 along the primary DQ path DQv at 912.
[0062] Concurrent with the first read transaction described above and continuing with reference to Figure 8 and Figure 9 , the second read transaction begins at 902 with an activation command "A" distributed by the memory control circuitry along the primary CA path CAx. The RCD component on the non-volatile module 804 receives the activation command "A" at 903 and re-transmits the information "ACT" on the secondary CA bus CAxa. The memory control circuitry 802 then distributes a read command "R" at 906, which is re-transmitted as an "RD" by the RCD component along the CAxa bus at 907. The read data is then accessed from the non-volatile component 210, accumulated as block data by the buffer, and transmitted to the memory control circuitry 802 as column data along the primary DQ path DQu at 910.
[0063] For the read transaction examples referred to above with reference to Figure 8 - Figure 17 shown, the commands and data for each transaction are typically pipelined. This means that they occupy fixed timing positions with respect to the transaction and also means that the transactions may overlap with other transactions. It should be noted that the write transactions corresponding to each configurable instance discussed in Figure 8 - Figure 17 are executed in a manner similar to the read operation, but the fixed timing positions of the commands and data are different.
[0064] The above read transaction also shows a timing interval that may be shorter compared to the time intervals associated with typical systems. For example, the interval tRCD from activation ACT to read command RD is shown as 6.25 ns, but is approximately 12.5 ns for typical DRAM components. This compression of the time scale is for clarity and does not affect the technical accuracy of the embodiments presented herein. For a tRCD delay of 12.5 ns, the pipeline timing also applies.
[0065] Regarding transaction granularity, the above example illustrates a granularity of 64 bytes. Thus, there are sufficient command slots to allow each primary DQu and DQv slot to be filled with data. Each transaction performs a random row activation and column access on each group of 64 bytes (“36x16b”). Additionally, the size of each byte is assumed to be 9 bits. This additional size is used for the syndrome of error detection and check code (EDC). If there is a bank conflict in the transaction stream and if the transaction stream switches between read and write operations, data slots may be skipped. This form of bandwidth inefficiency exists in all memory systems. The embodiments described herein do not introduce additional resource conflicts.
[0066] Figure 10 and Figure 11 shows an example of system operation similar to Figure 8 and Figure 9 shown, where 50% of the system bandwidth is allocated to the non-volatile memory module 1004 and 50% is allocated to the DRAM module 1006. However, in this example, instead of the upper DRAM devices 408 and 410 being accessed, the lower devices 414 and 412 are accessed. This is done by distributing commands on the secondary CA bus CAyb and taking advantage of a bypass formed by the “loop” connection configuration between various DRAM component stacks.
[0067] Therefore, now referring to Figure 11 , the read access to the non-volatile memory module 1004 includes receiving an activation command “A” on the CAx bus at 1102 and retransmitting this command as “ACT” on the secondary CA bus CAxa at 1104. Then, the corresponding read command is received at 1106 after the tR + tSCD interval and the read command is retransmitted at 1108 after the buffer delay tBUF. Then, the resulting read data is sent on the secondary DQ path DQyab at 1110 and then the read data is transferred to the memory control circuitry via the DQ primary nibble path DQu at 1112.
[0068] In parallel with the read transaction to the non-volatile memory module 1004, the DRAM memory module 1006 receives an activate command "A" at 1114 along the primary CA bus CAy, and re-transmits the command as "ACT" at 1116 along the secondary CAyb bus. Just before receiving the read command, a bypass control signal "B" is received at 1118 and re-transmitted at 1120 along the secondary CAya bus, thereby enabling the bypass data path between the upper and lower DRAM component stacks at 1010. Then a read command "R" is received at 1122 and re-transmitted at 1124 along the secondary CA bus CAyb. The resulting read data is driven along the bypass path 1010 and then the read data is transferred at 1126 along the primary DQ nibble path DQV to the memory control circuitry.
[0069] Now referring to Figure 12 and Figure 13 , another embodiment employs a non-volatile memory module 1204 and a DRAM memory module 1206, where the entire system bandwidth can be allocated to the DRAM module 1206. As Figure 12 shown, read data from the upper DRAM module stacks 408 and 410 is accessed and the read data is directly driven back to the memory control circuitry 1202 via the primary DQv nibble path. Data from the lower DRAM stacks 412 and 414 is driven onto the primary shared DQ path DQt and driven to the non-volatile memory module buffer 208, and re-transmitted by the buffer along the primary DQ path DQu to the memory control circuitry 1202.
[0070] Figure 13 Illustrated is for Figure 12Timing of various commands and data for the above example. As the system bandwidth is fully allocated to the DRAM module 1206, the corresponding activation commands "A" are sent by the memory control circuitry 1202 along the primary CA path CAy at 1302 and 1304 and received by the DRAM module 1206. Since the upper and lower DRAM stacks respond to commands sent on the independent secondary CA paths CAya and CAyb, the two activation commands "A" are re-sent along the two secondary CA paths at 1306 and 1308. Then the corresponding read commands "R" are received along the primary CA link CAy at 1310 and 1312, and the read commands are re-sent along the secondary paths CAya and CAyb at 1314 and 1316. Then, in response to the read command from the secondary link CAya, the read data from the upper DRAM component is directly transmitted to the memory control circuitry 1202 along the primary DQ path DQv at 1318. At 1320, the bypass command "B" activates the steering logic in the non-volatile data buffer 208 such that the buffer secondary interface (including the buffer memory SRAM) is bypassed. Then at 1322, in response to the read command "R" from the secondary link CAyb, the read data is transmitted to the non-volatile memory module 1204 along the primary DQ path DQt (thus causing a buffer delay due to the re-transmission of the read data from the buffer 208), and is transmitted to the memory control circuitry 1202 via the primary DQ path DQu at 1324.
[0071] In another system embodiment that employs both the non-volatile memory module 1404 and the DRAM memory module 1406 in a point-to-point topology, the two modules can directly transfer data between each other. This example is shown in Figure 14 and Figure 15 . Generally, as shown in the partial system diagram of Figure 14 , data from, for example, the upper DRAM stacks 408 and 410 can be read from the DRAM module 1406, while data can be read from the non-volatile memory module 1404 in parallel and transmitted as write data to the lower DRAM stacks 412 and 414 modules. In fact, three transactions occur concurrently.
[0072] Figure 15 Shows for Figure 14Timing of various commands and data for the above example. Each of the two read transactions includes an activation command "A" transmitted along primary CA links CAx and CAy at 1502 and 1504. These commands are then retransmitted along secondary CA path CAxa at 1506 and along path CAya at 1508. Then the corresponding read commands are received at 1510 and 1512, and the read commands are retransmitted along the secondary CA path at 1514 and 1516, respectively.
[0073] A single write transaction includes an activation command "A" at 1518, which is retransmitted at 1520. Then the write command is received at 1522 and retransmitted at 1524. For this example, the write data used is generated by the read transaction for the non-volatile memory. The timing of the write transaction is configured to match the read transaction with the interval from column command to column data. At 1526, data is transferred between the two modules on the shared DQ bus DQt. At 1528, additional read data is directly transferred to the memory control circuitry via the DQ path DQv. When the command-data interval for the write operation matches the read operation, the memory control circuitry 1402 is responsible for bank usage when a read transaction is performed on the same stack after a transfer transaction or a write transaction to the DRAM stack.
[0074] Figure 14 and Figure 15 The transfer examples of and may have different variations according to the application. Some of these variations include: [1] The transfer may include a column read operation from the NVM stack coupled (via the DQt bus) to a write operation on the DRAM stack (and an independent column read operation from another DRAM stack to the controller via the DQv bus) - this is the above in Figure 14 and Figure 15The examples described in the context of [2]. The transfer may include a column read operation from a DRAM stack (coupled via the DQt bus to a column read operation of a non-volatile memory stack) (and an independent column read operation from another DRAM stack to the controller via the DQv bus). [3] The above transfer [1] or [2], where the independent operation is a column write operation to another DRAM stack. [4] The above transfer [1], where the column read operation from the non-volatile memory stack is also driven on the DQu bus to the memory control circuitry (and a write operation on the DQt bus to the DRAM stack). [5] The above transfer [1], [2] or [3], where a second independent column read operation is performed from the non-volatile memory stack to the controller via the DQu bus. [6] The above transfer [1], [2] or [3], where a second independent column write operation is performed from the memory control circuitry to the non-volatile memory stack via the DQu bus. It should be noted that the transfer variations of [5] and [6] above involve the non-volatile memory module being able to perform two simultaneous column operations (like a DRAM module).
[0075] The direct transfer operation between the above non-volatile module and the DRAM module can also be used for alternative purposes. Dedicated physical space can be allocated in the DRAM device to be used as a temporary buffer for non-volatile memory read operations and non-volatile memory write operations. This will allow the size of the SRAM buffer space in the non-volatile memory module to be reduced. This alternative scheme will cause all non-volatile memory read and write operations to be carried out in two steps. For reading, the non-volatile memory reads data which is transferred on the DQt primary link to write to the temporary DRAM buffer. When the non-volatile memory read is complete, as described above, the data block can be accessed in the DRAM buffer via the DQt / DQu link. For writing, the write data is transferred on the DQt / DQu primary link to write to the temporary DRAM buffer. When the DRAM buffer has a complete block, as described above, it will be written to the non-volatile memory module via the DQt link.
[0076] For an alternative embodiment, the DRAM module employed in any of the above system diagrams can be of the low-load DIMM type, which is similar to the Figure 4 RDIMM DRAM memory module described in the reference, but also includes a data buffer circuit between the DRAM components and the module pin interface. Each DRAM stack can be connected to the buffer in a point-to-point configuration (instead of the ring configuration described in the previous reference). Figure 4 described ring configuration) connected to the buffer.
[0077] Now refer to Figure 16 and Figure 17, another system embodiment employs a non-volatile memory module 1604 and a DRAM memory module 1606 in a point-to-point topology, where in some cases, the full system bandwidth can be allocated to the non-volatile memory module 1604. Generally speaking and now referring to Figure 16 , when paired with a buffered DRAM module such as an LRDIMM, a nibble of data can be directly accessed from the non-volatile memory module 1604 and transmitted directly to the memory control circuitry 1602 via the primary DQ nibble path DQu, and a second nibble of data can be accessed from the non-volatile memory module 1604 (concurrently with the first access), transmitted to the buffered DRAM module 1606 via the shared DQ path DQt, then retransmitted by the DRAM buffer circuitry at 1608 and transmitted directly to the memory control circuitry 1602 via the DQ primary nibble path DQv.
[0078] Figure 17 Shows the timing of various commands and data for the Figure 16 above example. A plurality of activate commands "A" for reading a nibble are received along the primary CA bus CAx at 1702 and 1704, and retransmitted as "ACT" commands along the respective secondary CA buses CAxa and CAxb at 1706 and 1708. Then the corresponding read commands "R" are received and retransmitted at 1710 and 1712. Concurrent with the receipt of the read command, the DRAM module 1606 receives a bypass command "B" along the primary CA path CAy at 1714, indicating to the DRAM buffer to enable bypass of the read data sent from the non-volatile memory module 1604 along the shared primary data path DQt. This data is shown at 1716, and the generated data is transmitted along the data paths DQu and DQv shown at 1718 and 1720 respectively. It should be noted that the bypass control signal "B" can be distributed by the memory control circuitry 1602 or the non-volatile memory module 1604.
[0079] The above system example is shown and described in a dual-module context only for clarity, and is intended to convey a general point-to-point architecture for multiple memory modules that can be of the same type or mixed. Figure 18A particular embodiment for a three-module configuration is shown. The system, generally designated 1800, includes memory control circuitry 1802 coupled to paired DRAM modules 1804 and 1806 and a single non-volatile memory module 1808. For the particular embodiment shown, each of the DRAM modules 1804 and 1806 is an LRDIMM and employs a topology in which, for each nibble, one DQ nibble (e.g., DQu) is connected to one DRAM socket and the other DQ nibble (e.g., DQv) is connected to a second DRAM socket. A third set of motherboard connections (e.g., DQ) connects the other DQ nibbles of the two DRAM module sockets together. The third socket can be used for the non-volatile module 1808 or a third DRAM module. The third socket can be coupled to the memory control circuitry 1802 in a conventional topology in which, for each pair of nibbles, both DQ nibbles DQu and DQv connect the controller interface to the socket interface.
[0080] When received within a computer system via one or more computer-readable media, such data- and / or instruction-based representations of the above-described circuits can be processed by a processing entity (e.g., one or more processors) within the computer system in conjunction with the execution of one or more other computer programs, including but not limited to a netlist generation program, a placement and routing program, etc., to produce a representation or image of the physical manifestation of such circuits. Such a representation or image can thereafter be used in device fabrication, e.g., by implementing the generation of one or more masks for forming the various components of the circuit in a device fabrication process.
[0081] In the foregoing description and drawings, specific terms and reference symbols have been set forth to provide a thorough understanding of the present invention. In some instances, the terms and symbols may imply specific details that are not necessary to practice the present invention. For example, any one of a specific number of bits, signal path widths, signaling or operating frequencies, component circuits or devices, etc. may be different from those described in the alternative embodiments above. Additionally, the interconnection between circuit elements or circuit blocks shown or described as a multi-conductor signal link may alternatively be a single-conductor signal link, and a single-conductor signal link may alternatively be a multi-conductor signal link. Signal and signaling paths shown or described as single-ended may also be differential, and vice versa. Similarly, in alternative embodiments, signals described or depicted as having a high-level active or low-level active logic level may have the opposite logic level. Component circuits within an integrated circuit device may be implemented using metal-oxide semiconductor (MOS) technology, bipolar technology, or any other technology in which logic and analog circuits may be implemented. Regarding terminology, a signal is said to be "asserted" to indicate a particular condition when the signal is driven to a low or high logic state (or charged to a high logic state or discharged to a low logic state). Conversely, a signal is said to be "de-asserted" to indicate that the signal is driven (or charged or discharged) to a state other than the asserted state (including a high or low logic state, or a floating state that may occur when the signal driving circuit transitions to a high-impedance state such as an open-drain or open-collector state). A signal driving circuit is said to "output" a signal to a signal receiving circuit when the signal driving circuit asserts (or de-asserts, where expressly stated or indicated by context) a signal on a signal line coupled between the signal driving circuit and the signal receiving circuit. When a signal is asserted on a signal line, the signal line is said to be "activated", and when the signal is de-asserted, the signal line is "de-activated". Additionally, the prefix symbol " / " appended to a signal name indicates that the signal is an active-low signal (i.e., the asserted state is a logic-low state). A bar over a signal name (e.g., ) is also used to indicate an active-low signal. The term "coupled" is used herein to denote both direct connection and connection through one or more intermediate circuits or structures. "Programming" an integrated circuit device may include, for example but not limited to, loading control values into registers or other storage circuits within the device in response to host instructions and thus controlling aspects of the operation of the device, establishing a device configuration, or controlling aspects of the operation of the device through a single programming operation (such as blowing a fuse in a configuration circuit during device manufacture), and / or connecting one or more selected pins or other contact structures of the device to a reference voltage line (also referred to as tying off) to establish a particular device configuration or aspect of operation of the device. The term "example" is used to denote an example rather than a preference or requirement.
[0082] Although the present invention has been described with reference to specific embodiments thereof, it will be apparent that various modifications and changes can be made thereto without departing from the broader spirit and scope of the invention. For example, features or aspects of any embodiment can be used in combination with, or instead of, the corresponding features or aspects of any other embodiment, at least where practicable. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive.
Claims
1. A buffer circuit, comprising: a primary interface for coupling to a memory controller; a secondary interface for coupling to at least two non-volatile memory devices; and wherein the primary interface receives non-volatile memory access commands from the memory controller, the non-volatile memory access commands being sequentially interleaved with DRAM memory module access commands for execution by a dynamic random access memory (DRAM) memory module, the DRAM memory module being separate from the non-volatile memory devices; wherein the buffer circuit further comprises: an internal storage device; and logic for interpreting commands from the memory controller to perform block accesses to the at least two non-volatile memory devices and transfer data blocks between the internal storage device and the at least two non-volatile memory devices; wherein the logic further comprises bypass circuitry for guiding an input signal through the primary interface without accessing the at least two non-volatile memory devices.
2. The buffer circuit according to claim 1, wherein: the internal storage device is for temporarily storing data blocks transferred from the at least two non-volatile memory devices; and the logic is for interpreting commands from the memory controller to perform column accesses to data blocks stored in the internal storage device.
3. The buffer circuit according to claim 1, wherein: the internal storage device comprises a static random access memory (SRAM).
4. The buffer circuit according to claim 1, wherein, the logic generates status return signals to the memory controller for at least one of the following purposes: (1) enabling parallel access to the at least two non-volatile memory devices, (2) accommodating variable non-volatile memory accesses, and (3) accommodating a larger non-volatile memory access granularity.
5. The buffer circuit according to claim 1, wherein, the logic is configured to handle multiple accesses to the at least two non-volatile memory devices, wherein the accesses overlap in time.
6. The buffer circuit according to claim 1, wherein, the primary interface further comprises at least two data interfaces for coupling to the memory controller.
7. The buffer circuit according to claim 6, wherein, each of the at least two data interfaces comprises: a four-data-link group defining a nibble; and a timing link.
8. The buffer circuit according to claim 6, wherein: a first data interface from the at least two data interfaces provides access to a first group of non-volatile memory devices or a second group of non-volatile memory devices.
9. The buffer circuit according to claim 1, wherein the primary interface receives non-volatile memory access commands from the memory controller for a first data, the non-volatile memory access commands being sequentially interleaved with the DRAM memory module access commands for a second data, the second data being independent of the first data.
10. A method of operation in a buffer circuit according to any one of claims 1-9, the method comprising: receiving, via a primary interface, a first read data transfer command from a memory controller, the first read data transfer command for execution by a non-volatile memory and sequentially interleaved with a second data transfer command for execution by a DRAM memory module; accessing, in response to the first read data transfer command, data blocks from at least two non-volatile memory devices disposed on a non-volatile memory module via a secondary interface; aggregating the accessed data blocks into aggregated read data; and transmitting the aggregated data as column read data to the memory controller via the primary interface.
11. The method according to claim 10, wherein the first read data transfer command is according to a non-volatile memory protocol and the second data transfer command is according to a dynamic random access memory DRAM protocol.
12. The method according to claim 11, further comprising: selectively guiding a second data transfer operation through a buffer circuit between a second memory module and the memory controller in response to the second data transfer command.
13. The method according to claim 10, further comprising: generating a status signal upon completion of a predefined level of aggregation corresponding to parallel block reads from the at least two non-volatile memory devices; and sending the status signal to the memory controller.
14. A buffer integrated circuit IC chip, comprising: a primary interface for coupling to a memory controller; a secondary interface for coupling to at least two non-volatile IC memory chips; and wherein the primary interface receives a non-volatile memory access command from the memory controller, the non-volatile memory access command being sequentially interleaved with a DRAM memory module access command for execution by a dynamic random access memory DRAM memory module, the DRAM memory module being separate from the non-volatile IC memory chips; wherein the buffer integrated circuit IC chip further comprises: an internal storage device; and logic for interpreting commands from the memory controller to perform block accesses to the at least two non-volatile IC memory chips and to transfer data blocks between the internal storage device and the at least two non-volatile IC memory chips; wherein the logic further comprises bypass circuitry for guiding an input signal through the primary interface without accessing the at least two non-volatile IC memory chips.
15. The buffer integrated circuit IC chip according to claim 14, wherein: the internal storage device is for temporarily storing data blocks transferred from the at least two non-volatile IC memory chips; and the logic is for interpreting commands from the memory controller to perform column accesses to data blocks stored in the internal storage device.
16. The buffer integrated circuit IC chip according to claim 14, wherein: the internal storage device comprises a static random access memory SRAM.
Citation Information
Patent Citations
Cyclic buffer mechanism for receiving wireless data under varying data traffic conditions
US20060253642A1
Composite memory having a bridging device for connecting discrete memory devices to a system
US20100091536A1
System Including Hierarchical Memory Modules Having Different Types Of Integrated Circuit Memory Devices
US20100115191A1