Data chunk grouping for memory device with misaligned burst length

By grouping data chunks to align with the burst length of memory devices, the inefficiencies caused by misaligned burst lengths are addressed, resulting in improved efficiency, reduced latency, and enhanced bandwidth utilization.

WO2025137169A1PCT designated stage expired Publication Date: 2025-06-26RAMBUS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/060860
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-12-18
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Memory devices with misaligned burst lengths experience inefficiencies due to unused bandwidth, known as data bubbles, which result in wasted bandwidth and power.

Method used

Implementing data chunk grouping, where multiple data chunks are combined to form a larger chunk that aligns with the burst length of the memory device, thereby eliminating or reducing data bubbles.

Benefits of technology

This approach enhances memory device efficiency by reducing latency and power consumption while maximizing bandwidth utilization, particularly beneficial for applications prioritizing high bandwidth over reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024060860_26062025_PF_FP_ABST
    Figure US2024060860_26062025_PF_FP_ABST
Patent Text Reader

Abstract

A memory device includes an array of memory cells to store data and an interface to transfer at least a portion of the data associated with a plurality of column access operations, wherein a first column access operation of the plurality of column access operations has a shorter duration than a second column access operation the plurality of column access operations, and wherein the plurality of column access operations are performed successively in response to one or more memory access commands to generate a gapless data burst.
Need to check novelty before this filing date? Find Prior Art

Description

DATA CHUNK GROUPING FOR MEMORY DEVICE WITH MISALIGNED BURSTLENGTHBACKGROUND

[0001] Modem computer systems generally include a data storage device, such as a memory component. The memory component may be, for example a random access memory (RAM) or a dynamic random access memory (DRAM). The memory component includes memory banks made up of storage cells which are accessed by a memory controller through a command interface and a data interface within the memory component.

[0002] In double data rate (DDR) memory, a memory controller issues a memory access command, and corresponding data is transferred on both the rising and falling edges of a strobe clock signal (DQS), thereby doubling the data transfer rate relative to the DQS frequency. Such memory devices, which use a system clock frequency that is the same frequency as the strobe signal, can utilize a burst length that is a multiple of two so that the burst length is aligned with the system clock signal (i.e., the end of the data burst is aligned with a rising edge of the system clock signal). Certain device, however, such as low power memory devices or other devices, may use a lower clock frequency (e.g., one half of the strobe clock frequency) to reduce command and clock power consumption, and thus utilize a burst length that is a multiple of four in order to align the end of the data burst with a rising edge of the system clock signal (i.e., 2x2 to account for the half clock frequency). As newer memory technologies are implemented, however, there are challenges associated with increasing the data bus width by a power of two. For example, while many memory systems have a bus width of 4 or 8 data lines, other memory technologies, may utilize a data bus width of 12 data lines (i.e., “by 12” or “xl2”) to increase the bandwidth. A typical application may be configured to transfer (e.g., read or write) a data chunk of 32 bytes in a memory access operation. Thus, to transfer 32 bytes (i.e., 256 bits) of data, a xl2 memory device can utilize a burst length of approximately 21.3 (i.e., 256 / 12) to transfer the data in a single burst.

[0003] If the memory device uses a lower clock frequency, however, the burst length may be a multiple of four, and the next largest multiple of four equates to a burst length of 24. Accordingly, the burst length can be said to be “misaligned,” as approximately 11.1% (i.e., 2.7 / 24) of the data bus bandwidth would go unused during the data transfer. This unused bandwidth can be referred to as a data “bubble” and is preferably avoided as bandwidth and power utilization are wasted. Some memory systems fdl the bubbles withmetadata (e.g., cyclic redundancy check (CRC) data or error correction code (ECC) data) so that the full burst length of 24 is utilized in the data transfer. In certain implementations, however, the applications utilizing the memory prioritize high bandwidth over reliability. For example, certain machine learning applications or image processing applications may exhibit these characteristics. Thus, it may be preferrable for such application to maximize the data payload that is transferred by eliminating the bubbles, rather than fdling them with metadata.BRIEF DESCRIPTION OF THE DRAWINGS

[0004] The present disclosure is illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings.

[0005] Figure 1 is a block diagram illustrating a computing environment with a memory module configured for data chunk grouping, according to an embodiment.

[0006] Figure 2 is a block diagram illustrating a computing environment with a memory module configured for data chunk grouping, according to an embodiment.

[0007] Figure 3 is a block diagram a memory device configured for data chunk grouping, according to an embodiment.

[0008] Figure 4 is a flow diagram illustrating a method of data chunk grouping for a memory device with a misaligned burst length, according to an embodiment.

[0009] Figure 5 is a read command sequence diagram illustrating data chunk grouping for a memory device with a misaligned burst length, according to an embodiment.

[0010] Figure 6 is a timing diagram illustrating data chunk grouping for a memory device with a misaligned burst length, according to an embodiment.DETAILED DESCRIPTION

[0011] The following description sets forth numerous specific details such as examples of specific systems, components, methods, and so forth, in order to provide a good understanding of several embodiments of the present disclosure. It will be apparent to one skilled in the art, however, that at least some embodiments of the present disclosure may be practiced without these specific details. In other instances, well-known components or methods are not described in detail or are presented in simple block diagram format in order to avoid unnecessarily obscuring the present disclosure. Thus, the specific details set forth are merely exemplary. Particular implementations may vary from these exemplary details and still be contemplated to be within the scope of the present disclosure.

[0012] Aspects of the present disclosure include data chunk grouping for a memory device with misaligned burst length. In one embodiment, a memory module, such as a dual in-line memory module (DIMM), includes a number of memory devices, such as dynamic random access memory (DRAM) devices. The memory module may further include a memory interface chip, such as a memory buffer, to control certain communications between the memory module and an external device, such as a memory controller or host system. Depending on the implementation, the memory interface chip can be connected internally within the memory module to the individual memory devices via associated links, such as command / address (CA) lines and / or data (DQ) lines. The burst length of the memory device represents how many data words can be transferred either to or from the memory device in a single burst or sequence during a memory access operation (e.g., a read operation or a write operation). The number of data lines in the memory module (i.e., the width of a data bus including the data lines) determines the number of bits that can be used for each data word transfer. Accordingly, the total amount of data that can be transferred in a burst is calculated by multiplying the burst length of the memory device by the width of the data bus.

[0013] In certain embodiments, the memory system can group multiple data chunks together for transfer to or from a memory device with a misaligned burst length. For example, if three 32 byte chunks can be combined into one 96 byte chunk (i.e., 768 bits), a burst length of 64, which is a multiple of four, can be used for a xl2 memory device to transfer the data in a single burst. This results in 0% overhead as the full data payload fits evenly with the given burst length. Even combining two 32 byte chunks into one 64 byte chunk (i.e., 512 bits), allows for using a burst length of 44, which is a multiple of four, with the xl2 memory device. In this case, there is approximately 3% overhead (i.e., 1.3 / 44), but this still represents an improvement over the 11% overhead present for a single 32 byte chunk.

[0014] Benefits that can be realized with certain embodiments of the approach described herein include, but are not limited to, improved efficiency in memory devices with misaligned burst lengths. As described herein, different memory access commands can be utilized depending on a mode of operation, to either proceed with a conventional operation where data chunks are not grouped so that metadata can be included with the payload data, or to group data chunks to reduce or eliminate bubbles in the data transfer. This approach enables improved performance for certain applications that prioritize bandwidth over reliability. The reduction or elimination of bubbles can reduce latency and power utilization associated with the data transfers and increase the bandwidth utilization in the memorysystem. Additional details with respect to the data chunk grouping for a memory device with a misaligned burst length are provided below with respect to Figures 1-6.

[0015] Figure 1 depicts an environment 100 showing a memory module 120. As an option, one or more instances of environment 100 or any aspect thereof may be implemented in the context of the architecture and functionality of the embodiments described herein.

[0016] As shown in Figure 1, environment 100 comprises a memory controller 102 coupled to a memory module 120. In one embodiment, memory module 120 is a dual in-line memory module (DIMM). Such memory modules can be referred to as DRAM DIMMs or load reduced DIMMs (LRDIMMs), or as a Compression Attach Memory Module (CAMM), and can share a memory channel with other DIMMs.

[0017] In one embodiment, the memory controller 102 comprises a clock signal generator 104 and a memory interface circuit 105. Depending on the embodiment, memory controller 102 can comprise multiple instances each of clock signal generator 104 and memory interface circuit 105. The memory controller 102 can further include a cache memory (not shown), which can be dedicated to a single processing core or shared with other cores as a system on chip (SoC). Clock signal generator 104 may include a phase lock loop (PLL) or other circuit to generate one or more clock signals. In other embodiments, clock signal generator 104 may receive one or more clock signals from a source external to the memory controller 102. In either embodiment, memory interface circuit 105 may include a driver circuits to drive the one or more clock signals and strobe signals from clock signal generator 104 out of memory controller 102 (e.g., to components such as buffer chips on memory module 120).

[0018] In one embodiment, the memory interface circuit 105 of the memory controller 102 communicates with the memory module 120 via a communications bus 115. Specifically, the memory interface circuit 105 can transmit write data to and / or receive read data from multiple sets of DRAM devices 124i - 124& by transferring the data over the communications bus 115. In one embodiment, communication bus 115 can include two or more groups of multiple data signals (e.g., four data signals per group), with each group having a corresponding strobe signal or signals, generated by clock signal generator 104. Memory interface circuit 105 can transmit the data signals together with the strobe signals to memory module 120 over communication bus 115.

[0019] The DRAM devices 124i - 124& in memory module 120 can each comprise an array of memory devices (e.g., SDRAM) arranged in various topologies (e.g., A / B sides, single-rank, dual-rank, quad-rank, etc.). Although a certain number of DRAM devices 124i -124& are illustrated in Figure 1, this is merely an exemplary embodiment, and it should be understood that in other embodiments, memory module 120 can include any other number of DRAM devices. In one embodiment, as shown, the data signals sent to and / or from the DRAM devices 124i - 124& can be buffered by a memory interface chip, such as memory buffer 126. Such a memory buffer 126 can serve to redrive the signals (e.g., data or DQ signals, etc.) on communications bus 115 to help mitigate high electrical loads of large computing and / or memory systems. For example, the memory buffer 126 can include a signal transmitter circuit to transmit the signals and a signal receiver to receive the signals. In one embodiment, at least a portion of the DRAM devices (e.g., DRAM devices 124i - 124s) are connected to memory buffer 126 by command / address lines 127, while another portion of the DRAM devices (e.g., DRAM devices 1244 - 124&) are connected to memory buffer 126 by command / address lines 129. In one embodiment, each of DRAM devices DRAM devices 124i - 124 are connected to memory buffer 126 by respective data busses 1251 - 125&. In another embodiment, each of data busses 1251 - 125& carry data directly between each respective DRAM of DRAM devices 124i - 124& and memory controller 102.

[0020] In addition, command / address signals from the memory interface circuit 105 can be received by memory buffer 126 at the memory module 120 using communication bus 115. A memory buffer, such as command buffer 126, can comprise a logical register and a phase-lock loop (PLL) to receive and re-drive command and address input signals from the memory controller 102 to the DRAM devices on a DIMM (e.g., DRAM devices 1241 , DRAM devices 1242, etc.), reducing clock, control, command, and address signal loading by isolating the DRAM devices from the memory controller 102.

[0021] In one embodiment, each of DRAM devices 124i - 124& is configured with a default burst length representing how many data words can be transferred over the respective data busses 1251 - 125& in a single burst during a memory access operation. In addition, the input / output (I / O) width of the data busses 1251 - 125&, which determines the number of bits that can be used for each data word transfer, may be a non-integer power of two. For example, the I / O width may be 12 in one embodiment, rather than 2, 4, 8, 16, etc. Given the I / O width of the data busses 1251 - 125&, in order to transfer an individual chunk of data (e.g., in response to a request from memory controller 102 or other host system (not shown)), the requisite burst length to transfer the chunk of data may be misaligned with the default burst length of the DRAM devices 124i - 124&. In one embodiment, in order to reduce or eliminate bubbles in the data transferred on data busses 1251 - 125&, multiple data chunks can be grouped together to generate a gapless data burst.

[0022] The grouping of data chunks can be accomplished by identifying memory access requests corresponding to sequential data stored on DRAM devices 124i - 1246. Sequential data, including data stored at memory addresses in the same row of an array of memory cells of one of DRAM devices 124i - 124 , can be accessed using consecutive column access operations. When consecutive column access operations are performed, without other intervening operations, significant time savings can be realized as row activation and precharge steps can be reduced. Thus, when memory controller 102 receives, from a host system, a request to read a number of chunks of sequential data, the memory controller 102 can issue one or more memory access commands (e.g., read commands) to read the data from a given row of one or more of DRAM devices 1241 - 1246. Since the read commands are directed to the same row, only one row activation and one precharge step is needed to read multiple chunks of data, whereas reading chunks from different rows would involve separate row activation and precharge steps for each chunk. Thus, depending on the embodiment, memory controller 102 can issue one or more memory access commands, or cause one or more of DRAM devices 124i - 124& to generate the one or more memory access commands, to cause a number of column access operations to be performed on the DRAM devices 124i - 124& at specific durations.

[0023] For a regular read operation, the memory access command may be issued by memory controller 102 based on the expected duration (i.e., number of clock cycles) of the corresponding column access operations in order to transfer the read data words according to the default burst length. As noted above however, when the burst length is misaligned to the amount of data being transferred, the full number of clock cycles may not be needed. Accordingly, the one or more memory access commands can be issued with different timings in order to group the multiple chunks of data together in a gapless data burst without any, or with fewer, data bubbles. Thus, when multiple column access operations are performed to read and transfer multiple chunks of data together, at least one column access operation has a shorter duration (i.e., lasts fewer clock cycles) than at least one other column access operation. For example, when three column access operations are performed, the initial and middle column access operations may have a shorter duration (e.g., 5 clock cycles) than the last column access operation (e.g., 6 clock cycles). Additional details are provided below.

[0024] The memory module 120 shown in environment 100 presents merely one partitioning. In other embodiments, in addition or in the alternative, memory module 120 may include other volatile memory devices, such as synchronous DRAM (SDRAM), Rambus DRAM (RDRAM), static random access memory (SRAM), etc. The specific example shownwhere the memory buffer 126 and the DRAM devices 124i - 124& are separate components is purely exemplary, and other partitioning is possible. For example, any or all of the components comprising the memory module 120 and / or other components can comprise one device (e.g., system-on-chip or SoC), multiple devices in a single package or printed circuit board, multiple separate devices, and can have other variations, modifications, and alternatives. In addition, memory controller 102 may include additional and / or different components than those illustrated in Figure 1. Furthermore, the illustrated components may be arranged differently depending on the embodiment.

[0025] Figure 2 depicts an environment 200 showing a memory module 220. As an option, one or more instances of environment 200 or any aspect thereof may be implemented in the context of the architecture and functionality of the embodiments described herein.

[0026] As shown in Figure 2, environment 200 comprises a memory controller 102 coupled to a memory module 220 through one or more communications busses. In one embodiment, memory module 220 is a dual in-line memory module (DIMM). Such memory modules can be referred to as load-reduced DIMMs (LRDIMMs), multiplexed-rank DIMMs (MRDIMMs), or a Compression Attach Memory Module (CAMM) and can share a memory channel with other DIMMs.

[0027] In one embodiment, the memory interface circuit 105 of the memory controller 102 communicates with the memory module 220 through a signaling interface formed from one or more communications buses connecting the memory controller 102 with multiple buffer chips on the module. Specifically, the memory interface circuit 105 can write data to and / or read data from multiple sets of DRAM devices 124i - 124& using data busses 2141 - 214&, respectively. For example, the data busses 2141 — 2146 can be used to convey signals transmitted by the memory interface circuit 105, such as a data signal, a chip select signal, and / or a data strobe signal. In one embodiment, data busses 2141 — 214 can each include two or more groups of multiple data signals, with each group having a corresponding strobe signal or signals, generated by clock signal generator 104. Memory interface circuit 105 can transmit the data signals together with the strobe signals to memory module 120 over data busses 2141 — 2146-

[0028] In one embodiment, as shown, the data to and / or from the DRAM devices 124i - 124 can be buffered by a set of data buffers 2221 - 222&, respectively. Such data buffers (DBs) can serve to redrive the signals (e.g., data or DQ signals, etc.) or to combine the signals (e.g., data or DQ signals, etc.) on data busses 214i - 214& to help mitigate highelectrical loads of large computing and / or memory systems. For example, each data buffer can include a signal transmitter circuit to transmit the signals.

[0029] In addition, command / address signals from the memory interface circuit 105 can be received by a command buffer 226, such as a register clock driver (RCD), at the memory module 220 using a command and address (CA) bus 215. For example, the command buffer 226 might be an RCD such as included in registered DIMMs (e.g., RDIMMs, LRDIMMs, MRDIMMs, etc.). Command buffers, such as command buffer 226, can comprise a logical register and a phase-lock loop (PLL) to receive and re-drive command and address input signals from the memory controller 102 to the DRAM devices on a DIMM (e.g., DRAM device 124i, DRAM device 1242, etc.), reducing clock, control, command, and address signal loading by isolating the DRAM devices from the memory controller 102.

[0030] As described above with respect to memory module 120 of Figure 1, each of DRAM devices 124i - 124& is configured with a default burst length representing how many data words can be transferred over the respective data busses 2141 - 214& in a single burst during a memory access operation. In addition, the input / output (I / O) width of the data busses 2141 - 214&, which determines the number of bits that can be used for each data word transfer, may be a non-integer power of two. Given the I / O width of the data busses 2141 - 214 , in order to transfer an individual chunk of data (e.g., in response to a request from memory controller 102 or other host system (not shown)), the requisite burst length to transfer the chunk of data may be misaligned with the default burst length of the DRAM devices 124i - 124&. In one embodiment, in order to reduce or eliminate bubbles in the data transferred on data busses 214i — 214&, multiple data chunks can be grouped together to generate a gapless data burst.

[0031] Figure 3 is a block diagram a memory device configured for data chunk grouping, according to an embodiment. Memory device 300 can represent, for example, any of DRAM devices 1241 - 124&. In one embodiment, memory device 300 includes an array of memory cells 302 arranged in corresponding rows and columns. Each row of memory cells 302 is coupled to one of the wordlines 304 of the array, and each column of memory cells 302 is coupled to one of the bitlines 306. By specifying a wordline 304 and a bitline 306, a particular memory cell 302 can be accessed.

[0032] To enable access to the various memory cells 302, memory device 300 includes a row decoder 312 and a column decoder 314. The row decoder 312 receives a row address on a set of address lines 316, and a row address strobe (RAS) signal on a control line 318, and in response to these signals, the row decoder 312 decodes the row address to selectone of the wordlines 304 of the memory device 300. The selection of one of the wordlines 304 causes the data stored in all of the memory cells 302 coupled to that wordline 304 to be loaded into the sense amplifiers (sense amps) 308. That data, or a portion thereof, may thereafter be placed onto the data bus 310 (for a read operation). Data bus 310 can represent any of data busses 1251 - 1256 or data busses 214i - 214&, or data bus 310 can be further serialized to be connected to any of data busses 1251 - 1256 or data busses 2141 - 214&.

[0033] What portion of the data in the sense amps 308 is actually placed onto the data bus 310 in a read operation is determined by the column decoder 314. More specifically, the column decoder 314 receives a column address on the address lines 316, and a column address strobe (CAS) signal on a control line 320, and in response to these signals, the column decoder 314 decodes the column address to select one or more of the bitlines 306 of the memory device 300. The number of bitlines 306 selected in response to a single column address may differ from implementation to implementation, and is referred to as the base granularity of the memory device 300. For example, if each column address causes sixty- four bitlines 306 to be selected, then the base granularity of the memory device 300 is eight bytes. Defined in this manner, the base granularity of the memory device 300 refers to the amount of data that is read out of or written into the memory device 300 in response to each column address / CAS signal combination (i.e. each CAS or column command).

[0034] Once the appropriate bitlines 306 are selected, the data in the sense amps 308 associated with the selected bitlines 306 are loaded onto the data bus 310. Data is thus read out of the memory device 300. In one embodiment, this process can be referred to as a column access operation. Data may be written into the memory device 300 in a similar fashion. A point to note here is that in a typical memory device, the address lines 316 are multiplexed. Thus, the same lines 316 are used to carry both the row and column addresses to the row and column decoders 312, 314, respectively. That being the case, a typical memory access requires at least two steps: (1) sending a row address on the address lines 316, and a RAS on the control line 318; and (2) sending a column address on the address lines 316, and a CAS on the control line 320. The row and column address information can be part of memory access commands received from memory controller 102 or generated by control logic of the memory device 300.

[0035] Figure 4 is a flow diagram illustrating a method of data chunk grouping for a memory device with a misaligned burst length, according to an embodiment. The method 400 may be performed by processing logic that may comprise hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions run on aprocessing device to perform hardware simulation), or a combination thereof. In one embodiment, the method 400 is performed by control logic of memory controller 102. In another embodiment, the method 400 is performed by a control logic within memory module 120, such as by memory buffer 126, command buffer 226 or one of DRAM devices 124i - 1246.

[0036] At block 410, the processing logic receives a request to access data stored at a memory module, such as memory module 120 or memory module 220. In one embodiment, the request is sent by a host system and receive by memory controller 102. Depending on the embodiment, the request can include a read request to read data stored on the memory module or a write request to program data to the memory module. In one embodiment, the request identifies particular data stored on one or more of DRAM devices 124i - 1246, which may include one or more chunks of data of a particular size (e.g., 32 bytes).

[0037] At block 420, the processing logic determines whether to apply data chunk grouping to the received request. In one embodiment, the memory controller 102 analyzes the request to determine whether to apply data chunk grouping. In one embodiment, the request includes an explicit indication of whether to apply data chunk grouping. In another embodiment, the memory controller analyzes one or more characteristics of the request to determine whether to apply data chunk grouping. The one or more characteristics can include, for example, the source of the request. For example, if a particular requestor (e.g., host application) is known to prioritize high bandwidth over reliability, the memory controller 102 may determine to apply data chunk grouping. Conversely, if the requestor is known to prioritize reliability over bandwidth, the memory controller 102 may determine to not apply data chunk grouping.

[0038] If the processing logic determines not to apply data chunk grouping to the request, at operation 430, the processing logic generates one or more memory access commands without grouping. The memory controller 102 issues these memory access commands to the memory module 120 or memory module 220 and the requested data is returned from one or more of DRAM devices 124i - 1246, for example, via one or more of data busses 1251 - 125& or data busses 214i - 2146. As noted above, given the I / O width of the data busses 1251 - 1256, the requisite burst length to transfer the individual chunks of the requested data may be misaligned with the default burst length of the DRAM devices 124i - 124 . Accordingly, one or more data bubbles may be present in the data transfer.

[0039] If, however, the processing logic determines to apply data chunk grouping to the request, at operation 440, the processing logic generates one or more memory accesscommands using data chunk grouping. For example, the memory access commands may be modified to include an opcode to indicate that data chunk grouping should be applied. In one embodiment, the command syntax for a read command includes “RDXY”, where the variable X indicates how many data chunks are to be grouped together, and the variable Y indicates how many consecutive column access operations are to be performed to read the data from a given memory page.

[0040] Figure 5 is a read command sequence diagram 500 illustrating data chunk grouping for a memory device with a misaligned burst length, according to an embodiment. The command sequence for a write command can be the same as the read command sequence. The command sequence diagram 500 compares memory access commands that use data chunk grouping and those that do not. Command sequence 502 illustrates memory access commands “RD” that do not use data chunk grouping, such as those used at operation 430. As illustrated in command sequence 502, there is a separate read command issued to read each individual data chunk. The read commands are separated in time to permit respective column access operations to be performed on the memory array storing the requested data. Thus, even if the commands are directed to sequential memory addresses, the data will be returned using the default burst length of DRAM devices 124i - 124& and may include data bubbles, which are either left empty or can be filled with metadata, for example.

[0041] Command sequence 504, however, illustrates a memory access command “RD33”, a command encoding issued by the memory controller 102 indicating that 3 data chunks are to be grouped together using 3 consecutive column access operations. Similarly, command sequence 506 illustrates a memory access command “RD32” indicating that 3 data chunks are to be grouped together using 2 consecutive column access operations. This command is followed by a memory access command “RD31” which indicates the third data chunk in the group, but which is read from a non-sequential memory address using a separate column access operation. Finally, command sequence 508 illustrates a series of memory access commands “RD31” indicating that 3 data chunks are to be grouped together, where each chunk is read using a separate column access operation.

[0042] Referring again to Figure 4, in an alternative embodiment, the memory controller 102 sends the received request to memory module 120 or memory module 220 and either memory buffer 126, command buffer 226, or one or more of DRAM devices 1241 - 124& generates the memory access commands, such as those shown in timing diagram 500.

[0043] In an embodiment where memory controller 102 generates the memory access commands, at operation 450, the processing logic sends the one or more memory accesscommands to memory module 120 or memory module 220 to initiate a sequence of column access operations to read the requested data. Using memory device 300, as an example, when a wordline 304 and one or more bitlines 306 are selected (e.g., based on the received address information), the data in the corresponding memory cells 302 can be loaded into the sense amps 308. Depending on the addresses of the requested data, one or more sequential or consecutive column access operations can be performed. For example, if the requested data is stored in the memory cells associated with one wordline 304 and one or more adjacent bitlines 306, the consecutive column access operations can be performed. As noted above, the activation and precharge operations associated with reading data from the memory cells associated with one wordline 304 need only be performed once for all of the consecutive column access operations rather than once for each individual column access operation.

[0044] At operation 460, the processing logic transfers the data read during the column access operations as a gapless data burst. In one embodiment, the data is transferred from DRAM devices 1241 - 124& to memory buffer 126 or to memory controller 102 on data busses 125i - 1256 or data busses 2141 — 2146- As described below, with reference to Figure 6, the individual data chunks can be combined into a combined data chunk and transferred as a gapless data burst without data bubbles in a shorter amount of time, which reduces latency and improves system performance.

[0045] Figure 6 is a timing diagram 600 illustrating data chunk grouping for a memory device with a misaligned burst length, according to an embodiment. The timing diagram 500 compares memory access commands that use data chunk grouping and those that do not. Line 602 illustrates timing of the command sequence 502 using memory access commands “RD” that do not use data chunk grouping, such as those used at operation 430. As illustrated in line 602, there is a separate read command issued to read each individual data chunk and the read commands are separated in time to permit respective column access operations to be performed on the memory array storing the requested data. In the illustrated example, each column access operation has a duration of 6 clock cycles (CLKs), which equates to a burst length of 24. As shown, however, the amount of data transferred in each burst is less than 24 (e.g., 21.3) which leaves a data bubble in each burst.

[0046] Line 608, however, illustrates timing of the command sequence 508 using a series of memory access commands “RD31”. When three data chunks are grouped together, even when coming from non-sequential memory addresses, a time savings can be realized compared to conventional commands. For example, because the memory controller 102 or the memory device issuing the memory access commands is aware that a number of datachunks will be grouped together, the timing with which the commands are sent can be modified. For example, as shown in line 608, the processing logic can modify the burst length used so that the data read during a first column access operation associated with the first “RD31” command can be transferred on data busses 1251 - 1256 or data busses 214i — 214 in a duration of 5 clock cycles. Accordingly, the second “RD31” command can be sent after the 5 clock cycles have elapsed. Similarly, the data read during a second column access operation associated with the second “RD31” command can be transferred on data busses 125i - 1256 or data busses 2141 - 214& in a duration of 5 clock cycles. Accordingly, the third “RD31” command can be sent after the 5 clock cycles have elapsed. The remaining data is read during a third column access operation associated with the third “RD31” command can be transferred on data busses 1251 - 125& or data busses 2141 - 214& in a duration of 6 clock cycles. Thus, when arranged in a gapless data burst, the same amount of data that was transferred in 18 clock cycles using read commands without data grouping in line 602 can be transferred in 16 clock cycles instead.

[0047] Although the operations of the methods herein are shown and described in a particular order, the order of the operations of each method may be altered so that certain operations may be performed in an inverse order or so that certain operation may be performed, at least in part, concurrently with other operations. In certain implementations, instructions or sub-operations of distinct operations may be in an intermittent and / or alternating manner.

[0048] It is to be understood that the above description is intended to be illustrative, and not restrictive. Many other implementations will be apparent to those of skill in the art upon reading and understanding the above description. The scope of the disclosure should, therefore, be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled.

[0049] In the above description, numerous details are set forth. It will be apparent, however, to one skilled in the art, that the aspects of the present disclosure may be practiced without these specific details. In some instances, well-known structures and devices are shown in block diagram form, rather than in detail, in order to avoid obscuring the present disclosure.

[0050] Some portions of the detailed descriptions above are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work toothers skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of steps leading to a desired result. The steps are those requiring physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like.

[0051] It should be bome in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless specifically stated otherwise, as apparent from the following discussion, it is appreciated that throughout the description, discussions utilizing terms such as “receiving,” “determining,” “selecting,” “storing,” “setting,” or the like, refer to the action and processes of a computer system, or similar electronic computing device, that manipulates and transforms data represented as physical (electronic) quantities within the computer system's registers and memories into other data similarly represented as physical quantities within the computer system memories or registers or other such information storage, transmission or display devices.

[0052] The present disclosure also relates to an apparatus for performing the operations herein. This apparatus may be specially constructed for the required purposes, or it may comprise a general purpose computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a computer readable storage medium, such as, but not limited to, any type of disk including floppy disks, optical disks, CD-ROMs, and magnetic-optical disks, read-only memories (ROMs), random access memories (RAMs), EPROMs, EEPROMs, magnetic or optical cards, or any type of media suitable for storing electronic instructions, each coupled to a computer system bus.

[0053] The algorithms and displays presented herein are not inherently related to any particular computer or other apparatus. Various general purpose systems may be used with programs in accordance with the teachings herein, or it may prove convenient to construct more specialized apparatus to perform the required method steps. The required structure for a variety of these systems will appear as set forth in the description. In addition, aspects of the present disclosure are not described with reference to any particular programming language. It will be appreciated that a variety of programming languages may be used to implement the teachings of the present disclosure as described herein.

[0054] Aspects of the present disclosure may be provided as a computer program product, or software, that may include a machine-readable medium having stored thereon instructions, which may be used to program a computer system (or other electronic devices) to perform a process according to the present disclosure. A machine -readable medium includes any procedure for storing or transmitting information in a form readable by a machine (e.g., a computer). For example, a machine-readable (e.g., computer-readable) medium includes a machine (e.g., a computer) readable storage medium (e.g., read only memory (“ROM”), random access memory (“RAM”), magnetic disk storage media, optical storage media, flash memory devices, etc.).

Claims

CLAIMSWhat is claimed is:

1. A memory device comprising: an array of memory cells to store data; and an interface to transfer at least a portion of the data associated with a plurality of column access operations, wherein a first column access operation of the plurality of column access operations has a shorter duration than a second column access operation the plurality of column access operations, and wherein the plurality of column access operations are performed successively in response to one or more memory access commands to generate a gapless data burst.

2. The memory device of claim 1, wherein the plurality of column access operations comprises three column access operations, and wherein a third column access operation of the plurality of column access operations has the shorter duration than the second column access operation.

3. The memory device of claim 1, further comprising: a command interface to receive memory access commands, wherein the commands include respective indications of a number of consecutive columns of the array of memory cells to be read to generate the gapless data burst.

4. The memory device of claim 3, wherein the number of consecutive columns of the array of memory cells to be read represents a number of individual data chunks requested by a host system, and wherein the memory device is to combine the number of individual data chunks to form a combined data chunk to be transferred by the gapless data burst.

5. The memory device of claim 1, wherein the one or more memory access commands are generated by the memory device in response to a request from a host system for the portion of the data.

6. The memory device of claim 1, wherein at least the portion of the data is transferred over a data bus having an input / output (I / O) width that is not an integer power of two.

7. The memory device of claim 1, wherein a default burst length representing how many data words can be transferred over the data bus in a single burst is misaligned with a total size of the portion of the data to be transferred.

8. A memory controller comprising: a memory interface circuit to send one or more memory access commands to a memory device, wherein the one or more memory access commands are to initiate a plurality of column access operations associated with at least a portion of data stored in an array of memory cells of the memory device, wherein a first column access operation of the plurality of column access operations has a shorter duration than a second column access operation the plurality of column access operations, and wherein the plurality of column access operations are performed successively to generate a gapless data burst.

9. The memory controller of claim 8, further comprising: a host interface circuit to receive a request from a host system for the portion of the data, wherein the memory controller is to generate the one or more memory access commands in response to receiving the request.

10. The memory controller of claim 9, wherein the one or more memory access commands comprise respective indications of a number of consecutive columns of the array of memory cells to be read to generate the gapless data burst.

11. The memory controller of claim 10, wherein the number of consecutive columns of the array of memory cells to be read represents a number of individual data chunks requested by the host system, and wherein the memory device is to combine the number of individual data chunks to form a combined data chunk to be transferred by the gapless data burst.

12. The memory controller of claim 8, wherein the plurality of column access operations comprises three column access operations, and wherein a third column access operation of the plurality of column access operations has the shorter duration than the second column access operation.

13. The memory controller of claim 8, wherein at least the portion of the data is transferred over a data bus having an input / output (I / O) width that is not an integer power of two.

14. The memory controller of claim 8, wherein a default burst length representing how many data words can be transferred over the data bus in a single burst is misaligned with a total size of the portion of the data to be transferred.

15. A memory module comprising: one or more memory devices; and a memory interface chip coupled to the one or more memory devices via one or more communication links, wherein the memory interface chip is configured to: receive a request to access a portion of data stored on the one or more memory devices; and initiate a plurality of column access operations to transfer the portion of the data over a data bus coupled to the memory interface chip, wherein a first column access operation of the plurality of column access operations has a shorter duration than a second column access operation the plurality of column access operations, and wherein the plurality of column access operations are performed successively in response to the request to generate a gapless data burst.

16. The memory module of claim 15, wherein the plurality of column access operations comprises three column access operations, and wherein a third column access operation of the plurality of column access operations has the shorter duration than the second column access operation.

17. The memory module of claim 15, wherein to initiate the plurality of column access operations, the memory interface chip is to generate one or more memory access commands comprising respective indications of a number of consecutive columns of an array of memory cells of the one or more memory devices to be read to generate the gapless data burst.

18. The memory module of claim 17, wherein the number of consecutive columns of the array of memory cells to be read represents a number of individual data chunks requested bya host system, and wherein the individual data chunks are combined to form a combined data chunk to be transferred by the gapless data burst.

19. The memory module of claim 15, wherein the data bus has an input / output (I / O) width that is not an integer power of two.

20. The memory module of claim 15, wherein a default burst length representing how many data words can be transferred over the data bus in a single burst is misaligned with a total size of the portion of the data to be transferred.

Citation Information

Patent Citations

  • Burst clock control based on partial command decoding in a memory device

    US11211103B1

  • Memory Signal Buffers and Modules Supporting Variable Access Granularity

    US20130036273A1

  • Memory control component with inter-rank skew tolerance

    US20200349991A1

  • Double fetch for long burst length memory data transfer

    US20210286561A1

  • Semiconductor memory devices and methods of operating the same

    US20230124660A1