Memory device

By introducing multiple buffers and output circuits into the memory device, the access and output methods of data rows are optimized, and the problem of long reading delay of existing non-volatile memory devices is solved, thereby achieving higher operating speed and lower delay.

CN114171071BActive Publication Date: 2025-05-23ADESTO TECHNOLOGIES CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111512571.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2016-03-04
Filing Date
2017-02-28
Publication Date
2025-05-23
Estimated Expiration
2037-02-28

AI Technical Summary

Technical Problem

The existing nonvolatile memory devices have a long read delay, especially in communication between the microprocessor and the memory, resulting in limited system performance.

Method used

By introducing a plurality of buffers and output circuits into the memory device, the access and output mode of data rows are optimized, including using a first buffer and a second buffer to store data rows from different array planes respectively, and output data sequentially from the starting byte to the highest addressed byte through the output circuit.

Benefits of technology

It effectively reduces the read delay, improves the operating speed of the memory device, and adapts to the sensitive demands of the microprocessor for delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114171071B_ABST
    Figure CN114171071B_ABST
Patent Text Reader

Abstract

Memory device. A memory device may include: a memory array having memory cells arranged in data rows; an interface that receives a read command requesting data bytes in a sequentially addressed order from an address of a starting byte; a first buffer that stores a first data row from the memory array including the starting byte; a second buffer that stores a second data row from the memory array that is sequentially addressed relative to the first data row; an output circuit configured to access data from the buffer and sequentially output bytes from the starting byte through the highest addressed byte of the first data row and sequentially output bytes from the lowest addressed byte of the second data row until the requested data byte has been output; and a data strobe driver that clocks each data byte outputted via a data strobe on the interface.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application with the original application number 201780005814.6 (International application number: PCT / US2017 / 020021, application date: February 28, 2017, invention name: Reduction of read latency in memory devices). Technical Field

[0002] The present invention relates generally to the field of semiconductor devices. More particularly, embodiments of the present invention relate to memory devices, including both volatile and non-volatile memory devices, such as flash memory devices, resistive random access memory (ReRAM), and / or conductive bridge RAM (CBRAM) processes and devices. Background Art

[0003] Non-volatile memory (NVM) is increasingly found in applications such as solid-state hard drives, removable digital picture cards, etc. Flash memory is the mainstream NVM technology used today. However, flash memory has limitations such as relatively high power and relatively slow operating speed. Microprocessor performance is very sensitive to memory latency. Many non-volatile memory devices have relatively slow access times or latency compared to microprocessors. In addition, many implementations of various communication protocols (e.g., serial peripheral interface (SPI)) between microprocessors / hosts and memories can even add more latency than required by the memory array itself. Summary of the invention

[0004] An aspect of the present invention relates to a memory device, the memory device comprising: a) a memory array, the memory array comprising a plurality of memory cells arranged into a plurality of data rows, wherein each data row comprises a predetermined number of data bytes, and wherein the memory array comprises a first array plane and a second array plane; b) an interface, the interface being configured to receive from a host a read command requesting a plurality of data bytes in a sequentially addressed order from an address of a starting byte; c) a first buffer, the first buffer being configured to store a first data row of the plurality of data rows read from the first array plane of the memory array, wherein the first data row comprises the starting byte ; d) a second buffer configured to store a second data row of the plurality of data rows read from the second array plane of the memory array, wherein the second data row is addressed consecutively relative to the first data row; and e) an output circuit configured to access data from the first buffer and sequentially output each byte from the starting byte through the highest addressed byte of the first data row, wherein the output circuit is configured to access data from the second buffer and sequentially output each byte from the lowest addressed byte of the second data row until the predetermined number of data bytes have been output to execute the read command.

[0005] Another aspect of the present invention relates to a memory device, the memory device comprising: a) a memory array, the memory array comprising a plurality of memory cells arranged into a plurality of data rows, wherein each data row comprises a predetermined number of data bytes, and wherein the memory array comprises a first word line and a second word line; b) an interface, the interface being configured to receive from a host a read command requesting a plurality of data bytes in a sequentially addressed order from an address of a starting byte; c) a first buffer, the first buffer being configured to store a first data row of the plurality of data rows along the first word line of the memory array, wherein the first data row comprises the starting byte; d) a second buffer, the second buffer being configured for storing a second row of the plurality of rows of data along the first word line of the memory array, wherein the second row of data is addressed sequentially relative to the first row of data, and wherein the second row of data is repeated along the second word line of the memory array; and e) an output circuit configured to access data from the first buffer and sequentially output bytes from the starting byte through the highest addressed byte of the first row of data, wherein the output circuit is configured to access data from the second buffer and sequentially output bytes from the lowest addressed byte of the second row of data until the predetermined number of data bytes have been output to execute the read command. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Figure 1 is an example memory device and host arrangement according to an embodiment of the present invention.

[0007] Figure 2 is a block diagram of an example memory array and buffer arrangement for reading data in accordance with an embodiment of the present invention.

[0008] Figure 3 is a block diagram of an example data line and buffer arrangement according to an embodiment of the present invention.

[0009] Figure 4A and Figure 4B is a timing diagram of an example read access according to an embodiment of the present invention.

[0010] Figure 5 is a timing diagram of an example read access with reduced latency and data strobe timing in accordance with an embodiment of the present invention.

[0011] Figure 6 is a block diagram of an example memory device and host arrangement with data strobes and I / O paths in accordance with an embodiment of the present invention.

[0012] Figure 7 is a block diagram of an example data line and buffer arrangement for interleaved data line access according to an embodiment of the present invention.

[0013] Fig. 8A , Figure 8B and Figure 8C is a timing diagram of an example interleaved data row read access according to an embodiment of the present invention.

[0014] Fig. 9 is a block diagram of an example memory array and buffer arrangement with duplicate data rows for adjacent word lines in accordance with an embodiment of the present invention.

[0015] Fig.10 is a flow chart of an example method of reading data bytes from a memory array according to an embodiment of the present invention. DETAILED DESCRIPTION

[0016] Reference will now be made in detail to specific embodiments of the present invention, examples of which are shown in the accompanying drawings. Although the present invention will be described in conjunction with preferred embodiments, it will be understood that it is not intended to limit the present invention to these embodiments. On the contrary, the present invention is intended to cover substitutions, modifications and equivalents that may be included in the spirit and scope of the present invention as defined by the appended claims. In addition, in the following detailed description of the present invention, many specific details are set forth in order to provide a thorough understanding of the present invention. However, it will be readily apparent to those skilled in the art that the present invention can be practiced without these specific details. In other cases, known methods, processes, procedures, components, structures and circuits are not described in detail to avoid unnecessarily obscuring aspects of the present invention.

[0017] Some parts of the following detailed description are presented in terms of processes, procedures, logic blocks, function blocks, processing, schematic symbols and / or other symbolic representations of operations on data streams, signals or waveforms within a computer, processor, controller, device and / or memory. Technicians in the field of data processing typically use these descriptions and representations to effectively convey the essence of their work to other technicians in the field. Typically, although not necessarily, the amount being manipulated takes the form of an electrical, magnetic, optical or quantum signal that can be stored, transmitted, combined, compared and otherwise manipulated in a computer or data processing system. It has proven to be convenient sometimes (primarily for general reasons) to refer to these signals as bits, waves, waveforms, streams, values, elements, symbols, characters, terms, numbers, etc.

[0018] Particular embodiments may be directed to memory devices, including volatile memories (e.g., SRAM and DRAM) and including non-volatile memories (NVM) (e.g., flash memory devices) and / or resistive switching memories (e.g., conductive bridging random access memories [CBRAM], resistive RAM [ReRAM], etc.). Particular embodiments may include structures and methods of operating flash memories and / or resistive switching memories that can be written (programmed / erased) between one or more resistance and / or capacitance states. In a particular example, a CBRAM storage element may be configured such that an electrical property (e.g., resistance) of the CBRAM storage element may change when a forward or reverse bias greater than a threshold voltage is applied across electrodes of the CBRAM storage element. Regardless, certain embodiments are suitable for any type of memory device, particularly NVM devices such as flash memory devices, and in some cases may include resistive switching memory devices.

[0019] Now refer to Figure 1, shows an example memory device and host arrangement 100 according to an embodiment of the present invention. In this example, the host 102 can interface with the memory device 104 via a serial interface. For example, the host 102 can be any suitable controller (e.g., a CPU, an MCU, a general purpose processor, a GPU, a DSP, etc.), and the memory device 104 can be any type of memory device (e.g., SRAM, DRAM, EEPROM, flash memory, CBRAM, magnetic RAM, ReRAM, etc.). The memory device 104 can therefore be implemented in accordance with a variety of memory technologies, such as non-volatile types. In some cases, the memory device 104 can be a serial flash memory that can be implemented in a more traditional non-volatile memory or in a CBRAM / ReRAM resistive switch memory.

[0020] Various interface signals, such as in a serial peripheral interface (SPI), may be included for communication between the host 102 and the memory device 104. For example, a serial clock (SCK) may provide a clock to the device 104 and may be used to control the flow of data to the device. Commands, addresses, and input data (e.g., via I / O pins) may be latched by the memory device 104 on the rising edge of SCK, while output data (e.g., via I / O pins) may be clocked out from the memory device 104 via SCK or a data strobe (DS). A chip select (CS) (which may be active low) may be used to select the memory device 104, for example, from among multiple such memory devices sharing a common bus or circuit board, or as a way to access the device. When the chip select signal is de-asserted (e.g., at a high level), the memory device 104 may be de-selected and placed in a standby mode. Enabling the chip select signal (e.g., via a high-to-low transition on CS) may be used to start an operation, and returning the chip select signal to a high state may be used to terminate an operation. For internally self-timed operations (eg, program or erase cycles), if chip select is disabled during the operation, the memory device 104 may not enter the standby mode until the particular operation in progress is completed.

[0021] In an example interface, data may be provided to the memory device 104 (e.g., for write operations, other commands, etc.) and provided from the memory device 104 (e.g., for read operations, verify operations, etc.) via I / O signals. For example, input data on the I / O may be latched by the memory device 104 at an edge of SCK, and may be ignored if the device is deselected (e.g., when a chip select signal is disabled). Data may also be output from the memory device 104 via the I / O signals. For example, for timing consistency, data output from the memory device 104 may be clocked out at an edge of DS or SCK, and the output signal may be in a high impedance state when the device is deselected (e.g., when a chip select signal is disabled).

[0022] In one embodiment, a memory device may include: (i) a memory array having a plurality of memory cells arranged into a plurality of data rows, wherein each data row includes a predetermined number of data bytes; (ii) an interface configured to receive a read command from a host requesting a plurality of data bytes in a sequentially addressed order from an address of a start byte; (iii) a first buffer configured to store a first data row from the plurality of data rows from the memory array, wherein the first data row includes the start byte; and (iv) a second buffer configured to store a second data row from the plurality of data rows from the memory array, wherein the second data row is sequentially addressed relative to the first data row. ; (v) an output circuit configured to access data from a first buffer and sequentially output each byte from a starting byte through a highest addressed byte of a first data row; (vi) the output circuit configured to access data from a second buffer and sequentially output each byte from a lowest addressed byte of a second data row until a requested plurality of data bytes have been output to execute a read command; and (vii) a data strobe driver configured to clock each data byte output from a memory device on an interface through a data strobe, wherein the data strobe is enabled with reduced read latency when the starting address is aligned with the lowest addressed byte of the first data row.

[0023] Now refer to Figure 2, a block diagram of an example memory array and buffer arrangement for reading data according to an embodiment of the present invention is shown. For example, the memory device 104 may include a memory array 202 (e.g., a flash memory array) and buffers 204-0 and 204-1, which may be implemented in an SRAM or any other relatively fast access memory. In some arrangements, only one or more than two buffers 204 may be provided, such as multiple buffers for multi-layer buffering and deeper pipelining. The memory device 104 may be configured as a data flash memory and / or a serial flash memory device, and the memory array 202 may be organized into any suitable number or arrangement of data pages. The output circuit 206 may receive a clock signal and may perform various logic, multiplexing, and drive functions to drive I / O pins (e.g., 4, 8, or any other number of pins) and an optional data select pin (DS).

[0024] As used herein, a "data row" may be a group of data bytes that may include code for in-place execution and / or data used in code execution, or any other type of stored data. A data row may be a group of continuously addressed data bytes that can be accessed from a memory array in one memory access cycle and output from a memory device over multiple output cycles (e.g., 16 cycles, or 8 cycles of double data rate output) of a clock or data strobe. For example, memory cells in a data row may share a common word line and a selected bank of sense amplifiers. As a specific example, a data row may be equivalent to a cache row or data page that a host may request to fill. In addition, for example, a data row may be 16 bytes of data that are sequentially / continuously addressed. In addition, a data row may represent a boundary so that when a byte within a given data row is requested as part of a read operation, a subsequent memory array access to the next sequentially addressed data row may be used to fetch data for a complete data row (e.g., 16 sequential bytes) starting from the requested byte. In addition, in some cases, in addition to the number of bytes of data, a data row may also include additional bits.

[0025] Thus, in many cases, the two reads of memory array 202 may be performed prior to (e.g., pre-fetching) or in parallel with outputting data via output circuit 206. For example, data row 1000 (e.g., 16 bytes = 128b) may be accessed from memory array 202, provided to buffer 204-0, and output via output circuit 206. Then, data row 1010 may be accessed and provided to buffer 204-1 for output via output circuit 206. As labeled herein, data rows are identified by their example starting byte aligned addresses in hexadecimal. Thus, for a data row size of 16 bytes, "1000" may be the hexadecimal address of the lowest addressed byte of the corresponding data row (i.e., the byte corresponding to the lowest address of a given data row), and "1010" may be the hexadecimal address of the lowest addressed byte of the next sequentially addressed data row.

[0026] Buffering (e.g., via buffer 204) can be used to address memory array access latency and can allow chunks of 128b (e.g., data row size) to be output from the memory device every 8 clock cycles. For example, each of buffers 204-0 and 204-1 can store at least 128b of data. In standard SPI, there may be no way to notify the host 102 that buffer 204 may not have enough data (e.g., less than 128b of data) to satisfy the current read request (e.g., for a total of 16 bytes, from the starting address to the consecutive addressed bytes), and as a result, increased latency may occur. Therefore, 2 entities or data rows may be accessed (prefetched) in advance in a sequential and back-and-forth manner, such as providing data row 1000 to buffer 204-0 and then providing data row 1010 to buffer 204-1. This can ensure sufficient buffering to meet the output clock control requirements of the memory device. In this way, a read request can be issued by the host 102, for example, every 4 or 8 clock (e.g., SCK) cycles, and data can be efficiently streamed out sequentially (e.g., once the buffer 204 is sufficiently full) by prefetching, for example, 128b data chunks every 4 or 8 clocks, depending on the I / O and data line width / size configuration.

[0027] In an example operation, if a memory device receives a read request with a particular starting address byte of a 128b entity (e.g., a data row), such data may be output from the memory device, and a request or implied request may be sent from the host to read out the next sequentially / continuously addressed data row. If the read request includes a starting address toward the end of a given data row, the data (e.g., consecutively addressed bytes) that may be sequentially accessed from that data row may be insufficient, as will be discussed in more detail below. For example, one situation where only a single entity or data row needs to be accessed to satisfy a read request is when the first byte (i.e., the data byte at the lowest address) in a given data row is the starting address. For a 16-byte data row size, the probability of occurrence of this particular situation is 1 / 16.

[0028] However, due to utilizing this process of back-to-back reads from the memory array 202, a read latency bottleneck may occur. This bottleneck may be due to the requirement that the starting byte address can be any byte (byte-aligned addressing). In order to accommodate all addressing situations, including the extreme case where the last byte (i.e., the data byte at the highest address) of the sensed N bits (e.g., a data row) is requested as the starting byte, and then the first byte of the next N bits (e.g., the next continuously addressed data row) can be accessed, two memory array accesses must occur for each read request. In another approach, one or more mode bits may be used to change to word, double word, or even row aligned addressing, which can be used to increase the time between back-to-back reads and correspondingly reduce the apparent latency of the read operation.

[0029] Now refer to Figure 3 , a block diagram of an example data row and buffer arrangement according to an embodiment of the present invention is shown. This particular example shows that some memory (e.g., NVM) devices have relatively high read latency due to performing two read accesses before returning the first data item to the host (e.g., via output circuit 206). For a read operation with a starting byte address as shown in data row 1000, for example, the requested data (the amount may be equal to the amount stored in the data row (e.g., 16 bytes)) may include data bytes from the beginning or lower addressed byte portion of the next sequential (e.g., continuously addressed) data row 1010. Thus, two columns or data rows may be accessed, which effectively doubles the basic time for a read compared to a single data row access. As shown, this method accommodates starting byte addresses that appear anywhere within a data row, whereby sequentially addressed data bytes may overlap with data row boundaries to keep up with I / O speeds.

[0030] Now refer to Figure 4A and Figure 4B, shows a timing diagram of an example read access according to an embodiment of the present invention. In example 400, the starting address "X" may be equal to 1000 and may therefore be the first byte (e.g., the lowest addressed byte) of data row 1000. Accessing from memory array 202 is shown as accessing 402 data row 1000 (which may be provided to buffer 204-0) and then accessing 404 data row 1010 (which may be provided to buffer 204-1). Thus, buffer 204 may be filled via 406, and delay 408 may represent the access time from buffer 204 via the output via output circuit 206. For example, data 410 output at double data rate over 8 clock cycles may represent the complete data of data row 1000, and data 412 may represent the sequential / continuously addressed and less significant byte portions of data row 1010 to fill the read request. Thus in this example, 8 I / O lines may output a complete data row of 16 bytes of data and may be output via the DS strobe starting at 414.

[0031] Although the above examples show the starting byte address of the lowest addressed byte of a data row (e.g., 1000), example 450 shows the starting byte address as the last byte (e.g., the highest addressed byte) of a given data row. In this example, data 452 may represent data corresponding to the starting address (e.g., X=100F) contained within data row 1000. Additionally, data 454 may represent data from the next sequentially / continuously addressed data row 1010, and data 456 may represent data from the subsequent / sequentially addressed data row 1020. It should be noted that the data strobe for clocking out data is enabled at 414. Thus, in these examples, the same read latency occurs for the various starting addresses of a given data row, including Figure 4A The lowest byte address (X=1000) and Figure 4B The highest byte address (X=100F).

[0032] Now refer to Figure 5, a timing diagram 500 of an example read access with reduced latency and data strobe timing according to an embodiment of the present invention is shown. In a particular embodiment, when a read request is fully aligned (e.g., the starting byte address is aligned with the data row address), a single memory array read access can be performed, and / or the read latency (e.g., 404) can be reduced by not waiting for subsequent read accesses. Therefore, for a starting address that is byte aligned with a given data row, it may not be necessary to perform two complete reads from the memory array before starting to send data on the I / O line. Therefore, the requested data can be fully supplied, for example, by accessing 402 data row 1000 (which can be completed by 502). After the delay 504 propagates through the output circuit 206, the data is available at 510. Alternatively, the delay 504 can be less than the 2 cycles shown, or can be overlapped or pipelined with other data output processing so that the data is available at 502 (rather than at 510). In any case, the read latency can be reduced relative to the dual access method shown at 414 with DS enable timing discussed above.

[0033] In some embodiments, because the host sends an aligned address, the host 102 may know in advance that data (e.g., 506 representing data row 1000) is available, and / or DS may be used to communicate to the host that data is ready at 510. Even though in this particular example, the host may not need data from the next sequential data row 1010, at least a portion of the data may still be output at 508. In any case, DS may be relied upon not only for clocking data, but also for determining that data from the memory device is ready. Therefore, as part of its state machine function, the host may also use DS as a flow control signal to control the extraction of data by determining the data ready state. For example, a state machine in the host may count virtual cycles, etc. to determine whether data is available for reading from the buffer, and start collecting data from the memory device when available. Therefore, in some embodiments, DS may be used to clock data out, as well as provide a data ready indication to the host.

[0034] This approach can improve read latency (e.g., 1 / 16) for situations when the start or read request address is aligned with the beginning byte of the corresponding data row, and can be indicated by moving DS up (e.g., from 414 to 510) as shown for these situations. If the request address (i.e., the starting byte address) is naturally aligned with a "data row" that can be defined by the number of sense amplifiers (e.g., 16 bytes aligned in a device with 128 shared sense amplifiers), there can be a single memory array access before the data is returned to the host. In other (non-aligned) request situations, two memory array accesses can still be used as discussed above. In any case, DS can communicate data availability timing to the host by being enabled (e.g., toggled).

[0035] Control of the DS pin may also be used to support informing the host that the memory may need to pause data transfers on the I / O lines. This may be required when the memory may need additional latency due to "housekeeping" functions or any other reason. In some embodiments, DS may be used as a "back pressure mechanism" or "flow control mechanism" to inform the host when more time is needed (e.g., which may be accommodated by dummy cycles or other predefined wait states). For example, DS may stop switching while waiting for data to be fetched from the memory array, may be driven to a constant value when the address phase is complete, and may begin switching when the first data is ready to be output from the memory device.

[0036] In any case, the host may utilize DS (or SCK) switching in order to clock in data for receipt in the host device. In addition, in the case where a burst of data may not be sustained after the first batch of data (e.g., due to wraparound extraction), DS may be frozen until the memory device "recovers" from the wraparound operation and the data may then be streamed again. In wraparound extraction, "continuously addressed" data bytes may wrap around from the highest addressed byte to the lowest addressed byte within a given data row. It should be noted that on a memory device where the number of sense amplifiers enabled for a given memory array access matches the bus throughput, this "freeze" may only occur once (e.g., after the first batch of data is sent), and the probability of such a freeze in the case of sequential reads is relatively low. However, in reads that support wraparound functionality and depending on the cache line size, this probability may be slightly higher. In addition, as just one example, if DRAM is used in the memory implementation, a pause may be required to process a refresh operation.

[0037] In addition, in a specific embodiment, the variable DS function / timing can allow the memory device to reread in the event of a read error, which can potentially increase the maximum operating frequency. This is in contrast to, for example, operating the flash memory device at a frequency level that is essentially guaranteed to have no such data errors. Instead, higher frequencies can be allowed as long as the gain from such a frequency increase is higher than the time that may be lost in processing any rereads. In order to detect and correct read errors or other errors (for example, due to defective cells or radiation effects), a reread function and error correction code (ECC) can be used. An alternative to increasing the read speed is to reduce the read current, for example for a device that is not running at maximum speed. For example, this can be accomplished by using a lower read current, or by using a shorter read pulse at a lower cock speed. In this case, the variable DS can be used to reduce the total power consumption of reading at such a relatively low speed.

[0038] Now refer to Figure 6, a block diagram 600 of an example memory device and host arrangement with data strobe and I / O paths according to an embodiment of the present invention is shown. The figure shows an example timing propagation signal path including SCLK from the host 102, which can be used to clock data into the memory device 104. The output circuit 206 can receive the clock SCLK and can generate (e.g., switch) DS aligned with the data transition on the I / O line. The host 102 may include a receiver circuit 602, which can utilize DS to clock data in via the I / O line. This source synchronous clock control can be used to address clock skew and maintain the clock and data in phase. In addition, DS can be tri-stated by circuit 206 to, for example, accommodate bus sharing between multiple memory devices, but when no memory device is enabled, a system-level pull-down resistor can keep DS low.

[0039] In one embodiment, a memory device may include: (i) a memory array having a plurality of memory cells arranged into a plurality of data rows, wherein each data row includes a predetermined number of data bytes, and wherein the memory array includes first and second array planes; (ii) an interface configured to receive a read command from a host requesting a plurality of data bytes in a sequentially addressed order from an address of a starting byte; (iii) a first buffer configured to store a first data row from a plurality of data rows from a first array plane of the memory array, wherein the first data row includes the starting byte; (iv) a second buffer configured to store a second data row from a plurality of data rows from a second array plane of the memory array, wherein the second data row is sequentially addressed relative to the first data row; (v) an output circuit configured to access data from the first buffer and sequentially output each byte from the starting byte through the highest addressed byte of the first data row; and (vi) the output circuit configured to access data from the second buffer and sequentially output each byte from the lowest addressed byte of the second data row until the predetermined number of data bytes have been output to execute the read command.

[0040] Now refer to Figure 7, a block diagram 700 of an example data row and buffer arrangement for interleaved data row access according to an embodiment of the present invention is shown. In this case, the memory array 202 may include separate array "planes," "rows," "portions," or "regions." For example, one such plane (e.g., 702) may include even-addressed data rows (e.g., 1000, 1020, 1040, etc.) and another plane (e.g., 704) may include intermediate / odd-addressed data rows (e.g., 1010, 1030, 1050, etc.). In this way, any two sequential or consecutively addressed data rows may be found in separate array planes. Thus, for example, data row 1000 may be found in array plane 702, while the next sequential data row 1010 may be found in array plane 704.

[0041] Thus, in certain embodiments, the array may be organized into two separate (even and odd data row numbered) array planes. As discussed above, a data row may represent the number of bytes read by the memory array in a single memory array access and may be determined by the number of shared sense amplifiers or sense amplifiers enabled during such a memory access. For example, 128 or 256 sense amplifiers may be used to provide 16B of data throughput (128 bits in 8 cycles), resulting in a data row size of 16 bytes. With this configuration of even and odd data rows in separate array portions / planes, simultaneous reading from two array planes may be supported. However, in some cases, accesses may be interleaved to reduce noise, such as for initial filling of a buffer based on a requested starting byte address.

[0042] In this way, the memory can be interleaved so that consecutive data rows reside in alternating arrays. For example, if the data row size is 128 bits, then in the worst case when a read access targets one of the last four bytes of a given data row, the device may perform two such reads in the same first cycle (see, e.g. Figure 8C ). However, if the read target is a byte address between 0 and B (hexadecimal), the second access can be started after one or more cycles (see, for example Fig. 8A and Figure 8B ). Furthermore, in these cases, as discussed above, it may not be necessary to use DS as backpressure, since the number of virtual cycles can be fixed (but small).

[0043] In a particular embodiment, read latency can be reduced by performing two memory array accesses substantially in parallel. Data read from memory array plane 702 can be provided to buffer 204-0, and data read from memory array plane 704 can be provided to buffer 204-1. In addition, interleaving can be based on the number of bits read from the array in one array access cycle. For example, if 128 sense amplifiers are used in the array access, the memory array can be divided into two rows as shown, so that the 128-bit data rows addressed with even numbers reside in one row and the 128-bit data rows addressed with odd numbers reside in the other row. For Octal DDR operations that command addressing the second-to-last byte of a 128-bit data row, since data from the next data row that is one cycle later than the data from the first row may be needed, access to the second row can start a cycle after starting access to the first row.

[0044] In various cases, access to the second row may be delayed by 1 to 8 cycles, depending on the starting address. In other cases, both array planes may be accessed simultaneously in a fully parallel manner. Since data from the second row (e.g., 704) may be required to satisfy a read request in all cases except the aligned case where the command addresses the first (e.g., least significant) byte in the data row, read latency may be significantly improved in this approach. For example, in uniform addressing, this second array plane access may reduce read latency by a factor of 7 / 8. Additionally, throughput may be maintained without the need for additional sense amplifiers (e.g., maintained at 128 sense amplifiers) for reduced data row sizes (e.g., 64 bits) and other buffering applications.

[0045] Now refer to Fig. 8A , Figure 8B and Figure 8C , shows a timing diagram of an example interleaved data row read access according to an embodiment of the present invention. Fig. 8AAs shown in example 800 of , the starting address "X" may be equal to 1000 and may therefore be the lowest byte address of the data row 1000. Access from the memory array 202 is shown as accessing 802 the data row 1000 from the left array plane 702 (which may be provided to the buffer 204-0), then accessing 804 the data row 1010 from the right array plane 704 (which may be provided to the buffer 204-1), and then accessing 818 the data row 1020 from the left array plane 702 (which may be provided to the buffer 204-0). Thus, the buffer 204-0 may be filled by 806, and the delay 808 may represent the access time from the buffer 204-0 to the output via the output circuit 206. For example, the data 810 output at the double data rate over 8 clock cycles may represent the complete data of the data row 1000, and the data 812 may represent the sequentially addressed bytes of the data row 1010 to fill the read request. Thus in this example, 8 I / O lines may output a complete data row of 16 data bytes and may be output via the DS strobe beginning at 820. Additionally, buffer 204-1 may fill via 814 and delay 816 may represent the access time from buffer 204-1 to output via output circuit 206. Data 812 may represent partial data of data row 1010 that is addressed sequentially / continuously.

[0046] Figure 8B Example 830 shows that the starting byte address of a given data row (e.g., 1000) is not the lowest or highest byte address, but rather is somewhere in the middle (e.g., X=1008). In this example, access 832 may represent data from the memory array 202 corresponding to the starting address (e.g., X=1008) contained within the data row 1000 (which may be read from the left memory plane 702). Additionally, as shown, access 834 may represent data from the next sequentially addressed data row 1010 (which may be accessed at least partially in parallel with the right array plane 704), and access 836 may represent data from the continuously / sequentially addressed data row 1020 (which may be accessed from the left array plane 702). It should be noted that the data enable for clocking data out may be enabled at 856, which is in conjunction with Fig. 8A The same time point of 820.

[0047] Buffer 204-0 may be filled with data row 1000 via 838 and may be output as data 846 after delay 840. As shown, data 846 may represent consecutive data bytes starting at X=1008 until the end of data row 1000 or the highest addressed byte (e.g., 100F). Bytes from data row 1010 may be available via buffer 204-1 via 842 and may be output as data 852 after delay 844 as shown. Data from data row 1020 may be available in buffer 204-0 via 848 and may be output as data 854 after delay 850 as shown. As shown, in this particular example, interleaving accesses between an initial memory array access to data row 1000 containing starting address X=1008 from array plane 702 and subsequent accesses to data row 1010 from array plane 704 may allow for noise reduction. However, in some cases, these accesses may be performed in a fully parallel manner, whereby access 834 may begin at substantially the same time as access 832 .

[0048] Figure 8C Example 860 illustrates this parallel access method for the case where the starting byte address of a given data row (e.g., 1000) is the highest byte address of the data row (e.g., X=100F). In this example, access 862 may represent data from the memory array 202 corresponding to the starting address (e.g., X=100F) contained within the data row 1000 (which can be read from the left memory plane 702). In addition, as shown, access 864 may represent data from the next sequentially addressed data row 1010 (which can be accessed completely in parallel with the right array plane 704), and access 866 may represent data from the continuously and sequentially addressed data row 1020 (which can be accessed from the left array plane 702). In addition, a data strobe for clocking data out may be enabled at 882, which is in conjunction with Fig. 8A The same time point of 820.

[0049] Buffer 204-0 may be filled with data row 1000 and buffer 204-1 may be filled with data row 1010 via 868 and may be output as data 876 followed by data 878 after delay 870. As shown, data 876 represents the requested data byte at X=100F and data 878 represents the lowest addressed byte 1010 of data row 1010 up to the end or highest addressed byte (e.g., 101F). Bytes from data row 1020 may be available via buffer 204-0 via 872 and may be output as shown by data 880 after delay 874. Thus, in this particular example, fully parallel access may occur between memory array accesses from array plane 702 to data row 1000 containing starting address X=100F and memory array accesses from array plane 704 to data row 1010 followed by subsequent accesses of data row 1020.

[0050] In one embodiment, a memory device may include: (i) a memory array having a plurality of memory cells arranged into a plurality of data rows, wherein each data row includes a predetermined number of data bytes, and wherein the memory array includes first and second word lines; (ii) an interface configured to receive a read command from a host requesting a plurality of data bytes in a sequentially addressed order from an address of a starting byte; (iii) a first buffer configured to store a first data row among a plurality of data rows along a first word line of the memory array, wherein the first data row includes the starting byte; (iv) a second buffer configured to store a second data row among a plurality of data rows along the first word line of the memory array, wherein the second data row is sequentially addressed relative to the first data row, and wherein the second data row repeats along a second word line of the memory array; (v) an output circuit configured to access data from the first buffer and sequentially output each byte from the starting byte through the highest addressed byte of the first data row; and (vi) the output circuit configured to access data from the second buffer and sequentially output each byte from the lowest addressed byte of the second data row until the predetermined number of data bytes have been output to execute the read command.

[0051] Now refer to Fig. 9, a block diagram 900 of an example memory array and buffer arrangement with duplicate data rows for adjacent word lines according to an embodiment of the present invention is shown. In this particular example, the memory array 202 can be organized so that data is replicated at the end of one word line to match the data at the beginning of the next word line. For example, the data rows sharing a common WL 10 may include data row 1000, data row 1010, data row 1020, ... data row N, and data row 1100, where the next subsequent or consecutive data row 1100 also shares the same data at the beginning of the next word line (e.g., WL 11). As shown in the dashed box of data row 1100 at the end of WL 10, the replicated data can be configured so that only a single word line access may be required to satisfy a read request, regardless of the starting address of the read request.

[0052] Additionally, sense amplifiers 902-0 and 902-1 may be mapped to data rows along word lines so that one such amplifier bank or both amplifier banks 902 may be enabled to satisfy a given read request. Thus, for example, a given word line may be enabled to access data, and an amount of data consistent with the data row size (e.g., 128b) may be accessed in those bytes found along the common word line (e.g., 1Kb of data along a word line for 128b of data in a data row). The associated sense amplifier 902 may also be enabled to read the data and provided to the corresponding buffer 204. In this way, data from a starting point (starting address) onward (continuously addressed) to the data row size (e.g., 128b) may be accessed in one memory array access cycle so that the full amount of data (e.g., 128b) required to fill the associated buffer and satisfy the read request may be accommodated. Furthermore, the timing of such access may be consistent with the timing of starting or having a starting byte address aligned with the beginning of the data row.

[0053] Thus, in a particular embodiment, a word line stretch may be equal to the data row size (e.g., 128b), and the stretch may store a repetition of the first data row (e.g., 1100) of the next adjacent word line (e.g., WL 11). Thus, although the size of the memory array 202 is increased in this approach, a large read latency reduction may be achieved. However, one disadvantage of this approach is that two write cycles may be required in order to copy the appropriate data row (e.g., 1100). In this approach, reads may occur on a data row boundary (e.g., of 128b), so the first and second 128b data chunks (256b total) may be read in the same array access cycle. Additionally, the impact of writes on timing may be uncertain because the memory device may go into a busy state and then indicate when the write operation is complete, and in-line execution may be primarily a read operation.

[0054] In this approach, data "belonging to" two "columns" (rather than one) may be read, and, for example, 256 sense amplifiers may be enabled instead of 128 sense amplifiers. By adding a duplicate column or data row at the end of the array (e.g., identical to column 0 of WL+1), it may allow two rows of data to be read from the same WL. In this way, the memory array may only need to be read once to satisfy a read request, which may save about 50% of read latency in some cases. In an alternative approach, a single bank of sense amplifiers 902 (e.g., 128 sense amplifiers) may be retained, and two read cycles may be performed, for example, in such a case whereby reading two such columns from the same WL may not save a significant amount of time compared to reading one column from WLn and one column from WLn+1.

[0055] Now refer to Fig.10 , a flowchart 1500 of an example method for reading data bytes from a memory array according to an embodiment of the present invention is shown. At 1502, a read request may be received to read a plurality of bytes (e.g., equal to a data row size) from a memory array (e.g., 202) at a starting address of byte X. As discussed above, a memory array may be organized to accommodate interleaved accesses (see, e.g., Figure 7 ), with extended / copied data (see e.g. Fig. 9 ), or allow back and forth access (see e.g. Figure 2 ). Therefore, in certain embodiments, various memory array and data row configurations may be supported to reduce their read latency.

[0056] In any case, at 1504, a first data row containing byte X can be accessed from the memory array and stored in a buffer (e.g., 204-0). At 1506, a second data row that is sequential (e.g., adjacent, continuously addressed) to the first data row can be accessed and stored in another buffer (e.g., 204-1). Additionally, as described above with reference to Figure 7 As discussed, accesses 1504 and 1506 may be performed partially or completely in parallel. If, at 1508, byte X is the first byte or the lowest addressed byte of the first data line, then only the first data buffer (e.g., 204-0) need be used to satisfy the read request. In this case, at 1510, the bytes may be sequentially output from the first data line via the first buffer to satisfy the read request. An example of this is shown in Figure 5 Also as shown, the data strobe may be triggered in concert with the data output from the memory device to notify the host that the requested data is ready and to provide the clock with timing sufficient to receive / clock the data in the host.

[0057] If, at 1508, byte X is not the first least addressed byte of the first data row, the data required to satisfy the read request may be fetched across a data row boundary, thus requiring access to two data rows from memory array 202. In this case, at 1512, the first buffer (see, e.g., Figure 8B At 1514, bytes may be output from the second data row in sequential order via a second buffer (e.g., 204-1) until a number of bytes (e.g., the data row size) has been output from the memory device to satisfy the read request (see, e.g., Figure 8B 852). In this way, various memory array configurations can be supported so that through these configurations, and for appropriate situations (see, for example Figure 5 )Reduce read latency by switching the data strobe as soon as the data is ready.

[0058] Particular embodiments may also support options for operating on other byte boundaries (e.g., 2, 4, 8, etc.), which may allow for increased interface performance in some cases. Additionally, to accommodate higher interface frequencies, particular embodiments may support differential input (e.g., SCK) and output (e.g., DS) clocks (e.g., using an external reference voltage). Additionally or alternatively, synchronous data transfer may involve an option to specify a number of dummy cycles, which may define the earliest time that data may be returned to the host. However, if the controller (e.g., host 102) is able to process the data immediately, this value may be kept at the minimum setting, and the memory device may output data as quickly as possible.

[0059] When data is received, the host controller can count incoming DS pulses, continue clock control until as many DS clocks as expected have been received, and can no longer rely on counting the SCK clocks generated by the host. For example, a minimum number of wait states can be set in a register, such as a mode byte for specifying a minimum dummy cycle. The host can also stop the outgoing SCK for a number of cycles to give itself time to prepare for the incoming data. In one case, if operating at a relatively low frequency, the minimum number of dummy cycles can be 0. In a variable setting, in some cases, a read command can have 0 wait states up to a certain frequency, and then have one or more dummy cycles.

[0060] Although the above examples include circuits, operations and structural implementations of certain memory arrangements and devices, those skilled in the art will recognize that other technologies and / or architectures may be used depending on the implementation. In addition, those skilled in the art will recognize that other device circuit arrangements, architectures, components, etc. may also be used depending on the implementation. For purposes of illustration and description, the above description of a specific embodiment of the present invention has been presented. It is not intended to be exhaustive or to limit the present invention to the precise form disclosed, and it is apparent that many modifications and variations may be made in accordance with the above teachings. The embodiments are selected and described in order to best illustrate the principles of the present invention and its practical application, so that other persons skilled in the art can best utilize the present invention and various embodiments with various modifications suitable for the specific purposes contemplated. The scope of the present invention is intended to be limited by the appended claims and their equivalents.

Claims

1. A memory device, the memory device include: a) a memory array comprising a plurality of memory cells arranged into a plurality of data rows, wherein each data row comprises a predetermined number of data bytes, and wherein the memory array comprises a first array plane and a second array plane; b) an interface configured to receive from a host a read command requesting a plurality of data bytes in a sequentially addressed order from an address of a starting byte; c) a first buffer configured to store a first row of data among the plurality of rows of data read from the first array plane of the memory array, wherein the first row of data includes the start byte; d) a second buffer configured to store a second row of data from the plurality of rows of data read from the second array plane of the memory array, wherein the second row of data is addressed consecutively with respect to the first row of data; and e) an output circuit configured to access data from the first buffer and sequentially output bytes from the starting byte through the highest addressed byte of the first data row, wherein the output circuit is configured to access data from the second buffer and sequentially output bytes from the lowest addressed byte of the second data row until the predetermined number of data bytes have been output to execute the read command.

2. The memory device according to claim 1, in, The predetermined number of data bytes in each data row is 16 bytes. 3 . The memory device of claim 1 , further comprising a plurality of sense amplifiers enabled to read the predetermined number of data bytes.

4. The memory device according to claim 1, in: a) the memory array comprises non-volatile memory; and b) The interface comprises a serial interface.

5. The memory device according to claim 1, in, The first data row and the second data row are accessed from the first array plane and the second array plane completely in parallel.

6. The memory device according to claim 1, in, The first data row and the second data row are accessed partially in parallel from the first array plane and the second array plane.

7. The memory device according to claim 1, in, The first data row and the second data row are accessed sequentially from the first array plane and the second array plane.

8. A memory device, the memory device include: a) a memory array comprising a plurality of memory cells arranged into a plurality of data rows, wherein each data row comprises a predetermined number of data bytes, and wherein the memory array comprises a first word line and a second word line; b) an interface configured to receive from a host a read command requesting a plurality of data bytes in a sequentially addressed order from an address of a starting byte; c) a first buffer configured to store a first row of data among the plurality of rows of data along the first word line of the memory array, wherein the first row of data includes the start byte; d) a second buffer configured to store a second row of data from the plurality of rows of data along the first word line of the memory array, wherein the second row of data is addressed consecutively with respect to the first row of data, and wherein the second row of data is repeated along the second word line of the memory array; and e) an output circuit configured to access data from the first buffer and output bytes sequentially from the starting byte through the highest addressed byte of the first data row, The output circuit is configured to access data from the second buffer and sequentially output bytes from the lowest addressed byte of the second data row until the predetermined number of data bytes have been output to execute the read command.

9. The memory device according to claim 8, in, The predetermined number of data bytes in each data row is 16 bytes.

10. The memory device according to claim 8, in: a) the memory array comprises non-volatile memory; and b) The interface comprises a serial interface.

11. The memory device of claim 8, further comprising a plurality of sense amplifiers enabled to read the first row of data and a plurality of sense amplifiers enabled to read the second row of data.

12. The memory device according to claim 8, in, At least one of the first word line and the second word line is enabled to access the first data row and the second data row.

13. The memory device according to claim 8, in, The repeated second row of data is the lowest addressed row of data along the second word line of the memory array.

14. The memory device according to claim 8, in, The repeated second row of data is the highest addressed row of data along the first word line of the memory array.

Citation Information

Patent Citations

  • Phase-Change and Resistance-Change Random Access Memory Devices and Related Methods of Performing Burst Mode Operations in Such Memory Devices

    US20100124102A1

  • Memory devices and methods having instruction acknowledgement

    WO2015176040A1