Data reading method, flash memory particle, flash memory chip, storage device, electronic device and computer program product
By dividing the outer clock signal of the flash memory particles, the second clock signal is formed to synchronize the reading operation in the flash memory particles, the problem of interplane timing delay in the multi-plane reading operation is solved, and the data transmission efficiency is significantly improved.
Patent Information
- Application Number
- CN202510308907.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-17
Smart Images

Figure CN119806436B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a data reading method, a flash memory particle, a flash memory chip, a storage device, an electronic device and a computer program product. Background Art
[0002] As we all know, unlike the mechanical structure of "head + motor + disk" used by traditional mechanical hard disks (HDDs), solid-state drives (SSDs) use a semiconductor storage chip structure of "flash media + master control". Compared with traditional mechanical hard disks, solid-state drives have the advantages of good performance, low power consumption, shock and drop resistance, low noise, and small size. When using solid-state drives for storage, the data stored by users on the solid-state drive will eventually be stored in non-volatile storage media. The characteristics of the storage medium determine the master control design and firmware design of the solid-state drive. The typical storage medium used by solid-state drives is flash memory, which is a non-volatile memory.
[0003] Generally speaking, each flash chip (Target) is controlled by an independent chip select signal CE#, including multiple flash memory particles (flash memory Die or flash memory LUN), each flash memory particle has multiple planes (Plane), each plane includes multiple flash memory blocks, each flash memory block includes multiple flash memory pages, each flash memory page corresponds to a word line and is composed of thousands of basic storage units (for example, NMOS floating gate transistors). In a flash memory particle, the smallest unit of reading and writing is a flash memory page. On the other hand, in order to increase the concurrency of reading flash memory and improve the flash memory reading performance, multi-plane read operation (Multi-PlaneRead) is used to replace single-plane read operation (Single-Plane Read). For multi-plane read operations, flash memory page data on different planes will be loaded into their respective page caches within one flash memory read time. In this way, data of multiple flash memory pages can be read in one read time, and data reading is accelerated.
[0004] Figure 7 A schematic diagram showing the main components of a flash memory chip for multi-plane read operation, Figure 8 and Fig. 9 FIG. 1 shows a timing diagram of data reading based on an existing multi-plane reading operation, and only shows a timing diagram of the part where data is output from a storage device. Figure 7 As shown, it is assumed that Plane0, Plane1, Plane2 and Plane3 shown in the figure are all objects of this multi-plane reading operation, and it is assumed that Plane0 is the first plane for data reading and Plane1 is the second plane for data reading.
[0005] Under this assumption, Figure 8This is a timing diagram of the data of a flash memory page (set as flash memory page A) of the first plane, i.e., plane Plane0, being transmitted inside a flash memory particle. Specifically, the target device connected to the storage device sends a read request command to the storage device, and these read request commands request to read the data in multiple planes. After receiving these read request commands, the main control unit of the storage device determines the flash memory particle as the read target after parsing by the FTL module. Then, the main control unit of the storage device sends the corresponding read request to the flash memory particle. These read request commands are transmitted to the SRAM inside the flash memory Die, and then further sent from the SRAM to each plane. After receiving these read request commands, each plane determines the corresponding flash memory page, and first caches the data stored in the flash memory page to its own page cache (first-level cache). On the other hand, the crystal oscillator and the phase-locked loop (PLL) or delay-locked loop (DLL) built into the flash memory particle generate the first clock signal CLK_0. The first clock signal CLK_0 is first sent to the data buffer BF (second-level cache) located between each page cache and the I / O interface (for example, ONFI interface) for issuance. Then, the first clock signal CLK_0 is sent from the data buffer BF to all planes to be read. Here, the time from the first clock signal CLK_0 to the first data of the flash memory page A (to be precise, its corresponding page cache) of the first plane Plane0 is sent to the data buffer BF is set to tD_0. In this way, the data of the page cache corresponding to the flash memory page A is sent to the data buffer BF in sequence under the action of the first clock signal CLK_0 until the data cached in the data buffer BF reaches the maximum cache amount. In this process, the input of the external clock signal CLK_1 of the master control device of the storage unit to the data buffer BF is prohibited. When the data buffer BF is full, the input of the external clock signal CLK_1 is allowed. The external clock signal CLK_1 is divided by the clock division module to form a second clock signal CLK_2 with a lower frequency. The second clock signal CLK_2 is used to synchronize the data buffer BF, the planes as the data reading object, and the data transmission between the I / O interface. The reason for dividing the frequency of the external clock signal CLK_1 is that the data transmission rate on the data bus outside the flash memory particle is high, and correspondingly, the data transmission rate inside the flash memory particle is low. However, the external clock signal CLK_1 is designed based on the data transmission rate on the data bus outside the flash memory particle and cannot be used for data transmission inside the flash memory particle. Therefore, after the external clock signal CLK_1 enters the flash memory particle, the external clock signal CLK_1 needs to be down-converted (i.e., divided) to make the divided clock signal compatible with the data transmission rate inside the flash memory particle.Since the frequency of the first clock signal CLK_0 is different from the frequency of the second clock signal CLK_2, in order to ensure the pipeline operation in the use of the data buffer BF, at the same time as the data output starts, the second clock signal CLK_2 will also take over the work of the first clock signal CLK_0, and be responsible for reading all the data in the flash memory page A of the first plane Plane0 until all the data in the flash memory page A is read. When the data buffer BF is full, under the action of the second clock signal CLK_2, the data in the data buffer BF is output to the I / O interface, and further sent to the outside of the flash memory particle via the I / O interface. Here, the time from the main control unit sending a read request to the first flash memory page, that is, the plane Plane0 where the flash memory page A is located, to the first batch of data from the cache of the flash memory page A to the data buffer BF starts to be transmitted to the outside of the flash memory particle is set as the initial fixed delay timing tCCS_0.
[0006] Fig. 9 This is a timing diagram of the data of a flash memory page (set as flash memory page B) of the second plane, i.e., plane Plane1, being transmitted inside the flash memory Die. After all the data of flash memory page A of the first plane Plane0 are output, the second clock signal CLK_2 returns the work to the first clock signal CLK_0, and the data buffer BF continues to send the first clock signal CLK_0. As a result, the data of flash memory page B of the second plane Plane1 cached in the corresponding page cache are sequentially cached to the data buffer BF under the action of the first clock signal CLK_0 until the data buffer BF is full. In this process, the input of the outer clock signal CLK_1 to the data buffer BF is also prohibited. When the data buffer BF is full, the input of the outer clock signal CLK_1 is allowed. Then, in the same way as flash memory page A, the reading of the data in flash memory page B is completed under the action of the second clock signal CLK_2. Here, the time from the last data cached in the data buffer BF of the flash memory page A of the first plane Plane0 being output externally and the first clock signal CLK_0 being re-issued to the first data cached in the data buffer BF of the flash memory page B of the second plane Plane1 being output externally and the input of the external clock signal CLK_1 being allowed is set as the inter-plane delay timing tCCS.
[0007] However, in the above-mentioned existing multi-plane read operation, as described above, there are some fixed timing delays on the data bus of, for example, a NAND flash memory. These fixed timing delays will cause the data transmission efficiency on the data bus of the flash memory to decrease. Specifically, for example, in the latest ONFI protocol, the data transmission rate of the data bus located outside the flash memory particle and communicatively connected to the I / O interface of the flash memory particle is 2.4GBps, which can reduce the data output time of each Byte to 0.42ns. Therefore, for a 16KB flash memory page, even with 2KB of redundant data, the output time of an 18KB flash memory page is 7.5μs. In the above-mentioned existing multi-plane read operation, after one plane is output and before the data of the next plane starts to be output, there is the above-mentioned inter-plane timing delay tCCS. Take NAND flash memory as an example for explanation. Since the frequency inside the flash memory particle of the NAND flash memory is lower than the frequency of the external bus, it is usually necessary to set up an additional secondary cache (i.e., the above-mentioned data buffer BF) to convert the low-frequency data stream into a high-frequency data stream. Due to the requirements of signal integrity, there is usually only one data buffer BF in the circuit, which is shared by all planes of a flash memory particle. In this case, the data of the previous plane has been output from the data buffer BF, and the data of the next plane needs to be input into the data buffer BF after certain preparation. That is, when the two planes are connected, it is necessary to wait for the previous plane to release the data buffer BF before the next plane can take over and use this data buffer BF. The time for this release and takeover is the above-mentioned inter-plane timing delay tCCS. The tCCS is about 0.25μs in NAND flash memory, which will cause the transmission efficiency of the entire bus to be reduced by about 3.33%, and the transmission efficiency will increase with the increase of the bus rate. For example, with a bus capacity of 4.8GBps, the above tCCS will cause a 6.67% reduction in bus efficiency. As a result, the actual data throughput efficiency is lower than the expected efficiency. Summary of the invention
[0008] Technical problem to be solved by the invention
[0009] The present invention is formed to solve the above-mentioned technical problems, and its purpose is to provide a data reading method, flash memory particles, flash memory chips, and storage devices, electronic devices and computer program products using the reading method, which can greatly reduce the inter-plane timing delay generated when the two planes of the flash memory are connected during the reading process, and can effectively improve the data transmission efficiency on the high-speed bus.
[0010] Technical solutions adopted to solve technical problems
[0011] The present invention provides a data reading method, wherein the data is read from a storage device to an object device connected to the storage device, wherein the storage device comprises a main control unit and a flash memory chip having a plurality of flash memory particles, each of the flash memory particles having a data buffer and a plurality of planes, and the method comprises:
[0012] A read request step, in which a read request command is sent from the target device to the storage device, wherein the read request command requests to read data in multiple planes of one flash memory particle;
[0013] A generating and sending step, in which the main control unit generates an outer clock signal based on the read request command, and sends the outer clock signal to the inside of one of the flash memory particles;
[0014] A frequency division step, in which the external clock signal entering one of the flash memory particles is frequency-divided to form a second clock signal;
[0015] a caching step, in which the second clock signal is sent to a plane of one of the flash memory particles to cache the data of the one plane into the data buffer; and
[0016] A judgment execution step, in which, when it is judged that the remaining cache amount of the data cached in the data buffer of the one plane is less than a non-zero cache amount threshold, or when it is judged that the remaining output time of the data cached in the data buffer of the one plane is less than a non-zero time threshold, the second clock signal is sent to another plane of the flash memory particle to cache the data of the other plane to the data buffer.
[0017] According to the data reading method described in the above technical solution, the above fixed timing delay generated when the two planes of the flash memory are handed over during the reading operation can be greatly shortened. Specifically, as described above, in the prior art, during the handover process of data reading of one plane and data reading of the next plane, after all the data of the one plane in the data buffer are output, the data of the next plane begins to be cached in the data buffer, thereby causing the above-mentioned longer inter-plane timing delay tCCS. In order to overcome the above technical problems, in the above technical solution of the present application, the outer clock signal sent from the outside of the flash memory particle is divided to form a clock signal with a lower frequency (i.e., a second clock signal) for reading data from a plane of the flash memory part, and the second clock signal is used to synchronize the reading-related operations in the flash memory part particle. In addition, in order to perform the "plane flash memory page → data buffer" operation in the flash memory part by the storage side clock signal formed by dividing the outer clock signal, the input of the outer clock signal to the inside of the flash memory particle is always allowed. In this way, the next plane can be allowed to send data to the data buffer for caching before the last batch of data cached in the data buffer of the previous plane is completely output. Specifically, since the second clock signal always synchronizes the operation of the data buffer, the flash memory unit and the I / O interface, when the remaining cache amount or the remaining output time of the data cached in the data buffer meets certain conditions, under the action of the second clock signal, the next plane can send data to the data buffer without having to wait until the amount of data in the data buffer becomes zero. Therefore, compared with the above-mentioned prior art, the fixed delay timing tCCS can be greatly shortened. For example, if the tCCS generated in the above-mentioned prior art is about 250ns, the tCCS generated based on the present technical solution can be controlled within 20ns.
[0018] Preferably, it is determined whether a portion of the data cached from the one plane to the data buffer is overwritten by the data cached from the other plane to the data buffer. If it is determined that a portion of the data cached from the one plane to the data buffer has been overwritten by the data cached from the other plane to the data buffer, the target device sends the read request command to the storage device again to request to re-read the data in the one plane.
[0019] According to the data reading method described in the technical solution, sometimes due to improper selection of the cache threshold or time threshold, the data cached in the data buffer of the previous plane may be overwritten by the data from the next plane before being output from the data buffer, thereby causing an error in the data output of the previous plane. As a result, the object device cannot obtain the correct data, causing an error in the object device. To this end, it is judged whether a part of the data cached from the previous plane to the data buffer has been overwritten by the data from the next plane. If it is judged to have been overwritten, the object device re-issues a read request command for the data in the plane to re-acquire the data in the plane. Therefore, the above-mentioned abnormal situation can be prevented from occurring.
[0020] Preferably, the buffer amount threshold and the time threshold are determined according to a physical transmission distance of the another plane relative to the data buffer.
[0021] According to the data reading method described in the technical solution, it is possible to avoid as much as possible that the data cached in the data buffer of the previous plane is overwritten by the data from the next plane. Specifically, due to factors such as process, design, and layout, the delay from the clock signal sent from the data buffer to each plane to the first data of each plane being cached in the data buffer is different. In addition, these delays mainly depend on the physical transmission distance of the data buffer relative to each plane. Therefore, by determining the cache amount threshold and time threshold based on the physical transmission distance of each plane relative to the data buffer, it can at least ensure that the output of the last data cached in the data buffer of the previous plane is staggered in time with the input of the first data of the next plane, thereby avoiding data overwriting.
[0022] Preferably, the buffer amount threshold and the time threshold are determined according to a physical transmission distance between a plane among a plurality of planes that is closest to the data buffer and the data buffer.
[0023] According to the data reading method described in the technical solution, it is possible to avoid to the greatest extent possible that the data cached in the data buffer of the previous plane is overwritten by the data from the next plane. Specifically, since the data of the plane with the shortest physical transmission distance from the data buffer has the shortest delay in reaching the data buffer, by setting the cache threshold and time threshold according to the physical transmission distance between the plane and the data buffer, it is possible to fully ensure that the output of the last data cached in the data buffer of the previous plane is staggered in time with the data of the first data of the next plane, thereby avoiding the occurrence of data overwriting to the greatest extent possible.
[0024] Preferably, the storage device is a solid state drive, and the flash memory chip is a NAND flash memory chip.
[0025] The data buffer is any one of a static random access memory, a dynamic random access memory, and a double rate synchronous dynamic random access memory.
[0026] The present invention also provides a flash memory particle, the flash memory particle includes an I / O interface, multiple planes and a data buffer, characterized in that:
[0027] The plurality of planes are configured to cache the data in each of the plurality of planes into a page cache of each of the plurality of planes when a read request command for reading the data in the plurality of planes is received,
[0028] The data buffer is configured such that, when data in each of the plurality of planes is cached in a page cache possessed by each of the plurality of planes:
[0029] receiving a second clock signal, where the second clock signal is formed by dividing the frequency of an external clock signal sent from outside the flash memory particle;
[0030] sending the second clock signal to one of the plurality of planes to cache data of the one plane into the data buffer;
[0031] When it is determined that the remaining cache amount of data cached in the data buffer of the one plane is less than a non-zero cache amount threshold, or when it is determined that the remaining output time of data cached in the data buffer of the one plane is less than a non-zero time threshold, the second clock signal is sent to another plane among the multiple planes to cache the data of the other plane to the data buffer.
[0032] The present invention also provides a flash memory chip, characterized in that the flash memory chip includes the above-mentioned flash memory particles.
[0033] The present invention also provides a storage device, comprising a flash memory chip and a main control unit, wherein the flash memory chip has a plurality of flash memory particles, each of the flash memory particles has a data buffer and a plurality of planes, and is characterized in that:
[0034] The storage device further includes a counter or a timer, wherein the counter calculates the cache amount of the data cached in the data buffer, and the timer calculates the remaining output time of the data cached in the data buffer.
[0035] The main control unit is designed as follows:
[0036] sending a second clock signal to one of the plurality of flash memory particles according to a read request command sent from an object device connected to the storage device, so as to cache data of one of the plurality of planes of the one flash memory particle to the data buffer, wherein the read request command requests to read the data in the plurality of planes, and the second clock signal is formed by dividing the frequency of an outer clock signal generated by the main control unit;
[0037] Determine whether the remaining buffer amount of the one plane cached in the data buffer calculated by the counter is less than a non-zero buffer amount threshold, or determine whether the remaining output time of the data cached in the data buffer of the one plane calculated by the timer is less than a non-zero time threshold; and
[0038] When it is determined that the remaining cache amount is less than a non-zero cache amount threshold, or when it is determined that the remaining output time is less than a non-zero time threshold, the second clock signal is sent to another plane among the multiple planes of the one flash memory particle to cache the data of the another plane to the data buffer.
[0039] The present invention also provides an electronic device, characterized in that it comprises:
[0040] The storage device mentioned above; and
[0041] The target device is connected to the storage device, and the target device sends the read request command and the external clock signal to the storage device.
[0042] A computer program product, comprising a computer program, characterized in that
[0043] When the computer program is executed by a processor, any of the above data reading methods is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 A functional schematic diagram showing data exchange between a solid state drive (SSD) as an example of a storage device and an operating system (OS) of an object device.
[0045] Figure 2 A schematic diagram of an electronic device for implementing a data reading method according to an embodiment is shown.
[0046] Figure 3 This is a timing diagram showing the transmission of two read commands for reading data in two different planes from a target device to a storage device. The clock signal for achieving synchronization is omitted in the diagram.
[0047] Figure 4It is a schematic diagram of a main module for data reading in a flash memory particle included in a storage device.
[0048] Figure 5 A timing diagram showing that data in the first flash memory page is cached in the data buffer and outputted externally via the data buffer.
[0049] Figure 6 A timing diagram showing data in the second flash memory page being cached in the data buffer and outputted externally via the data buffer.
[0050] Figure 7 A schematic diagram showing the main modules of a flash memory particle using a multi-plane read operation in the prior art is shown.
[0051] Figure 8 It is a timing diagram showing data reading of a flash memory page of the first plane based on an existing multi-plane read operation, and only shows the portion of the timing diagram where data is output from the storage device.
[0052] Fig. 9 A timing diagram of data reading of a flash memory page of the second plane based on an existing multi-plane read operation is shown, and only the timing diagram of the part where data is output from the storage device is shown. DETAILED DESCRIPTION
[0053] First refer to Figures 1 to 6 , a data reading method according to an embodiment of the present invention is described. More specifically, in this embodiment, a method for reading data from a storage device such as a solid state drive (SSD) through a multi-plane operation is described.
[0054] Figure 1 The functional diagram of data interaction between a solid-state drive SSD as an example of a storage device and the operating system OS of the target device is shown. The solid-state drive SSD mainly includes an input and output interface, a main control unit with a cache such as an on-chip SRAM, and a flash memory storage array composed of multiple flash memories (for example, NAND flash memories). First, starting from the target device, the user sends a request to the solid-state drive SSD from the operating system application layer. The file system converts the read request or write request into corresponding protocol-compliant read, write and other commands through the driver. The solid-state drive SSD receives the command, performs the corresponding operation, and then outputs the result. The input and output of each command are standardized by the protocol standard organization. From the perspective of functional composition, the solid-state drive SSD mainly includes three functional modules: the front-end interface and related protocol modules, the intermediate FTL (Flash Translation Layer) module, and the back-end and flash communication modules.
[0055] The following describes the write and read operations based on the solid state drive (SSD). It is known that the smallest unit of data transmission between the object device and the solid state drive (SSD) is a logical block. When the object device formats the solid state drive (SSD), the size of the data logical block has been determined. When performing a write operation, the object device sends a write request command to the fixed hard disk SSD via the input interface of the solid state drive (SSD). After receiving the write request command, the solid state drive (SSD) executes the command and receives the data to be written by the object device. Data is usually cached in an example of a cache inside the main control unit, namely RAM (on-chip SRAM). The FTL module assigns a flash address to each logical data block. When the data reaches a certain amount, the FTL module sends a write flash request to the back end, and then the back end writes the data in the cache to the corresponding flash space according to the write request. When performing a read operation, the object device sends a read request command and an instruction containing an address code to the solid-state drive SSD via the input interface of the solid-state drive SSD. The FTL module sends a read request to the corresponding data storage area in the flash memory unit according to the address code in the instruction. The back end caches the data in the flash memory space to an example of a cache inside the main control unit, namely, RAM, based on the read request. Then, the data is sent from the cache to the output interface and transmitted to the main memory (i.e., internal memory) of the object device via the data bus.
[0056] Figure 2A schematic diagram of an electronic device for implementing the data reading method of the present embodiment is shown. Specifically, data stored in the flash memory unit FU of the storage device 1 is read to the object device 2. The object device 2 includes a central processing unit CPU and a main memory M, and the central processing unit CPU includes an arithmetic unit and a controller (not shown). The main memory M includes a storage body M0 composed of multiple storage units, a memory address register MAR (Memory Address Register) and a memory data register MDR (Memory Data Register). The storage device 1 is a solid state drive (SSD), including a main control unit 3 including a flash memory controller, a flash memory unit 4, a connector 5 that acts as an input and output interface, an I / O interface 6 connected between the main control unit 3 and the flash memory unit 4, a power supply for power supply (not shown), and a DRAM (not shown) for storing L2P mapping tables and user write data, etc. The flash memory unit 4 has multiple flash memory chips, each of which has multiple flash memory particles, and the flash memory particles are basic units for receiving and executing commands. Each flash memory particle includes an I / O interface, multiple planes, and a data buffer BF located between the I / O interface and the multiple planes as a secondary cache. The connector 5 is, for example, a PCIe interface, which is controlled by a PCIe controller and an NVMe controller, NVMe being a protocol standard running on the PCIe interface. The central processing unit CPU and the main memory M of the storage device 1 are communicatively connected to the object device 2 through the root complex RC and via the control bus CB, the address bus AB and the data bus DB.
[0057] Next, refer to Figures 2 to 6 , the process of implementing the data reading method of this embodiment is explained. Figure 3 FIG. 1 is a timing diagram showing two read commands for reading data from two different planes, respectively, transmitted from the flash memory controller included in the main control unit 3 of the storage device 1 to the flash memory unit 4. The clock signal used to achieve synchronization is omitted in the diagram. In this embodiment, a multi-plane read operation is adopted. For the sake of convenience, Figure 3The multi-plane read operation for two planes is shown in FIG. 1 , but the number of planes is not limited thereto and may be more than three. First, the main control unit 3 of the storage device 1 generates multiple (two in this embodiment) read commands based on the read request command from the object device 2, and each of these read commands requests to read data from one of the multiple planes of the same flash memory particle of a flash memory chip of the flash memory unit 4. Each read command includes a specific logical block address, which includes a row logical address (Row Address) and a column logical address (Column Address), and the column logical address includes a flash memory particle logical address, a flash memory block logical address, and a flash memory page logical address, wherein the plane is addressed by the low-order address of the flash memory block logical address, and the row logical address represents the offset bit within the flash memory page. It should be noted that for NAND flash memory, the minimum unit of data reading is a flash memory page. As described above, each flash memory chip includes multiple flash memory particles, each flash memory particle includes multiple planes, each plane includes multiple flash memory blocks, and each flash memory block includes multiple flash memory pages. That is, each plane includes multiple flash memory pages, and the data reading operation is performed with the flash memory page as the minimum unit. In addition, each plane has its own independent cache, namely, a page cache, and the size of each page cache is usually equal to the size of a flash memory page. Since a multi-plane read operation is adopted in this embodiment, a flash memory page data on different planes of the same flash memory particle will be loaded into the respective page caches within a flash memory read time, and then the data in the page cache will be transferred to the data buffer BF in turn, and then transferred to the outside of the flash memory unit 4 in turn via the data buffer BF and the I / O interface 6. Specifically, the CPU of the object device 2 writes multiple (two in this embodiment) read commands for multi-plane read operations to the submission queue SQ (Submission Queue) located in the main memory M via RC. After completing the writing of these read commands, the CPU writes the doorbell queue DB (DoorBell register) of the submission queue SQ to notify the storage device 1 to fetch instructions. After using the doorbell queue DB to notify the storage device 1 to fetch instructions, the storage device 1 sends a request to the main memory M to read the above-mentioned multiple read commands written to the submission queue SQ. Under the action of the device-side clock signal generated by the CPU of the object device 2 for synchronizing the sending and receiving of these read commands, the read operation code and logical block address of the first read command are sent to the storage device 1 via the control bus CB and the address bus AB respectively. The first read command is a read command requesting to read a flash memory page A of a plane of a flash memory particle of the flash memory unit 4.As an NVMe read command, it includes a read opcode, a start logical block address (Start Logical Block Address), the amount of data to be read (for example, how many bytes), the memory address in the main memory M where the data to be read is stored, the numbering information of the read command in the submission queue, and the numbering information in the completion queue CQ described later. The start logical block address and the amount of data to be read (for example, how many bytes) constitute the logical page address of the flash memory page A. After the first read command enters the storage device 1 via the above-mentioned system bus and connector 5, it is temporarily stored in the command queue (i.e., the internal command queue) of the NVMe controller. Then, under the action of the device-side clock signal, the read opcode and logical block address of the second read command are sent to the storage device 1 via the control bus CB and the address bus AB, respectively, and the second read command is a read command requesting to read the flash memory page B of another plane of the same flash memory particle of the same flash memory unit 4. Similar to the first read command, after the second read command enters the storage device 1 via the above-mentioned system bus and connector 5, it is temporarily stored in the internal command queue of the NVMe controller. Next, the NVM controller sequentially transmits the first read command and the second read command to the CPU (firmware) system of the main control unit 3 through the internal command queue. On the other hand, the FTL (flash translation layer) module in the firmware system uses the pre-written L2P mapping table to map the logical block address in the first read command (to be precise, the logical page address of flash page A) and the physical address corresponding to the logical block address in the second read command (to be precise, the logical page address of flash page B). According to the ONFI protocol V5.1, each physical address includes a row address and a column address, and the row address includes a flash page address code, a flash block address code, and a flash particle address code. Flash pages A and B are respectively addressed by the low-order address of the flash block address code of their respective physical addresses, and the column address represents the address within the flash page (i.e., the offset within the page). That is, the row address in the physical address of flash page A includes the flash page address code R1 from the low bit to the high bit. A 、Flash memory block address code R2 A And the flash memory particle address code R3 A The column address from low to high includes the first column address code C1 A And the second column address code C2 A Similarly, the row address in the physical address of flash memory page B includes the flash memory page address code R1 from low to high. B 、Flash memory block address code R2 B And the flash memory particle address code R3 B The column address from low to high includes the first column address code C1 B And the second column address code C2 B .
[0058] like Figure 3As shown, the command 00h and the column address C1 of the flash memory page A are sent to start the first round of read commands. A and C2 A , row address R1 of flash page A A ~R3 A , and the command 30h indicating the preparation for sending the second round of read commands are sent to the flash controller back-end subsystem in sequence. Then, after a specified time interval (including tWB and tPLRBSY, see the ONFI protocol for specific meanings), the command 32h indicating the start of sending the second round of read commands, the column address C1 of flash page B, B and C2 B , row address R1 of flash page B B ~R3 B , and the command 30h indicating the preparation for the first round of data reading are sent to the flash memory controller back-end subsystem in sequence. Then, through the task scheduling module, data processing unit and flash memory driver of the flash memory controller back-end subsystem, the data of flash memory page A and flash memory page B are read into the corresponding page cache. That is, within one flash memory read time, the data of two flash memory pages, namely flash memory page A and flash memory page B, on different planes of the same flash memory particle are loaded into their respective page caches. After confirming that the data of each flash memory page is loaded into their respective page caches, the data in the above-mentioned flash memory page A and flash memory page B are read in sequence. Therefore, first, a command including "06h" and the column address C1 of flash memory page A are sent to the plane where flash memory page A is located. A and C2 A , row address R1 of flash page A A ~R3 A How the data cached in the page cache of the plane where the flash page A is located and the page cache of the plane where the flash page B is located are further cached in the shared data buffer BF in sequence will be discussed later in conjunction with Figure 4~Figure 6 Provide specific instructions.
[0059] Figure 4 This is a schematic diagram of the main components of the flash memory particle that is the read object. Figure 5 FIG. 1 shows a timing diagram of data in the first flash memory page, namely, flash memory page A (to be precise, data of flash memory page A cached in the corresponding page cache) being cached in the data buffer BF and outputted externally via the data buffer BF. Figure 4Four planes are shown, namely Plane0, Plane1, Plane2 and Plane4, but for ease of understanding and explanation, it is assumed that the planes that are the objects of the multi-plane read operation only include Plane0 and Plane1, and Plane0 includes flash memory page A, and Plane1 includes flash memory page B. After the data in flash memory page A and the data in flash memory page B are cached to the page cache of Plane0 and the page cache of Plane1 respectively based on the above read command, an external clock signal CLK_1 is sent to the flash memory Die that is the read object. After the external clock signal CLK_1 enters the interior of the flash memory particle through the I / O interface, it is divided by the clock division module to form a second clock signal CLK_2 with a lower frequency, and the second clock signal CLK_2 is sent to the data buffer BF, and then sent by the data buffer BF to reach each plane. Assume that the flash memory page A of a certain plane is the first object to be read. Since there is a certain physical distance between the data buffer BF and the plane where the flash memory page A is located, there is a certain delay from the issuance of the second clock signal CLK_2 to the arrival of the first data in the flash memory page A (to be precise, in the page cache corresponding to the flash memory page A) in the data buffer BF, which is recorded as tD_0. Under the action of the second clock signal CLK_2, the data in the page cache corresponding to the flash memory page A is sequentially cached to the data buffer BF until the data cache amount in the data buffer BF reaches the maximum cache amount. Figure 5 "D00, D01...D0X...D0N..." in the figure represents the data cached in the page cache corresponding to the flash memory page A. When the data cache amount in the data buffer BF reaches the maximum cache amount, under the action of the second clock signal CLK_2, the data in the data buffer BF is sequentially output to the outside of the flash memory particle via the I / O interface, and further transmitted from the connector 5 of the storage device 1 to the target memory address of the main memory M of the object device 2 via the data bus DB. Here, the time from the first issuance of the second clock signal CLK_2 to the first batch of data cached in the flash memory page A to the data buffer BF starts to be transmitted to the outside of the flash memory particle is set as the initial fixed timing delay tCCS_0.
[0060] Figure 6The timing diagram shows that the data in the second flash memory page, i.e., flash memory page B (to be precise, the data of flash memory page B cached in the corresponding page cache) is cached in the data buffer BF and outputted outwardly via the data buffer BF. Specifically, when a portion of the data in flash memory page A is still cached in the data buffer BF and certain conditions are met, the data buffer BF sends a second clock signal CLK_2 to the plane Plane1 where the second flash memory page, i.e., flash memory page B, is located. At the same time, the flash memory controller determines whether the data currently cached in the data buffer BF is the last batch of data of flash memory page A, and determines whether the cache amount of these data is less than a non-zero cache amount threshold or whether the remaining output time of these data is less than a non-zero time threshold. Specifically, a counter or a timer (not shown) is also provided in the storage device 1, and the counter calculates the cache amount of the last batch of data of the flash memory page cached in the data buffer BF, and the timer calculates the remaining output time of the last batch of data of the flash memory page cached in the data buffer BF. The cache amount threshold and the time threshold are determined in such a way that the remaining data of the current flash page (flash page A) cached in the data buffer BF will not be overwritten by the data from the next flash page (flash page B). As described above, due to factors such as process, design, and layout, the delay from the data buffer BF sending a clock signal to each plane and a read request to the first data cached in each plane to the data buffer is different. In other words, the size of the delay depends on the physical transmission distance of the data buffer BF relative to each plane. Therefore, it can also be said that the cache amount threshold and the time threshold are determined according to the physical transmission distance of each plane relative to the data buffer BF. In this embodiment, the case where the number of planes that are the objects of the multi-plane read operation is two is explained, therefore, the cache amount threshold and the time threshold can be determined according to the physical transmission distance of the data buffer BF relative to the plane Plane0 where the flash page A is located and the plane Plane1 where the flash page, i.e., the flash page B is located. That is, the cache amount threshold and the time threshold can be determined based on the smaller of the delay from the data buffer BF sending the second clock signal CLK_2 to the plane Plane0 where the flash memory page A is located to the first data in the flash memory page A (to be precise, its corresponding page cache) reaching the data buffer BF, and the delay from the data buffer BF sending the second clock signal CLK_2 to the plane Plane1 where the flash memory page B is located to the first data in the flash memory page B (to be precise, its corresponding page cache) reaching the data buffer BF. On the other hand, in the case where the number of planes used as the multi-plane read operation is three or more, it is preferred that the cache amount threshold and the time threshold are determined based on the physical transmission distance between the plane among the multiple planes that is closest to the data buffer B and the data buffer. Assume that the number of planes that are the objects of the multi-plane read operation is Figure 4 The four planes in the data buffer B are Plane0, Plane1, Plane2, and Plane3, and the delays of the above planes relative to the data buffer B are set to tD_0, tD_1, tD_2, and tD_3, respectively, and assuming that tD_0>tD_1>tD_2>tD_3, then the cache threshold and the time threshold are determined according to the minimum delay tD_3. In this way, it can be ensured that the remaining data of the previous flash memory page cached in the data buffer BF will not be overwritten by the data from the next flash memory page. When it is determined whether the cache amount of the last batch of data cached in the data buffer BF of the flash memory page A is less than the cache threshold determined according to the above method or whether the remaining output time of these data is less than the time threshold determined according to the above method, the data buffer BF sends a second clock signal CLK_2 to the plane Plane1 where the flash memory page B is located. Under the action of the second clock signal CLK_2, the data in the page cache corresponding to the flash memory page B is sequentially cached to the data buffer BF until the data cache amount in the data buffer BF reaches the maximum cache amount. Figure 6 The "D10, D11...D1X...D1N..." in the figure represent the data cached in the page cache corresponding to the flash memory page B. Then, also under the action of the second clock signal CLK_2, these data are transmitted from the data buffer BF to the outside of the flash memory unit 4 via the I / O interface 6, and further stored from the connector 5 of the storage device 1 to the target memory address of the main memory M of the object device 2 via the data bus DB. Here, the delay from the last data of the flash memory page A cached in the data buffer BF to the outside of the flash memory particle to the first data of the flash memory page B cached in the data buffer BF to the outside of the flash memory particle is set to the fixed inter-plane timing delay tCCS.
[0061] Next, continue to read data from the flash memory page of plane Plane0 that is different from the above-mentioned flash memory page A and from the flash memory page of plane Plane1 that is different from the above-mentioned flash memory page B in the above-mentioned manner until the read request command and other commands in the submission queue SQ are all completed. It should be noted that each time a command in the submission queue SQ is completed, the storage device 1 will write the command execution result to the completion queue of the storage body M0 that is also stored in the main memory M until all the commands in the submission queue SQ are completed. Then, the storage device 1 will send an interrupt notification to the object device 2 to notify the object device 2 that all commands have been completed. After receiving the interrupt notification, the object device 2 processes the command execution result, and after completing the processing, it replies to the storage device 1 through the doorbell DoorBell.
[0062] According to the data reading method described in this embodiment, by comparing with the corresponding Figure 6 The timing diagram shown and the corresponding prior art Fig. 9As shown in the timing diagram, compared with the prior art, the time for the two planes to handover is greatly shortened, and the new fixed inter-plane timing delay tCCS will be equal to (tD_Max-tD_Min) / CLK_2 and rounded up. This tCCS will be much less than 250nS and is expected to be within 20nS.
[0063] In addition, in this embodiment, the case where Plane0 where flash page A is located is used as the first plane for data reading and Plane1 where flash page B is located is used as the second plane for data reading is described as an example, but the reading order of the planes is not limited to this. Alternatively, Plane1 where flash page B is located is used as the first plane for data reading and Plane0 where flash page A is located is used as the second plane for data reading. The reading order of the planes is determined by the actual needs of the target device 2.
[0064] In addition, as the data buffer BF, a FIFO (First In First Out) type memory can be used, and various types of RAM can be used, such as static random access memory SRAM, dynamic random access memory DRAM or double data rate synchronous dynamic random access memory DDR.
[0065] In addition, considering that the setting of the cache amount threshold and / or the time threshold may be inappropriate, it may cause the data cached in the data buffer BF of the previous flash memory page to be overwritten by the data from the next flash memory page, thereby causing a data reading error. Therefore, in this embodiment, it is preferred to also determine whether at least one of the data cached in the data buffer BF of the previous flash memory page has been overwritten by the data from the next flash memory page. When it is determined that an overwrite occurs, the object device 2 sends the corresponding read request command to the storage device 1 again, requesting to re-read the data of the previous flash memory page.
[0066] In addition, in this embodiment, the storage device 1 is described as a solid state drive SSD as an example, but it is not limited to this. Any storage device with a similar structure to the storage device 1 can adopt the above data reading method.
[0067] It should be understood that the present invention is not limited to the exact construction that has been described above and shown in the drawings and that various modifications and changes may be made without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.
Claims
1. A data reading method, wherein data is read from a storage device to an object device connected to the storage device, wherein the storage device comprises a main control unit and a flash memory chip having a plurality of flash memory particles, each of the flash memory particles having a data buffer and a plurality of planes, characterized in that: include: A read request step, in which a read request command is sent from the target device to the storage device, wherein the read request command requests to read data in multiple planes of one flash memory particle; A generating and sending step, in which the main control unit generates an outer clock signal (CLK_1) based on the read request command, and sends the outer clock signal (CLK_1) to the inside of one of the flash memory particles; A frequency division step, in which the external clock signal entering one of the flash memory particles is frequency-divided to form a second clock signal (CLK_2); A caching step, in which the second clock signal (CLK_2) is sent to a plane of one of the flash memory particles to cache the data of the one plane into the data buffer; as well as A judgment execution step, in which, when it is judged that the remaining cache amount of the data cached in the data buffer of the one plane is less than a non-zero cache amount threshold, or when it is judged that the remaining output time of the data cached in the data buffer of the one plane is less than a non-zero time threshold, the second clock signal (CLK_2) is sent to another plane of the flash memory particle to cache the data of the other plane to the data buffer.
2. The data reading method according to claim 1, characterized in that: determining whether a portion of the data cached from the one plane to the data buffer is overwritten by the data cached from the other plane to the data buffer, When it is determined that a portion of the data cached from the one plane to the data buffer has been overwritten by the data cached from the other plane to the data buffer, the object device sends the read request command to the storage device again, requesting re-reading of the data in the one plane.
3. The data reading method according to claim 1, characterized in that: The buffer amount threshold and the time threshold are determined according to a physical transmission distance of the other plane relative to the data buffer.
4. The data reading method according to claim 1, characterized in that: The buffer amount threshold and the time threshold are determined according to a physical transmission distance between a plane among a plurality of planes and the data buffer that is closest in physical transmission distance to the data buffer.
5. The data reading method according to claim 1, characterized in that: The storage device is a solid state drive, and the flash memory chip is a NAND flash memory chip.
6. The data reading method according to claim 1, characterized in that: The data buffer is any one of a static random access memory, a dynamic random access memory, and a double rate synchronous dynamic random access memory.
7. A flash memory particle, comprising an I / O interface, a plurality of planes and a data buffer, characterized in that: The plurality of planes are configured to cache the data in each of the plurality of planes into a page cache of each of the plurality of planes when a read request command for reading the data in the plurality of planes is received, The data buffer is configured such that, when data in each of the plurality of planes is cached in a page cache possessed by each of the plurality of planes: receiving a second clock signal (CLK_2), where the second clock signal (CLK_2) is formed by dividing the frequency of an external clock signal sent from outside the flash memory particle; sending the second clock signal (CLK_2) to one of the plurality of planes to cache data of the one plane into the data buffer; as well as When it is determined that the remaining cache amount of the data cached in the data buffer of the one plane is less than a non-zero cache amount threshold, or when it is determined that the remaining output time of the data cached in the data buffer of the one plane is less than a non-zero time threshold, the second clock signal (CLK_2) is sent to another plane among the multiple planes to cache the data of the other plane to the data buffer.
8. A flash memory chip, characterized in that: Includes the flash memory particles described in claim 7.
9. A storage device, comprising a flash memory chip and a main control unit, wherein the flash memory chip has a plurality of flash memory particles, each of the flash memory particles has a data buffer and a plurality of planes, characterized in that: The storage device further includes a counter or a timer, wherein the counter calculates the cache amount of the data cached in the data buffer, and the timer calculates the remaining output time of the data cached in the data buffer. The main control unit is designed as follows: sending a second clock signal (CLK_2) to one of the plurality of flash memory particles according to a read request command sent from an object device connected to the storage device, so as to cache data of one of the plurality of planes of the one flash memory particle to the data buffer, wherein the read request command requests to read the data in the plurality of planes, and the second clock signal (CLK_2) is formed by dividing the frequency of an outer clock signal generated by the main control unit; Determine whether the remaining buffer amount of the one plane cached in the data buffer calculated by the counter is less than a non-zero buffer amount threshold, or determine whether the remaining output time of the data cached in the data buffer of the one plane calculated by the timer is less than a non-zero time threshold; as well as When it is determined that the remaining cache amount is less than a non-zero cache amount threshold, or when it is determined that the remaining output time is less than a non-zero time threshold, the second clock signal (CLK_2) is sent to another plane among the multiple planes of the one flash memory particle to cache the data of the other plane to the data buffer.
10. The storage device according to claim 9, characterized in that The buffer amount threshold and the remaining output time are determined according to a physical transmission distance between a plane among a plurality of planes and the data buffer that is closest in physical transmission distance to the data buffer.
11. An electronic device, characterized in that: include: The storage device according to claim 9 or 10; as well as The target device is connected to the storage device, and the target device sends the read request command to the storage device.
12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the data reading method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Sequential NAND type flash memory, flash memory device and operating method of sequential NAND type flash memory
CN104425014A
Data processing method, device and equipment and readable storage medium
CN118427121A