Threshold voltage distribution data acquisition method and solid state drive

CN122470429BActive Publication Date: 2026-09-15MAXIO TECHNOLOGY (HANGZHOU) CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610959061.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-09-15
Estimated Expiration
2046-06-29

AI Technical Summary

Technical Problem

[0004]然而,上述技术方案存在明显的缺陷:首先,数据采集严重滞后,从错误发生到固件触发中断并执行采集任务,两者时间间隔往往较长,在此期间,由于NAND型存储单元的初始电压漂移效应、环境温度变化、数据保持特性等因素,阈值电压分布信息受到污染,导致采集到的数据无法真实反映错误发生时的现场状态,严重增加了分析难度;其次,若尝试在错误发生时立即执行采集,由于需要对发生不可纠错误的位置进行上千次不同读偏置电压下的读取操作,任务繁重且集中,将长时间占用大量资源,导致长达秒级的处理延迟,使用户端感受到明显的操作卡顿,影响正常使用

Benefits of technology

[0016] Compared to existing technologies, this application generates multiple independent read tasks with different read bias voltages for uncorrectable errors and stores them in a task pool. Utilizing both firmware idle-time scheduling and timeout-forced scheduling mechanisms, it ensures that the original data acquisition is completed within the contamination window of the storage unit where the uncorrectable error occurs. This effectively avoids information contamination caused by factors such as the IVS effect, temperature changes, and data retention, ensuring the on-site authenticity and timeliness of the acquired original data. In a further embodiment, the task pool is divided into a storage pool and an execution pool: read tasks in the storage pool can be scheduled to the execution pool or the storage medium. The execution pool can obtain read tasks from the storage medium or the storage pool. These operations can be executed in parallel, thereby avoiding the risk of cache unit overflow due to read task accumulation. Simultaneously, the dual-pool strategy improves the real-time response of the system. Furthermore, by performing read tasks in units of the smallest read unit (typically 4K), the instantaneous occupation of firmware resources is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122470429B_ABST
    Figure CN122470429B_ABST
Patent Text Reader

Abstract

Provided is a method for acquiring threshold voltage distribution data, comprising: in response to an uncorrectable error occurring in a read operation, generating a plurality of read tasks using different read bias voltages for the storage unit in which the uncorrectable error occurs, and storing the plurality of read tasks in a task pool; polling the state of firmware, if the firmware is in a busy state and the busy duration exceeds a preset waiting time, forcibly starting a processing task, otherwise, starting the processing task when the firmware is idle; counting the number of bit 1 in the original data, and flushing the statistical value and the original data as threshold voltage distribution data from the cache unit to the storage medium for subsequent generation of a threshold voltage distribution graph. By polling the firmware state, the processing flow is immediately launched when the firmware is idle, and a preset waiting time is set as an upper limit for waiting, so that threshold voltage distribution data can be collected within the pollution window period of the flash memory grain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of solid-state drive technology, specifically to a method for obtaining threshold voltage distribution data and a solid-state drive. Background Technology

[0002] For solid-state drives (SSDs), especially NAND SSDs, obtaining the threshold voltage distribution information of storage cells is crucial during firmware development and troubleshooting. This information can effectively reflect the data storage status and allow for analysis of the threshold voltage distribution of the storage medium and whether any abnormalities exist.

[0003] In existing technologies, a mainstream method for obtaining threshold voltage distribution data is as follows: When an uncorrectable error occurs during a read operation of the solid-state drive (SSD), the firmware records the physical location information of the flash memory corresponding to the error. After the firmware is interrupted or stops running due to an anomaly, a dedicated data retention task is started to scan and collect data from the previously recorded error location, and the obtained raw data (without ECC correction) is flushed to the storage medium. Finally, the data is extracted from the storage medium, and after offline calculation and plotting, the required threshold voltage distribution map is obtained.

[0004] However, the above technical solutions have obvious drawbacks: First, data acquisition is severely delayed. The time interval between the occurrence of an error and the firmware triggering an interrupt and executing the acquisition task is often long. During this period, due to factors such as the initial voltage drift effect of NAND storage cells, changes in ambient temperature, and data retention characteristics, the threshold voltage distribution information is contaminated, resulting in the acquired data failing to accurately reflect the on-site state at the time of the error, which greatly increases the difficulty of analysis. Second, if an attempt is made to execute acquisition immediately when an error occurs, thousands of read operations at different read bias voltages are required at the location where the uncorrectable error occurred. The task is heavy and concentrated, which will occupy a large amount of resources for a long time, resulting in a processing delay of up to seconds. This causes users to experience obvious operational lag and affects normal use. Summary of the Invention

[0005] To address the aforementioned issues, this application provides a method for obtaining threshold voltage distribution data and a solid-state drive.

[0006] In a first aspect, a method for acquiring threshold voltage distribution data is applied to a solid-state drive, wherein the firmware responsible for handling read and write operations in the controller resides in a cache unit, and the acquisition method includes: In response to an uncorrectable error occurring during a read operation, multiple read tasks with different read bias voltages are generated for the memory cell where the uncorrectable error occurred, and the multiple read tasks are stored in a task pool. The firmware status is polled. If the firmware is busy and the busy duration exceeds a preset waiting time, a processing task is forcibly started. Otherwise, when the firmware is idle, a processing task is started. The processing task executes the read task to obtain threshold voltage distribution data temporarily stored in the cache unit. The preset waiting time is less than the pollution window. The threshold voltage distribution data is refreshed from the cache unit to the storage medium for subsequent generation of the threshold voltage distribution map.

[0007] In some embodiments, the processing task includes: requesting firmware resources, executing at least one read task, counting the number of bits 1 in the raw data obtained by each read task, and using the count value and the raw data as the threshold voltage distribution data.

[0008] In some embodiments, the task pool is divided into a storage pool and an execution pool, and the acquisition method includes scheduling read tasks in the storage pool and the execution pool, including: When the number of read tasks accumulated in the storage pool reaches the upper limit and the execution pool is not empty, the read tasks in the storage pool are refreshed to the storage medium, and the number of refreshes in the storage pool is accumulated. When the execution pool is empty and the number of refreshes in the storage pool is zero, the read tasks in the storage pool are scheduled to the execution pool. When the execution pool is empty but the storage pool refresh count is not empty, read tasks in the storage medium are preferentially scheduled to the execution pool.

[0009] In some embodiments, generating multiple read tasks using different read bias voltages for the memory cell where the uncorrectable error occurred includes: Determine the total number of read bias voltages required to obtain the threshold voltage distribution based on the type of memory cell where the uncorrectable error occurred; Based on the total number of read bias voltages, a corresponding number of read tasks are generated, and each read task corresponds to a unique read bias voltage.

[0010] In some embodiments, the size of the storage cell where the uncorrectable error occurs is equal to the smallest read cell in a host read operation.

[0011] In some embodiments, the preset waiting time is determined based on the duration of the contamination window of the storage cell where the uncorrectable error occurred and the number of read tasks generated for the storage cell where the uncorrectable error occurred.

[0012] In some embodiments, the forced-start processing task processes only one read task in the execution pool at a time, and the processing task started when the firmware is idle continues to process the read tasks in the execution pool until all read tasks in the execution pool are processed or the firmware is no longer idle.

[0013] In some embodiments, the storage pool is refreshed to the storage medium using a single-bit programming method.

[0014] Secondly, embodiments of this application provide a solid-state drive, including: a coupled controller and a storage medium, wherein the controller executes the acquisition method described above.

[0015] Thirdly, embodiments of this application provide a controller for a solid-state drive, the controller being coupled to a storage medium and executing the acquisition method described above.

[0016] Compared to existing technologies, this application generates multiple independent read tasks with different read bias voltages for uncorrectable errors and stores them in a task pool. Utilizing both firmware idle-time scheduling and timeout-forced scheduling mechanisms, it ensures that the original data acquisition is completed within the contamination window of the storage unit where the uncorrectable error occurs. This effectively avoids information contamination caused by factors such as the IVS effect, temperature changes, and data retention, ensuring the on-site authenticity and timeliness of the acquired original data. In a further embodiment, the task pool is divided into a storage pool and an execution pool: read tasks in the storage pool can be scheduled to the execution pool or the storage medium. The execution pool can obtain read tasks from the storage medium or the storage pool. These operations can be executed in parallel, thereby avoiding the risk of cache unit overflow due to read task accumulation. Simultaneously, the dual-pool strategy improves the real-time response of the system. Furthermore, by performing read tasks in units of the smallest read unit (typically 4K), the instantaneous occupation of firmware resources is reduced. Attached Figure Description

[0017] The above and other objects, features and advantages of the present application will become clearer from the following description of embodiments of the present application with reference to the accompanying drawings, in which: Figure 1 This is a schematic block diagram of the host system; Figure 2 This application is for use with Figure 1 The flowchart shown illustrates a method for obtaining threshold voltage distribution data on a solid-state drive. Figure 3 This is a schematic block diagram of the method for obtaining threshold voltage distribution data using a dual-pool strategy proposed in this application. Detailed Implementation

[0018] The following description of embodiments of this application is based on examples, but the embodiments of this application are not limited to these embodiments. In the detailed description of the embodiments of this application below, some specific details are described in detail. Those skilled in the art can fully understand the embodiments of this application without these details. To avoid obscuring the essence of the embodiments of this application, well-known methods, processes, and flows are not described in detail. Furthermore, the accompanying drawings are not necessarily drawn to scale.

[0019] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions, and operations of the systems, methods, and apparatuses according to embodiments of this application. The blocks in the flowcharts and block diagrams may represent a module, program segment, or simply a piece of code. These modules, program segments, and code are all executable instructions used to implement a specified logical function. It should also be noted that the executable instructions implementing the specified logical function can be recombined to generate new modules and program segments. Therefore, the blocks and their order in the accompanying drawings are only used to better illustrate the processes and steps of the embodiments and should not be construed as limiting the invention itself.

[0020] The following concepts are involved in this article.

[0021] Threshold voltage (Vth): For NAND flash memory cells (typically floating gate transistors), this refers to the gate voltage required to switch it from the off state to the on state. The actual Vth value of a memory cell determines whether it stores "0" or "1" (for SLC) or more states (for MLC / TLC / QLC).

[0022] Threshold Voltage Distribution: Due to factors such as manufacturing process deviations, wear and tear, data retention, and temperature variations, the Vth value of a large number of memory cells is not a fixed value, but rather exhibits a probability distribution. A reference voltage (Vread) is applied to the word line to determine whether the memory cell is conducting. If the Vth of the memory cell is lower than Vread, the cell is conducting (bit 1); otherwise, it is turning off (bit 0). By systematically scanning a series of continuously changing Vread voltages and recording the statistical values ​​of bits 0 and 1 in the raw data read at each voltage, the Vth distribution of the memory cell group can be deduced.

[0023] The contamination window represents the time from the occurrence of an uncorrectable error (UNC) to before the threshold voltage distribution information is contaminated.

[0024] This application may be presented in various forms, some of which will be described below.

[0025] Figure 1 An example diagram of a host system 100 is shown. The host system 100 is, for example, a personal computer, a laptop, or a server.

[0026] The host system 100 includes a host device 110 and a storage device, the storage device consisting of a controller 120 and a storage medium 130. The host device 110 issues host commands to the storage device, enabling the storage device to manage host data stored in the storage device according to the commands. For example, the host device 110 issues read (READ) or write (WRITE) commands to the storage device through the host interface 121, and the storage device reads or writes host data according to the address range indicated by the read or write command.

[0027] The controller 120 includes a host interface 121, a processor 123, a cache unit 124, and a storage medium interface 128. The host interface 121 of the controller 120 is connected to the host device 110 and is used to cache and transmit host commands and host data. The processor 123 is connected to the host interface 121, the cache unit 124, and the storage medium interface 128; the processor 123 parses host commands and executes corresponding operations. The storage medium interface 128 includes interface circuitry for data transmission between the controller 120 and the storage medium 130.

[0028] Cache unit 124 is a key component of controller 120. Its core function is to act as a temporary high-speed storage area for data and instructions, including but not limited to: temporarily storing data to be written from the host and data to be returned from the storage medium; maintaining critical metadata such as the mapping table of frequently accessed logical addresses to physical addresses in the Flash Translation Layer (FTL); and loading and executing various programs (such as programs in the Flash Translation Layer (FTL)) as working memory. Cache unit 124 may use SRAM and / or DRAM (DRAM is typically not part of the controller), but due to manufacturing cost limitations, the capacity of cache unit 124 is usually limited.

[0029] In response to the problems raised in the background art, this application proposes a method applied to Figure 1 The method for obtaining threshold voltage distribution data on a solid-state drive (SSD) is shown. This method can generate firmware that resides in cache unit 124. Cache unit 124 also contains firmware responsible for processing read / write commands from host device 110. Given that the firmware resources provided by the SSD have an upper limit (e.g., the SRAM size that the firmware can use cannot exceed a certain limit), the two firmware types need to compete for the resources allocated to the firmware. Under this premise, the flowchart of this acquisition method is as follows: Figure 2 As shown.

[0030] In step S201, in response to an uncorrectable error occurring during a read operation, multiple read tasks using different read bias voltages are generated for the memory cell where the uncorrectable error occurred, and the multiple read tasks are stored in a task pool, which is set in a cache unit.

[0031] In step S202, the status of the firmware responsible for processing read and write commands from the host device 110 is polled. If the firmware is in a busy state and the busy duration exceeds a preset waiting time, the processing task is forcibly started. Otherwise, when the firmware is idle, the processing task is started. The processing task includes requesting firmware resources and executing at least one read task in the task pool, counting the number of bits 1 in the raw data of each read task, and storing the raw data and statistical values ​​obtained by each read task in the corresponding cache unit.

[0032] In step S203, the statistical values ​​and the original data are refreshed from the cache unit to the storage medium as threshold voltage distribution data for subsequent generation of threshold voltage distribution map.

[0033] For example, consider a Quad-Level Cell (QLC) memory cell. It has 16 different charge states, each representing a 4-bit binary value (e.g., 0000, 0001). Since these 16 states are characterized by different voltage levels (or threshold voltages) within the memory cell, a decision voltage point needs to be set between two adjacent voltage states to determine the current state during a read operation. This point is called a read level / read threshold. There are 15 read voltages, and each read voltage performs 120 read operations, requiring a total of 1800 reads. Because host reads data in 4KB units, uncorrectable errors triggered by host reads only require collecting the threshold voltage distribution data of the current 4KB unit, not the threshold voltage distribution data of the entire physical page. Therefore, for each uncorrectable error encountered during host reading, 1800 read tasks are generated for the 4K unit where the uncorrectable error occurred. This minimizes the firmware resources consumed by the read tasks. These 1800 read tasks are stored in the task pool in the cache unit 124. Then, a preset waiting time is set, for example, 5ms, and the status of the firmware (responsible for handling read and write commands from the host device 110) is continuously polled. When the firmware's busy state lasts for more than 5ms, a processing task is forcibly started. The forcibly started processing task will preempt firmware resources and use them to process a small number of read tasks in the task pool. If the firmware's busy state does not last for 5ms and becomes idle, a normal processing task is started. The processing task normally requests firmware resources, uses firmware resources to continuously process each read task in the task pool and obtains the raw data of the read task. It counts the number of bits 1 and / or bits 0 in the raw data of each read task, and combines the statistical value with the raw data to form threshold voltage distribution data, which is temporarily stored in the cache unit 124, until all read tasks in the task pool are processed or until the firmware state is no longer idle. Finally, the threshold voltage distribution data in cache unit 124 is refreshed from cache unit 124 to storage medium 130 for subsequent generation of threshold voltage distribution map.

[0034] Under normal circumstances, when the firmware (responsible for handling read and write commands from host device 110) generates multiple uncorrectable errors while processing host read operations, the read tasks will be queued in the task pool to prioritize the response time of the host.

[0035] Generally, steps S203 and S202 are performed alternately. For example, after each preset number of read tasks in step S202 are completed, step S203 is executed once to refresh the threshold voltage distribution data stored in cache unit 124 from cache unit 124 to storage medium 130.

[0036] In some embodiments, a preset waiting time is determined based on the duration of the contamination window of the memory cell where the uncorrectable error occurred and the number of read tasks generated for the memory cell where the uncorrectable error occurred. For example, if the contamination window duration of a 4K cell is 1 second and 1800 read tasks are generated, then the preset waiting time is approximately 1 second / 1000.

[0037] Figure 3 A schematic block diagram is given for the method of obtaining threshold voltage distribution data using the dual-pool strategy proposed in this application.

[0038] As shown in the figure, the storage pool 301 and execution pool 302 work together. Storage pool 301 receives and temporarily stores multiple read tasks, each corresponding to a unique read bias voltage. Each read task is identified by a task identifier plus the read bias voltage; for example, tasks D0 to DN in the figure represent read tasks targeting the same location with different read bias voltages. Execution pool 302 and storage pool 301 are located in cache unit 124, and each can only store a fixed number of task information.

[0039] As shown in the figure, steps S10 to S12 indicate that after the firmware of controller 120 sends the read command to storage medium 130, it receives an uncorrectable error from storage medium 130. At this time, the firmware decomposes the corresponding location information, generates multiple simple read tasks using different read bias voltages, and stores them in storage pool 301.

[0040] Step S13 is to check whether the number of read tasks accumulated in storage pool 301 has reached the upper limit. If the number of read tasks accumulated in storage pool 301 has reached the upper limit and execution pool 302 is not empty at this time, the read tasks in storage pool are refreshed to storage medium 130, and the number of storage pool refreshes is accumulated.

[0041] Step S14 is to check whether execution pool 302 is empty. If execution pool 302 is empty and the number of times the storage pool is refreshed is 0, then the read tasks in storage pool 301 are scheduled to execution pool 302.

[0042] Step S15 is to check whether execution pool 302 is empty. If execution pool 302 is empty and the number of times the storage pool is refreshed is greater than 0, then the read tasks in storage medium 130 are scheduled to execution pool 302 first.

[0043] Step S16 involves continuously polling the status of the firmware responsible for processing read / write commands from host device 110 to initiate a processing task: if the firmware busy time has remained at a preset waiting time (e.g., 250ms), the processing task is forcibly started; otherwise, in the idle state after the firmware becomes busy, the processing task is started normally. The processing task is responsible for executing at least one processing task in execution pool 302. More specifically, the processing task executes at least two of steps S17 to S20. S17 is to request firmware resources. S18 is to execute task A0, the operations of which include setting the read bias voltage 0, obtaining raw data based on the read bias voltage 0, and counting the number of bits 1 in the raw data. Step S19 transfers the raw data and the statistical value to the storage medium. Step S20 releases the requested firmware resources.

[0044] To avoid impacting firmware responsiveness to the host, forced-start processing tasks can execute only one read task, while processing tasks started during the firmware's idle period can execute as many read tasks as possible until all read tasks in execution pool 302 are completed or the firmware is detected to be busy again. To achieve this, execution pool 302 employs a task scheduling mechanism. When the firmware has an idle time slice, execution pool 302 immediately requests firmware resources to execute a single read operation. If the firmware remains busy, execution pool 302 proactively initiates a task execution at preset waiting times (e.g., every 0.25 seconds) based on a timer to prevent tasks from being blocked for extended periods. After each read task is completed, execution pool 302 immediately releases the occupied firmware resources.

[0045] Meanwhile, to improve system efficiency, a single-bit programming method can be used during refresh operations to refresh the read task information in storage pool 301 to storage medium 130. The single-bit programming method treats multi-level storage units (such as QLC and MLC) as single-level storage units (SLC), writing only one bit to each. Since single-bit programming is faster than multi-bit programming, it improves processing efficiency.

[0046] Accordingly, embodiments of this application also provide, for example... Figure 1 The solid-state drive controller and solid-state drive shown have readable computer instructions stored in the cache unit of the controller or the storage medium of the solid-state drive. When these instructions are executed by the processor, they can implement the acquisition method provided in the above embodiments.

[0047] Accordingly, embodiments of this application also provide a computer-readable storage medium for storing readable computer instructions of a software program constructed based on the acquisition method provided in embodiments of this application.

[0048] As used herein, the term "module" may refer to, be part of, or include the following: application-specific integrated circuits (ASICs), electronic circuits, processors (shared, dedicated, or grouped) and / or memories (shared, dedicated, or grouped) that execute one or more software or firmware programs, combinational logic circuits, and / or other suitable components that provide the described functionality.

[0049] Those skilled in the art will understand that the various modules or units of the data processing system according to the present invention can be implemented by hardware, firmware, or software. Software includes, for example, coded programs written in various programming languages ​​such as JAVA, C / C++ / C#, and SQL. Although the steps and their order are given in the methods and method diagrams of embodiments of the present invention, the executable instructions that implement the specified logical functions of the steps can be recombined to generate new steps. The order of the steps should not be limited to the order shown in the methods and method diagrams, and can be adjusted at any time according to functional needs. For example, some steps can be executed in parallel or in reverse order.

Claims

1. A method for acquiring threshold voltage distribution data, applied to a solid-state drive (SSD), wherein the firmware responsible for handling read / write operations in the SSD controller resides in a cache unit, and the acquisition method includes: In response to an uncorrectable error occurring during a read operation, multiple read tasks with different read bias voltages are generated for the memory cell where the uncorrectable error occurred, and the multiple read tasks are stored in a task pool. The status of the firmware is polled. If the firmware is in a busy state and the busy duration exceeds a preset waiting time, the processing task is forcibly started. Otherwise, when the firmware is idle, the processing task is started. The processing task executes the read task to obtain the threshold voltage distribution data temporarily stored in the cache unit. The preset waiting time is less than the pollution window. The pollution window represents the time from the occurrence of the uncorrectable error to before the threshold voltage distribution information is polluted. as well as The threshold voltage distribution data is refreshed from the cache unit to the storage medium of the solid-state drive for subsequent generation of the threshold voltage distribution map.

2. The acquisition method according to claim 1, wherein the processing task includes: The process involves requesting firmware resources, executing at least one read task, counting the number of 1 bits in the raw data obtained from each read task, and using the count and the raw data as the threshold voltage distribution data.

3. The acquisition method according to claim 1, wherein, The task pool is divided into a storage pool and an execution pool. The acquisition method includes scheduling read tasks in the storage pool and the execution pool, including: When the number of read tasks accumulated in the storage pool reaches the upper limit and the execution pool is not empty, the read tasks in the storage pool are refreshed to the storage medium, and the number of refreshes in the storage pool is accumulated. When the execution pool is empty and the number of refreshes in the storage pool is zero, the read tasks in the storage pool are scheduled to the execution pool. When the execution pool is empty but the storage pool refresh count is not empty, read tasks in the storage medium are preferentially scheduled to the execution pool.

4. The acquisition method according to claim 1, wherein, The process of generating multiple read tasks using different read bias voltages for the memory cell where the uncorrectable error occurred includes: Determine the total number of read bias voltages required to obtain the threshold voltage distribution based on the type of memory cell where the uncorrectable error occurred; Based on the total number of read bias voltages, a corresponding number of read tasks are generated, and each read task corresponds to a unique read bias voltage.

5. The method for obtaining according to claim 1, wherein, The size of the storage cell where the uncorrectable error occurs is equal to the smallest read cell in a host read operation.

6. The method for obtaining according to claim 1, wherein, The preset waiting time is determined based on the duration of the contamination window of the storage cell where the uncorrectable error occurred and the number of read tasks generated for the storage cell where the uncorrectable error occurred.

7. The method for obtaining according to claim 1, wherein, A forced-start processing task processes only one read task from the task pool at a time. A processing task started when the firmware is idle continuously processes read tasks in the task pool until all read tasks in the task pool are processed or the firmware is no longer idle.

8. The method for obtaining according to claim 3, wherein, The read tasks in the storage pool are flushed to the storage medium using a single-bit programming method.

9. A solid-state drive, comprising: A coupled controller and storage medium, the controller performing the acquisition method as described in any one of claims 1 to 8.

10. A controller for a solid-state drive, the controller being coupled to a storage medium and performing the acquisition method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Operating method of memory controller, memory device including memory controller, and operating method of memory device

    CN115129630A

  • Data recovery method and device, solid state disk, electronic equipment and storage medium

    CN117666967A