Processing unit, controller of storage device, and storage device

By designing processor solutions and co-processing units for multiple hardware threads in the controller of the storage device, the resource conflicts and waste problems when managing multiple flash memory particles are solved, and efficient flash memory control and performance improvement are achieved.

CN120196572APending Publication Date: 2025-06-24MAXIO TECHNOLOGY (HANGZHOU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510208206.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

When existing storage devices manage multiple flash memory particles under the same channel, they face problems such as resource conflicts, waste of hardware resources and system performance.

Method used

A processing unit is designed, using a processor scheme of multiple hardware threads, and multiple physical threads are designed in each channel. The instructions are prioritized by arbitration and allocation through the parsing unit. The co-processing unit specializes in processing atomic operations of flash operations, and multiple physical threads share the co-processing unit in time-sharing.

Benefits of technology

It realizes efficient control of multiple flash memory particles, reduces waste of hardware resources, improves system performance, and optimizes the utilization of chip resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196572A_ABST
    Figure CN120196572A_ABST
Patent Text Reader

Abstract

The invention discloses a processing unit, a controller of a storage device, and the storage device. The processing unit is applied to a controller of the storage device and comprises an instruction fetching unit used for reading a host command from an instruction space and transmitting the read host command to an analysis unit; the analysis unit is used for generating an instruction sequence based on an operation performed by executing a host command, arbitrating instruction priorities of a plurality of physical threads according to resource conditions required by execution of each instruction, and allocating each instruction to the corresponding physical thread based on an arbitration result, each instruction is a universal instruction or a flash memory operation instruction, and the instruction sequence is used for executing the host command. The flash memory operation instruction relates to atomic operation for input and output of flash memory particles; the execution unit is used for executing the flash memory operation instruction and the universal instruction in a physical thread; and the co-processing unit is specially used for processing atomic operations aiming at input and output of the flash memory particles, and the plurality of physical threads share the co-processing unit in a time-sharing manner. The processing unit can improve the processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of storage technologies, and particularly to a processing unit, a controller of a storage device, and a storage device. Background Art

[0002] In today's era, storage devices that use flash memory storage media (especially NAND flash memory media) such as solid-state drives have gradually become mainstream. As the demand for capacity in the market application of such products continues to climb, higher requirements are put forward for the integrated utilization of flash memory storage media. Usually, a single flash memory controller undertakes the important task of connecting and controlling a large number of flash memory particles. From the perspective of cost control, when designing the topology of the controller, an economical and efficient mode is often followed. Specifically, as Figure 1 shown, multiple channels are built inside each controller, and each channel is connected to and controls multiple flash memory particles to achieve the effective integration and efficient utilization of flash memory resources, and finally achieve the goal of saving chip costs while meeting the large-capacity storage requirements. Therefore, each channel of the controller shoulders the important mission of precisely controlling multiple flash memory particles.

[0003] However, when the controller controls multiple flash memory particles under the same channel, it faces many challenges. Since the internal operation speeds of each particle, whether it is the data read speed or the write speed, may have slight differences, it means that the chip selection of flash memory particles cannot be simply switched at fixed time intervals. In addition, in actual applications, data access requests do not appear in a fixed order or pattern. For example, in a server storage system, different application programs may read or write data on each flash memory particle at any time. This random access mode is difficult to predict and requires the controller to be able to respond quickly and reasonably arrange the operations of flash memory particles. If random access cannot be effectively processed, it may lead to an increase in access latency and affect the overall performance of the system.

[0004] Meanwhile, when multiple processors are used to manage multiple flash memory particles under the same channel in parallel, resource conflicts often occur during the operation process. At the same time, since each flash memory particle corresponds to a separate processor, the chip resources cannot be effectively reused, resulting in a large waste of hardware resources. For example, in a data storage system, if flash memory particle 1, flash memory particle 2, and flash memory particle 3 are connected under channel A, when three independent processors are used to manage these three particles respectively, once they perform operations such as programming, erasing, or reading simultaneously, conflicts may occur due to competing for channel bandwidth, control signal resources, etc. Moreover, from the perspective of resource utilization, each processor may be in an idle waiting state for most of the time, while the resources such as the chip area and power consumption occupied by it are continuously consumed, greatly reducing the overall utilization efficiency of the hardware resources, increasing the system cost and design complexity, and being unfavorable for the large-scale integration and optimization of the system.

[0005] In summary, how to manage the operations of multiple flash memory particles in the same channel is a relatively complex problem. Summary of the Invention

[0006] In view of this, embodiments of the present disclosure provide a processing unit to solve the problems in the prior art.

[0007] According to a first aspect of the embodiments of the present disclosure, a processing unit is provided, which is applied to a controller of a storage device. The controller is coupled to multiple flash memory particles through a first channel. The processing unit is equipped with multiple physical threads, and the processing unit includes:

[0008] An instruction fetch unit for reading a host command from an instruction space and transmitting the read host command to a parsing unit;

[0009] A parsing unit for generating an instruction sequence based on the operations performed for executing the host command, arbitrating the instruction priorities of the multiple physical threads according to the resource status required for each instruction execution, and allocating each instruction to the corresponding physical thread based on the arbitration result, where each instruction is a general instruction or a flash operation instruction, and the flash operation instruction involves atomic operations for input and output of the flash memory particle;

[0010] An execution unit for executing the flash operation instruction and the general instruction when the multiple physical threads are respectively executed;

[0011] A coprocessing unit dedicated to processing atomic operations for input and output of the flash memory particle, where the multiple physical threads share the coprocessing unit in a time-sharing manner.

[0012] In some embodiments, when the coprocessing unit processes the atomic operations for the input / output of the flash memory cells in the flash memory operation instruction, it locks the corresponding flash memory resources by using a thread lock.

[0013] In some embodiments, when the parsing unit decomposes the host command, it combines the waiting operation for the execution of the atomic operations for the input / output of the flash memory cells and the operation for judging the execution result of the atomic operations for the input / output of the flash memory cells in the flash memory operation instruction.

[0014] In some embodiments, the processing unit further includes timers respectively corresponding to a plurality of the physical threads, and the timers are used to time the waiting for the execution of the atomic operations for the input / output of the flash memory cells.

[0015] In some embodiments, the processing unit further includes a write-back unit, and the write-back unit writes the data returned by the coprocessing unit into a general register so that the parsing unit can obtain the data for the judgment operation.

[0016] In some embodiments, the flash memory operation instruction is one of a simple instruction type, an instruction plus waiting type, and a loop waiting type.

[0017] In some embodiments, each of the plurality of physical threads is equipped with a set of mutually independent general registers so that each thread can store and call data.

[0018] In some embodiments, the instruction format output by the parsing unit includes one or more of the following items: instruction type, coprocessing unit pointer, operation type, timer waiting time, completion condition, priority of the instruction, and thread locking condition.

[0019] In some embodiments, when the condition returned by the coprocessing unit does not meet the requirements, the parsing unit directly re-executes the current instruction.

[0020] According to a second aspect of the embodiments of the present disclosure, there is provided a controller of a storage device, including the processing unit described in any one of the above.

[0021] According to a third aspect of the embodiments of the present disclosure, there is provided a storage device, including: a controller and a plurality of flash memory cells, the controller is coupled to the plurality of flash memory cells through a first channel, and the controller is the controller described above.

[0022] In summary, the processing unit provided by the embodiments of the present disclosure adopts a processor design solution with multiple hardware threads, designs multiple physical threads in each channel to achieve efficient control of multiple flash memory particles in each channel, and for the input / output operations of protocol conversion, a dedicated coprocessor unit is designed, and multiple physical threads use the coprocessor unit in a time-sharing and sharing manner to complete. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, the above and other objects, features and advantages of the present invention will become clearer. In the drawings:

[0024] Figure 1 is the topological structure of the controller and flash memory particles in the storage device;

[0025] Figure 2 shows the process executed by the controller in the storage device for the read command received from the host;

[0026] Figure 3A shows the IO operation of directly sending a simple instruction;

[0027] Figure 3B shows the IO operation of directly sending an instruction plus waiting for the tR time;

[0028] Figure 3C shows the IO operation of directly sending an instruction plus loop waiting for the tR time;

[0029] Figure 4 shows the schematic structural diagram of the processing unit proposed by the embodiments of the present disclosure;

[0030] Figure 5 shows Figure 4 the flowchart of the branch judgment when the parsing unit in

[0031] Figure 6 shows Figure 4 the schematic diagram of the instruction format output by the parsing unit in DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] The present invention will be described in more detail below with reference to the accompanying drawings. In each of the drawings, the same elements are denoted by similar reference numerals. For clarity, the various parts in the drawings are not drawn to scale. In addition, some well-known parts may not be shown.

[0033] The present invention will be described based on embodiments, but the present invention is not limited to these embodiments. In the following detailed description of the present invention, some specific details are described in detail. Those skilled in the art can fully understand the present invention without the description of these details. In order to avoid obscuring the essence of the present invention, well-known methods, processes, procedures, components, and circuits are not described in detail.

[0034] Unless the context clearly requires otherwise, the words "including", "comprising", and similar words throughout the specification and claims should be construed in an inclusive sense rather than an exclusive or exhaustive sense; that is, the meaning of "including but not limited to". In the description of the present invention, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. In addition, in the description of the present invention, unless otherwise specified, the meaning of "a plurality of" is two or more.

[0035] In the control of flash memory storage media, most of its operations mainly focus on input-output (IO) operations for protocol conversion. Such operations have relatively low requirements for computing performance and are not compute-intensive tasks. Based on this characteristic, the core idea of the embodiments of the present disclosure is: adopting a processor design scheme with multiple hardware threads, designing multiple physical threads in each channel to achieve efficient control of multiple flash memory particles under each channel. At the same time, for input-output operations of protocol conversion, a dedicated coprocessor unit is designed, and multiple physical threads use the coprocessor unit in a time-sharing and sharing manner. This design method can fully exploit the potential of hardware computing resources, effectively reduce the chip area while improving the control efficiency, realize the optimal integration and efficient utilization of resources, and bring significant advantages and improvements to the controller in terms of performance, cost, and other aspects.

[0036] For example. Figure 2 The flow executed by the controller for a read command received from the host is given. As Figure 2 shown, it includes the following steps.

[0037] Step S210, send a read instruction to the flash memory particle to start the flash memory particle to move the data in the storage array to the page cache. The read instruction contains the physical address of the data to be read in the flash memory particle.

[0038] Step S220, the timer waits for time tR. During this period, the flash memory particle reads the data; during the waiting stage, the controller can send read instructions to other flash memory particles in the same channel or complete other computing operations.

[0039] Step S230, after the waiting time is completed, send a read status instruction to the flash memory particle.

[0040] Step S240: Determine whether the read instruction has been completed based on the feedback of the read status instruction (i.e., whether the flash memory cell has completed the operation of moving the data in the storage array to the page cache). If not, execute Step S250; if so, execute Step S260.

[0041] Step S250: The timer waits for time tR, and then Step S230 is executed again. The controller can perform other operations during the waiting phase.

[0042] Step S260: Perform a data transfer operation, that is, read the data in the page cache into the controller.

[0043] It should be noted that there are two types of interface protocols between the current controller and the flash memory cell: one is the traditional flash interface protocol (e.g., ONFi or Toggle), in which commands / addresses and data are transmitted using the same bus; the other is the SCA (Separate Command Address, abbreviated as SCA) protocol, in which command / address and data transmissions are separated, that is, each channel includes a DQ bus and a CA (Command and Address) bus. Corresponding to Figure 2 in this case, if the first type of protocol is adopted, Steps S210, S230, and S260 all involve IO operations on the same bus; if the second type of protocol is adopted, Steps S210 and S230 use IO operations on the CA bus, and Step S260 involves IO operations on the DQ bus.

[0044] The process executed by the controller for the write command received from the host is also Figure 2 similar. It can be seen that the process executed for the read / write command received from the host can be regarded as consisting of a series of IO atomic operations, conditional judgments, and waiting times for the flash memory cell. Therefore, the process executed for the read / write command received from the host can be re-decomposed and combined into three types of instructions such as Figure 3A 、 Figure 3B and Figure 3C shown: Figure 3A What is shown is the IO atomic operation of directly sending a simple instruction; Figure 3B The command plus waiting type shown is the IO atomic operation of directly sending an instruction plus waiting for time tR; Figure 3CThe cyclic waiting type shown is an IO atomic operation that directly sends instructions plus a cyclic wait for time tR, and also includes, after sending the instructions, making a judgment based on the instruction feedback result to determine whether to continue waiting. These three combinations are all classified as flash memory operation instructions. At the same time, 3 types of flash memory operation instructions are defined according to these three combinations. Moreover, since these flash memory operation instructions require waiting operations, a timer is set inside the processor to handle the delay. In addition to flash memory operation instructions, there are also general instructions, which are used to complete operations such as basic arithmetic, general register access, and memory access that do not involve the IO operations of the flash memory storage medium.

[0045] Correspondingly, the embodiment of the present disclosure proposes to design a processing unit in each channel of the controller of the storage device, and this processing unit processes host commands in the manner of an instruction pipeline. Figure 4 The structural schematic diagram of the processing unit proposed by the embodiment of the present disclosure is given. As shown in the reference figure, the processing unit 400 is constructed as a processor architecture with a 5-stage pipeline, and the processing unit is equipped with N physical threads. The processing unit 400 includes an instruction fetch unit 401, an instruction parsing unit 402, an execution unit 403, a write-back unit 404, N timers 405, and N groups of general register files 408. The specific number of N can be flexibly configured according to parameters.

[0046] After receiving the start signal sent by the upper-layer module, the instruction fetch unit 401 starts the process of reading the host command from the instruction space and transmits the read host command to the instruction parsing unit 402.

[0047] The instruction parsing unit 402 is responsible for parsing and processing the host command, which is used to decompose the host command into an instruction sequence composed of flash memory operation instructions and / or general instructions, and will arbitrate the instruction priorities of each physical thread according to the resource status required for each instruction execution. According to the arbitration result, each flash memory operation instruction and / or general instruction is allocated to the corresponding physical thread, and some instructions are specially processed.

[0048] The execution unit 403 executes in each physical thread and is responsible for executing each flash memory operation instruction and general instruction. Moreover, the execution unit 403 forwards the atomic operation of the input and output (IO) of the flash memory particles in the flash memory operation instruction to the coprocessing unit 406 for execution. The coprocessing unit 406 is dedicated to completing the atomic operation of the input and output (IO) of the flash memory particles. During the operation of multiple physical threads, through the use of a priority arbitration mechanism and thread locks, the competition for the control right of each thread to the coprocessing unit 406 is realized.

[0049] The function of the write-back unit 404 is to write the calculated result back to the general-purpose register 408 or the data buffer. At the same time, it performs calculation processing on data such as the flash status value (NAND Status) or the set value (NAND SetFeature) returned by the coprocessor unit 406, and writes it back to the internal general-purpose register 408 through the bus 407.

[0050] Preferably, each physical thread is equipped with an independent timer 405. After the instruction is executed, if it is determined to be a loop-wait type instruction and the judgment condition is not met, the timer 405 will promptly notify the parsing unit 402 to resend the previous instruction. In this way, there is no need to perform the fetch operation again, effectively avoiding the pipeline flush phenomenon caused by code jumps due to the flash status not meeting the requirements.

[0051] Preferably, each physical thread is equipped with a set of independent general-purpose registers 408 to facilitate the independent storage and invocation of data by each thread. At the same time, each physical thread divides an independent data cache access space in the data buffer area to ensure the independence and efficiency of data access for each thread.

[0052] In some embodiments, the branch judgment process when the parsing unit 402 issues an instruction is as Figure 5 shown.

[0053] In step S501, prepare to parse the next instruction.

[0054] In step S502, determine whether the current instruction is a flip type. If so, execute step S504; otherwise, execute step S503. The flip type refers to an instruction type that includes loop waiting. In this step, taking Figure 4 as an example, the coprocessor unit 406 returns data such as the flash status value (NAND Status) or the set value (NAND Set Feature). The write-back unit 404 writes this data into the corresponding register, and the parsing unit 402 reads this data from the corresponding register and makes a judgment based on this data.

[0055] In step S503, parse and issue the next instruction.

[0056] In step S504, wait for the instruction to execute and return the status

[0057] In step S505, it is determined whether the requirements are met. If so, step S503 is executed; otherwise, step S506 is executed. Generally, it refers to whether the flash memory status is ready, such as whether the erase / program is completed, whether the data to be read is ready, whether data can be read, whether the previous operation is completed, whether a command can be sent, and so on.

[0058] In step S506, the current instruction is retransmitted.

[0059] The embodiments of the present disclosure also provide a Figure 6 special instruction format as shown. After the host command enters the parsing unit 402, the parsing unit 402 decomposes the host command into multiple instructions, and each instruction consists of the following entries as shown Figure 6 shown: Instruction type: including general instructions for basic operations, general register access, memory access, etc., and flash memory operation instructions. The flash memory operation instructions include atomic operations of input and output (IO) involving protocol conversion completed by the coprocessor unit, such as operations of transmitting read instructions, reading status instructions, and data transmission; Coprocessor unit pointer: the function pointer corresponding to the flash memory atomic operation executed by the coprocessor unit; Operation type: the execution type of the atomic operation, such as simple instruction type, instruction plus wait type, loop wait type, etc.; Timer waiting time (timer): if the operation type is command plus wait type or loop wait type, this field represents the waiting time for command execution; Ready condition: if the operation type is loop wait type, this field represents the judgment condition for the return status of the coprocessor unit. If the return status is equal to the ready condition, the condition is met and the loop is exited. For example, the judgment condition of the return status refers to the flash memory status, and the ready condition is that the flash memory status is ready; Priority: the priority of this instruction. When multiple threads simultaneously meet the execution conditions, a priority arbitration mechanism is used for screening to determine a thread with the highest priority to execute the command. If the priorities of these threads are the same, then a polling method is used to perform the arbitration operation to ensure the orderliness and fairness of instruction execution; Lock: A thread locking condition. After this command is executed, the corresponding flash memory resources will be locked as required. During this period, commands related to the flash memory resources of other threads cannot be executed. Only after the flash memory resources are released can these commands be executed. For example, when a thread executes two consecutive atomic operations A and B, if both operations need to be performed on the same physical address, and the continuity of the operations between the two atomic operations A and B must be ensured, and it is not allowed to switch to other physical addresses under the same chip enable signal during the execution of these two atomic operations. In this case, for the instruction of atomic operation A, its lock needs to be set to lock an enable signal. In this way, this enable signal can only be controlled by the currently executing thread, and other threads that need to operate on this resource will be locked. When the instruction of atomic operation B sets the lock to release the resources, after the B instruction is executed, other locked threads will be unlocked, thereby restoring the operation permission of this enable resource.

[0060] In summary, the processing unit provided by the embodiments of the present disclosure has the following advantages:

[0061] First, integrate the timer module inside the processor. In this way, the waiting process for operations related to the flash memory can be directly processed, reducing the access frequency to the peripherals, and thus optimizing and improving the performance of the processor;

[0062] Second, the parsing unit performs judgment and processing operations on the type of circular waiting. When the condition returned by the coprocessor unit does not meet the requirements, the current instruction is directly re-executed, effectively avoiding the performance loss caused by branch judgment and pipeline flushing;

[0063] Third, build a two-level processor execution system. The upper-level processor focuses on the parallel execution management except for the atomic operations of the input and output of the flash memory particles, while the coprocessor unit is responsible for specifically implementing the atomic operations of the input and output of the flash memory particles;

[0064] Fourth, by designing a dedicated instruction format, integrate the thread locking function, priority, and the judgment condition for exiting the loop instruction into the instruction. Thus, threads with conflicting flash memory resources can be directly locked, priority arbitration operations can be directly carried out on multiple threads at the instruction level, and flash memory operations and status judgments can be completed in a single instruction. If the condition is not met, only this instruction needs to be re-executed, without experiencing the resource consumption caused by a series of operations such as jumps, pipeline flushing, and re-fetching instructions.

[0065] Correspondingly, the embodiments of the present disclosure also provide a storage device. Refer to Figure 1As shown, the storage device includes a controller and a flash memory storage medium composed of flash memory particles. The controller includes a plurality of channels, each channel is respectively coupled to at least one flash memory particle, and a processing unit 400 as shown in Figure 4 is provided in each channel.

[0066] As described above in accordance with the embodiments of the present invention, these embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the above description. These embodiments are selected and specifically described in this specification in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can make good use of the present invention and its modifications based on the present invention. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. A processing unit, applied to a controller of a storage device, wherein the controller is coupled to a plurality of flash memory particles through a first channel, the processing unit is equipped with a plurality of physical threads, and the processing unit comprises: An instruction fetch unit, used for reading a host command from the instruction space and transmitting the read host command to the parsing unit; A parsing unit, configured to generate an instruction sequence based on an operation performed by executing the host command, arbitrate instruction priorities of the plurality of physical threads according to resource conditions required for executing each instruction, and assign each instruction to a corresponding physical thread based on an arbitration result, wherein each instruction is a general instruction or a flash memory operation instruction, and the flash memory operation instruction involves an atomic operation on an input and output of the flash memory particle; An execution unit, which is executed by the multiple physical threads respectively and is used to execute the flash memory operation instruction and the general instruction; The co-processing unit is dedicated to processing atomic operations for input and output of the flash memory particles, wherein the multiple physical threads share the co-processing unit in a time-sharing manner.

2. The processing unit according to claim 1, wherein: The co-processing unit uses a thread lock to lock the corresponding flash memory resource when processing the atomic operation of the flash memory operation instruction for the input and output of the flash memory particle.

3. The processing unit according to claim 1, wherein: When decomposing the host command, the parsing unit incorporates the waiting operation for the execution of the atomic operation of the input and output of the flash memory particle and the judgment operation on the execution result of the atomic operation of the input and output of the flash memory particle into the flash memory operation instruction.

4. The processing unit according to claim 3, wherein: The processing unit further includes timers corresponding to the plurality of physical threads respectively, and the timers are used to time the waiting for the execution of the atomic operation of the input and output of the flash memory particles.

5. The processing unit according to claim 3, further comprising a write-back unit, wherein the write-back unit writes the data returned by the co-processing unit into a general register so that the parsing unit obtains the data for the determination operation.

6. The processing unit according to claim 3, wherein: The flash memory operation instruction is one of a simple instruction type, an instruction plus wait type, and a loop wait type.

7. The processing unit according to claim 5, wherein: Each of the plurality of physical threads is equipped with a set of mutually independent general registers so that each thread can store and call data.

8. The processing unit according to claim 1, wherein: The instruction format output by the parsing unit includes one or more of the following items: instruction type, co-processing unit pointer, operation type, timer waiting time, completion condition, instruction priority and thread locking condition.

9. The processing unit according to claim 1, wherein: When the condition returned by the co-processing unit does not meet the requirement, the parsing unit directly re-executes the current instruction.

10. A controller of a storage device, comprising the processing unit according to any one of claims 1 to 9.

11. A storage device comprising: A controller and a plurality of flash memory particles, wherein the controller is coupled to the plurality of flash memory particles via a first channel, and the controller is the controller according to claim 10.