Device and method for improving read-write performance of solid-state drive

By introducing multiple virtual direct memory access units and input and output processors into the SSD controller, the read and write process is dynamically adjusted, and the problem of poor read and write performance of traditional SSDs is solved, achieving more efficient IO processing and lower latency.

CN120295546APending Publication Date: 2025-07-11T-HEAD (SHANGHAI) SEMICON CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410044433.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

During read and write operations, traditional solid-state drivers (SSDs) have problems such as low performance and high delay in read and write commands, especially when mixing size IO and high priority IO, it is difficult for the prior art to effectively optimize the average delay of IO writes and processing delay of high priority IO.

Method used

By introducing multiple virtual direct memory access units and input and output processors into the SSD controller, the virtual direct memory access unit caches the context information of IO requests, and dynamically adjusts the read and write process according to the priority strategy to achieve hardware-level IO command scheduling optimization.

Benefits of technology

It improves the read and write efficiency of SSD, reduces the read and write delay, optimizes the processing delay of high-priority IO, and achieves lower average delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295546A_ABST
    Figure CN120295546A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a device and method for improving read-write performance of a solid-state drive, and the device comprises a host driving unit, a control unit, a plurality of virtual direct memory access units, and an input / output processor which receives a state queue entry corresponding to a first input / output request in a state queue, and carries out the read-write performance of the solid-state drive on the basis of the state queue entry. Determining that context information corresponding to the first input / output request is placed in an idle virtual direct memory access unit, and after determining that the priority is greater than the priority of a currently executed second input / output request, interrupting a read-write process corresponding to the second input / output request, and executing a read-write process corresponding to the first input / output request based on the context information. According to the technical scheme, after the current execution process is interrupted, the execution process of the read-write request with the higher priority can be executed, so that the read-write efficiency is improved, and the read-write time delay is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of chip technology, and particularly to an apparatus and method for improving the read and write performance of a solid-state drive. Background Art

[0002] When performing read and write operations on a Solid-State Drive (SSD), that is, when multiple Input / Output (IO) commands are sent to the SSD for execution, there is a mixture of large and small IOs. At the same time, generally some high-priority IOs have the characteristic of small data volume. In order to optimize the average latency of IO writes and the processing latency of high-priority IOs, the SSD controller generally adjusts the execution order of IO commands through software to achieve QoS optimization.

[0003] However, the implementation of the traditional method requires the software to perform cumbersome control, and there are still problems with low execution efficiency. Therefore, there is an urgent need for a technical solution to improve the performance of SSD read and write commands. Summary of the Invention

[0004] Multiple aspects of this application provide an apparatus and method for improving the read and write performance of a solid-state drive, so as to improve the performance of SSD read and write commands.

[0005] In a first aspect, an embodiment of this application provides an apparatus for improving the read and write performance of a solid-state drive. The apparatus includes: a solid-state drive master controller, and a host drive unit that interacts with the solid-state drive master controller; the solid-state drive master controller includes: a control unit, a plurality of virtual direct memory access units connected to the control unit, and input / output processors respectively connected to the plurality of virtual direct memory access units;

[0006] The input / output processor is configured to receive a status queue entry corresponding to a first input / output request placed in a status queue by the host drive unit, and determine context information corresponding to the first input / output request based on the status queue entry corresponding to the first input / output request;

[0007] The input / output processor is further configured to place the context information corresponding to the first input / output request in any one of the idle virtual direct memory access units, where the context information includes the priority of the first input / output request;

[0008] The control unit is configured to interrupt the read / write process corresponding to the second input / output request in the input / output processor after determining that the priority of the first input / output request is greater than the priority of the second input / output request currently being executed in the input / output processor;

[0009] The control unit is further configured to control the input / output processor to execute the read / write process corresponding to the first input / output request based on the context information of the first input / output request.

[0010] In a possible implementation, the virtual direct memory access unit corresponding to the second input / output request is configured to record the breakpoint information corresponding to the second input / output request into the context information corresponding to the second input / output request after the read / write process corresponding to the second input / output request is interrupted;

[0011] The control unit is further configured to control the input / output processor to continue to execute the read / write process corresponding to the second input / output request based on the breakpoint information corresponding to the second input / output request after the read / write process corresponding to the first input / output request is completed.

[0012] In a possible implementation, the control unit is further configured to determine, according to the context information in the virtual direct memory access unit corresponding to the second input / output request, that the priority of the second input / output request is not interruptible;

[0013] The control unit is further configured to, after the read / write process corresponding to the second input / output request is completed, according to the context information in each virtual direct memory access unit, control the input / output processor to execute the read / write process corresponding to the first input / output request based on the context information of the first input / output request after determining that the priority of the first input / output request is the highest.

[0014] In a possible implementation, the solid-state drive master controller further includes: a physical direct memory access unit connected to the host drive unit, and the physical direct memory access unit is connected to the control unit;

[0015] The input / output processor is further configured to read data from the memory corresponding to the host drive unit through the physical direct memory access unit based on the context information of the first input / output request, and write the data into the physical space.

[0016] In a possible implementation, the input / output processor is further configured to generate a completion queue entry corresponding to the input / output request after the read / write process corresponding to any input / output request is completed;

[0017] The host drive unit is further configured to release the input / output resources corresponding to the input / output request after receiving the completion queue entry corresponding to the input / output request.

[0018] In a possible implementation, the virtual direct memory access unit corresponding to the first input / output request is further configured to release the context information corresponding to the first input / output request after the read / write process corresponding to the first input / output request is completed.

[0019] In a second aspect, an embodiment of the present application provides a method for improving the read / write performance of a solid-state drive, which is applied to a solid-state drive controller in the device involved in the first aspect. The method includes:

[0020] Receiving a status queue entry corresponding to a first input / output request placed in a status queue by a host drive unit, and determining context information corresponding to the first input / output request based on the status queue entry corresponding to the first input / output request;

[0021] Placing the context information corresponding to the first input / output request in any idle virtual direct memory access unit, where the context information includes the priority of the first input / output request;

[0022] After determining that the priority of the first input / output request is greater than the priority of a second input / output request currently being executed in the input / output processor, interrupting the read / write process corresponding to the second input / output request in the input / output processor;

[0023] Controlling the input / output processor to execute the read / write process corresponding to the first input / output request based on the context information of the first input / output request.

[0024] In a possible implementation, the method further includes:

[0025] After the read / write process corresponding to the second input / output request is interrupted, recording breakpoint information corresponding to the second input / output request into the context information corresponding to the second input / output request;

[0026] After the read / write process corresponding to the first input / output request is completed, controlling the input / output processor to continue executing the read / write process corresponding to the second input / output request based on the breakpoint information corresponding to the second input / output request.

[0027] In a third aspect, an embodiment of the present application provides a chip, including a digital integrated circuit, where the digital integrated circuit is configured to execute the method involved in the second aspect.

[0028] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed, they are used to implement the method involved in the second aspect.

[0029] The device and method for improving the read and write performance of a solid-state drive provided by an embodiment of the present application. The device includes: a solid-state drive master controller, and a host drive unit that interacts with the solid-state drive master controller. The solid-state drive master controller includes: a control unit, a plurality of virtual direct memory access units connected to the control unit, and input / output processors respectively connected to the plurality of virtual direct memory access units. The input / output processor is configured to receive a status queue entry corresponding to a first input / output request placed in a status queue by the host drive unit, and determine context information corresponding to the first input / output request based on the status queue entry corresponding to the first input / output request. The input / output processor is further configured to place the context information corresponding to the first input / output request in any one of the idle virtual direct memory access units, where the context information includes the priority of the first input / output request. The control unit is configured to interrupt the read / write process corresponding to the second input / output request in the input / output processor after determining that the priority of the first input / output request is greater than the priority of the second input / output request currently being executed in the input / output processor. The control unit is further configured to control the input / output processor to execute the read / write process corresponding to the first input / output request based on the context information of the first input / output request. In this technical solution, by setting up virtual direct memory access units, it is possible to switch the operations of the input / output processor to the virtual direct memory access unit corresponding to the read / write request with a higher priority when executing the read / write process, and execute the execution process of the read / write request with a higher priority, so as to improve the read / write efficiency and reduce the read / write latency. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0031] Figure 1 is a schematic diagram of a hardware structure in the prior art;

[0032] Figure 2 is a schematic diagram of the solution steps provided by the traditional solution - I;

[0033] Figure 3 is a schematic diagram of the solution steps provided by the traditional solution - II;

[0034] Figure 4 is a schematic diagram of the solution steps provided by the traditional solution - III;

[0035] Figure 5 is a schematic structural diagram of a device for improving the read and write performance of a solid-state drive provided by an embodiment of the present application Figure 1 ;

[0036] Figure 6Structural schematic of a device for improving the read and write performance of a solid-state drive provided by an embodiment of the present application Figure 2 ;

[0037] Figure 7 Flow chart of a method for improving the read and write performance of a solid-state drive provided by an embodiment of the present application;

[0038] Figure 8 Schematic diagram of solution steps provided by an embodiment of the present application Detailed implementation manners

[0039] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part rather than all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application. The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions, and corresponding operation entrances are provided for users to select authorization or rejection.

[0040] First, the terms involved in the embodiments of the present application are explained:

[0041] Solid-state drive (SSD): A large-capacity data storage device composed of NAND Flash chips and SSD Controller chips;

[0042] SSD Controller: Also known as the main control chip or main controller, it is one of the key components of an SSD, a dedicated chip with built-in firmware for managing NAND Flash;

[0043] Nand Flash: A non-volatile data storage device, commonly used as a storage device in SSDs and memory cards;

[0044] PCIe: PCI Express, a high-speed serial data transmission protocol commonly used in computer systems. Most interfaces between SSDs and central processing units (CPUs) use this protocol;

[0045] NVMe: Non-Volatile Memory Express, an interface protocol carried on PCIe, used to define the interaction interface between host software and NVM devices;

[0046] Submission Queue (SQ): The space used to cache the commands issued by the host software (English: Commands). Among them, the Submission Queue Entry (SQE) is the Commands issued by the host software;

[0047] Completion Queue (CQ): The space used to cache the Completion events reported by the SSD controller. Among them, the Completion Queue Entry (CQE) is the Completion event reported by the SSD controller;

[0048] Doorbell: The host notifies the SSD controller that the Producer Index (PI) or the Consumer Index (CI) has been updated by writing the Doorbell register;

[0049] PI: The producer writes data according to PI for the consumer to use;

[0050] CI: The consumer reads data according to CI and performs corresponding operations;

[0051] Scatter Gather List (SGL): Used to identify the memory addresses of data.

[0052] Next, the scenarios involved in the embodiments of the present application will be introduced:

[0053] In practical applications, when multiple Input / Output (IO) commands are sent to the SSD for execution, there is a mixture of large and small IOs. At the same time, generally, some high-priority IOs have the characteristic of small data volume. For example, when the system generates some Log flushes, the IOs are generally 4KB or 8KB in size. In order to optimize the average latency of IO writes and the processing latency of high-priority IOs, the SSD Controller generally adjusts the execution order of IO commands through software to achieve the optimization of Quality of Service (QoS).

[0054] In the traditional solution, there is only one Direct Memory Access (DMA) channel at the bottom layer, and software needs to execute cumbersome control.

[0055] First, the specific execution process of the IO Write Command in the traditional solution is described, and the average latency of command processing and the absolute latency of single command processing under different solutions are described. The execution of a single IO Write command in the NVMe protocol can be simplified into the following steps:

[0056] In the first step, the Host Driver fills the Submission Queue Entry (SQE) of the IO Write into the Submission Queue (SQ) space and rings the Doorbell to the SSD controller to indicate that there is a new command to be executed;

[0057] In the second step, the SSD controller sends a PCIe Read operation to read the SQE;

[0058] In the third step, after the SSD controller obtains the SQE, it initiates an operation to obtain the Scatter Gather List (SGL) (in this article, SGL is used to generally refer to the address space), and sends a PCIe Direct Memory Access (DMA) Read operation to obtain data from the Host space according to the SGL information;

[0059] In the fourth step, after the SSD controller obtains the IO Data, it performs the IO operation; after completion, it assembles the Completion Queue Entry (CQE) and sends it to the Completion Queue (CQ) space;

[0060] In the fifth step, the Host Driver obtains the CQE, releases the resources related to the IO command, and sends a Doorbell to the SSD controller.

[0061] Correspondingly, Figure 1 is a schematic diagram of the hardware structure in the prior art, as Figure 1 shown, the hardware structure includes: a Host and an SSD controller. In this SSD controller, there are a hardware DMA and a CPU core.

[0062] For the above execution process, three methods in the prior art are specifically exemplified for illustration. The data volume written by IO A is greater than that written by IO B, and the data volume written by IO B is greater than that written by IO C; IO A and IO B are low-priority operations, and IO C is a high-priority operation. For the convenience of analysis, it is assumed that except for the processing time of the above third step during the execution of IO A / IO B / IO C, the processing times of other steps are the same. It is assumed that the execution time of IO A / IO B / IO C in the above first step is 3 time instants (3T); the execution time of IO A / IO B / IO C in the second step is 3 time instants; the execution time of IO A / IO B / IO C in the fourth step is 3 time instants; the execution time of IO A / IO B / IO C in the fifth step is 3 time instants; the execution time of the third step for IO A is 20 time instants, for IO B is 12 time instants, and for IO C is 4 time instants. It should be noted that the PCIe bottom layer can only complete the traffic in the same direction (read and write) serially. In the following diagrams, only the serial operations within each step are highlighted. The parallel operations between different steps refer to different operation types, and do not represent the physical traffic of SerDes.

[0063] The first type is that the SSD controller does not care about the size of the written data volume of the IO command and the priority of the command. After fetching the IO command, it executes the IO operation according to the IO order filled in by the Host Driver. The third step between multiple IOs needs to be executed serially. Figure 2 It is a schematic diagram of the solution steps provided by the traditional solution - I. As Figure 2 shown, the schematic diagram includes: time (from node t00 to node t56), Host, and SSD controller.

[0064] For example, at node t02, the Host writes the IO A command to the SQ and writes a Doorbell to mark the SQE, ending at node t05. The SSD main controller fetches the SQE corresponding to the IO A command. At node t10, the read / write process is implemented by reading the SGL and Data. After the process ends at node t30, a CQE is generated and the value CQ is written to indicate the completion of the operation for the IO A command. At node t35, the CQE is read, and then a Doorbell is used to release the CQE. At node t06, the Host writes the IO B command to the SQ and writes a Doorbell to mark the SQE, ending at node t09. The SSD main controller fetches the SQE corresponding to the IO B command. At node t30, the read / write process is implemented by reading the SGL and Data. After the process ends at node t42, a CQE is generated and the value CQ is written to indicate the completion of the operation for the IO B command. At node t47, the CQE is read, and then a Doorbell is used to release the CQE. At node t10, the Host writes the IO C command to the SQ and writes a Doorbell to mark the SQE, ending at node t13. The SSD main controller fetches the SQE corresponding to the IO C command. At node t42, the read / write process is implemented by reading the SGL and Data. After the process ends at node t46, a CQE is generated and the value CQ is written to indicate the completion of the operation for the IO C command. At node t51, the CQE is read, and then a Doorbell is used to release the CQE.

[0065] In this mode, the latency of low-priority IO A is T33, the latency of low-priority IO B is T41, the latency of high-priority IO C is T41, and the average latency reaches 38.33.

[0066] The second method: The SSD main controller cares about the write data volume size of the IO command and the command priority. After fetching the IO command, it starts to execute the operation. There is only one unit in the chip to execute the operation in the third step. Only after the third step of a single command is completed can the third step of subsequent IO commands be executed. The SSD main controller can flexibly adjust the order of subsequent unexecuted IO commands according to the write data volume size and command priority of the subsequent unexecuted commands. Figure 3 In this case, IO C will be executed before IO B. Figure 3 Schematic diagram of the solution steps provided for the traditional solution - II, as Figure 3 shown. This schematic diagram includes: time (from node t00 to node t56), Host, and SSD main controller.

[0067] For example, at the t02 node, the Host writes the IO A command to the SQ and writes a Doorbell to mark the SQE, ending at the t05 node. The SSD main controller fetches the SQE corresponding to the IO A command. At the t10 node, the read / write process is implemented by reading the SGL and Data. After the process ends at the t30 node, a CQE is generated and the value CQ is written to indicate that the operation for the IO A command is completed. At the t35 node, the CQE is read, and then the Doorbell is used to release the CQE. At the t06 node, the Host writes the IO B command to the SQ and writes a Doorbell to mark the SQE, ending at the t13 node. The SSD main controller fetches the SQE corresponding to the IO B command. At the t34 node, the read / write process is implemented by reading the SGL and Data. After the process ends at the t46 node, a CQE is generated and the value CQ is written to indicate that the operation for the IO B command is completed. At the t51 node, the CQE is read, and then the Doorbell is used to release the CQE. At the t10 node, the Host writes the IO C command to the SQ and writes a Doorbell to mark the SQE, ending at the t13 node. The SSD main controller fetches the SQE corresponding to the IO C command. At the t28 node, the read / write process is implemented by reading the SGL and Data. After the process ends at the t34 node, a CQE is generated and the value CQ is written to indicate that the operation for the IO C command is completed. At the t39 node, the CQE is read, and then the Doorbell is used to release the CQE.

[0068] In this mode, the latency of low-priority IO A is T33, the latency of low-priority IO B is T45, and the latency of high-priority IO C is T29. The average latency reaches 35.67.

[0069] The third method: The SSD main controller cares about the write data volume size of the IO command and the priority of the command. After fetching the IO command, it waits for a period of time inside the chip before starting to execute the IO operation. There is only one unit in the chip to execute the operation in the third step. After the third step of a single command is completed, the third step of subsequent IO commands can be executed. The SSD main controller can flexibly adjust the order of unexecuted IO commands according to the write data volume size and command priority of the unexecuted commands. Figure 4 Among them, IO C will be executed before IO B, and IO B will be executed before IO A. Figure 4 It is a schematic diagram of the solution steps provided for the traditional solution - III, as Figure 4 shown. This schematic diagram includes: time (from the t00 node to the t56 node), Host, and SSD main controller.

[0070] For example, at node t02, the Host writes the IO A command to the SQ and writes a Doorbell to mark the SQE, ending at node t05. The SSD controller fetches the SQE corresponding to the IO A command. At node t34, the read / write process is implemented by reading the SGL and Data. After the process at node t54 ends, a CQE is generated and the value CQ is written to indicate the completion of the operation for the IO A command. At node t59, the CQE is read, and then a Doorbell is used to release the CQE. At node t06, the Host writes the IO B command to the SQ and writes a Doorbell to mark the SQE, ending at node t13. The SSD controller fetches the SQE corresponding to the IO B command. At node t55, the read / write process is implemented by reading the SGL and Data. After the process at node t34 ends, a CQE is generated and the value CQ is written to indicate the completion of the operation for the IO B command. At node t39, the CQE is read, and then a Doorbell is used to release the CQE. At node t10, the Host writes the IO C command to the SQ and writes a Doorbell to mark the SQE, ending at node t13. The SSD controller fetches the SQE corresponding to the IO C command. At node t18, the read / write process is implemented by reading the SGL and Data. After the process at node t22 ends, a CQE is generated and the value CQ is written to indicate the completion of the operation for the IO C command. At node t27, the CQE is read, and then a Doorbell is used to release the CQE.

[0071] In this mode, the latency of low-priority IO A is T57, the latency of low-priority IO B is T33, and the latency of high-priority IO C is T17. The average latency reaches 35.67.

[0072] In any of the traditional solutions provided above, there is a problem of high latency, which has a certain negative impact on the performance of SSD read / write commands.

[0073] For the technical solution provided in this application, the technical concept of the inventor is as follows: If the hardware can provide multiple virtual data DMA channels, the scheduling priorities of different channels can be specified, the channel pause operation is supported, and the command execution of each DMA channel can be switched without blocking, so as to optimize the average latency of IO writes and the processing latency of high-priority IOs. At the same time, this solution can also use hardware to replace software to adjust the execution order of host IO commands, specify the scheduling priorities when different IOs execute data DMA, set the maximum number of pauses, the maximum pause latency, etc., reduce the CPU Core computing power requirements inside the SSD controller, and thus solve the problems existing in the prior art.

[0074] The technical solution shown in this application will be described in detail through specific embodiments. It should be noted that the following several embodiments can exist independently or be combined with each other. For the same or similar content, it will not be repeated in different embodiments.

[0075] Figure 5 The structural schematic of a device for improving the read and write performance of a solid-state drive provided by an embodiment of this application Figure 1 . As Figure 5 shown, the device includes: a solid-state drive master controller (i.e., SSD controller) 12, and a host drive unit (i.e., Host) 11 that interacts with the solid-state drive master controller; the solid-state drive master controller includes: a control unit (i.e., Scheduler) 121, a plurality of virtual direct memory access units (i.e., virtual DMA channels) connected to the control unit, and an input / output processor (i.e., IO Processor) 123 or CPU Core respectively connected to the plurality of virtual direct memory access units.

[0076] Among them, this solution takes 2 virtual direct memory access units (for example, 1221, 1222) as an example for illustration, and the principle for 3 or more is similar.

[0077] Optionally, the input / output processor 123 is configured to receive a status queue entry (Submission Queue Entry, SQE) corresponding to a first input / output request placed in the status queue by the host drive unit 11, and determine context information corresponding to the first input / output request based on the status queue entry corresponding to the first input / output request;

[0078] In this implementation, the host drive unit 11 fills the SQE of the first input / output request into the status queue (Submission Queue, SQ) space, and rings the Doorbell to the SSD controller to indicate that there is a new command to be executed. The input / output processor 123 sends a PCIe Read operation to read the SQE corresponding to the first input / output request, and generates context information corresponding to the first input / output request based on the SQE corresponding to the first input / output request. At this time, the context information is initial context information.

[0079] Among them, the input / output processor 123 determines a pointer corresponding to a scatter gather list (SGL) based on the SQE corresponding to the first input / output request, and then initiates an operation to obtain the SGL based on this pointer. The table records address information of data to be read and written, etc. And, after receiving the SQE, it also carries the priority of the first input / output request. The above but not limited to the above data is used as the context information corresponding to the first input / output request.

[0080] In a possible implementation, the priority of the input / output request can be the priority of the corresponding I / O request itself, or the priority corresponding to the data volume of the I / O request, etc.

[0081] Among them, the priority can be the level of priority, whether it can be interrupted, and the number of times it can be interrupted, etc.

[0082] The input / output processor 123 is further configured to place the context information corresponding to the first input / output request in any idle virtual direct memory access unit (for example, 1221), and the context information includes the priority of the first input / output request;

[0083] At this time, any idle virtual direct memory access unit is determined among multiple virtual direct memory access units, for example, 1221, and the context information is placed in the virtual direct memory access unit 1221.

[0084] Optionally, the control unit 121 is configured to interrupt the read / write process corresponding to the second input / output request in the input / output processor 123 after determining that the priority of the first input / output request is greater than the priority of the second input / output request currently being executed in the input / output processor;

[0085] In this implementation, the control unit 121 can read the priorities of the corresponding input / output requests in each virtual direct memory access unit. After reading that the priority of the first input / output request corresponding to the virtual direct memory access unit 1221 is greater than the priority of the second input / output request currently being executed in the input / output processor, the read / write process corresponding to the second input / output request in the input / output processor 123 is interrupted, so that the input / output processor 123 executes the read / write process corresponding to the first input / output request.

[0086] Optionally, the control unit 121 is further configured to control the input / output processor 123 to execute the read / write process corresponding to the first input / output request based on the context information of the first input / output request.

[0087] In this implementation, there is information related to the read / write process in the context information of the first input / output request, and the input / output processor 123 executes the read / write process corresponding to the first input / output request based on the context information.

[0088] It should be understood that the implementation of the input / output processor 123 executing the read / write process corresponding to the input / output request is given by Figure 6 the embodiments shown, and will not be elaborated here.

[0089] In the above embodiments, multiple virtual DMA channels are added inside the chip. Each virtual DMA channel caches the context information of the DMA operation of an IO or the context information of the DMA operation of another event, and sends a DMA request operation to the Scheduler according to a policy. The Scheduler executes the command scheduling of different virtual DMA channels according to the configured scheduling policy (i.e., the priority-related policy), which can be to suspend a certain virtual channel, or to enable a certain virtual channel for a certain period of time, or to perform interval scheduling between different virtual channels, etc.

[0090] Under the above implementation, the virtual direct memory access unit corresponding to the second input / output request is, for example, 1222.

[0091] Furthermore, the virtual direct memory access unit 1222 corresponding to the second input / output request is used to record the breakpoint information corresponding to the second input / output request into the context information corresponding to the second input / output request after the read / write process corresponding to the second input / output request is interrupted.

[0092] Under this implementation, after the read / write process corresponding to the second input / output request is interrupted, the virtual direct memory access unit 1222 corresponding to the second input / output request records the relevant information at the time of interruption and records the breakpoint information corresponding to the second input / output request into the context information corresponding to the second input / output request.

[0093] It should be understood that in the embodiments of the present application, the context information changes continuously as the read / write operation process corresponding to the entire input / output request is executed.

[0094] Optionally, the control unit 121 is further configured to, after the read / write process corresponding to the first input / output request is completed, control the input / output processor 123 to continue executing the read / write process corresponding to the second input / output request based on the breakpoint information corresponding to the second input / output request.

[0095] Under this implementation, after the read / write process corresponding to the first input / output request is completed, the control unit 121 opens the channel of the virtual direct memory access unit 1222 corresponding to the second input / output request, and the input / output processor 123 continues to execute the subsequent unfinished read / write process based on the breakpoint information corresponding to the second input / output request.

[0096] In addition, when the input / output processor 123 continues to execute the subsequent unfinished read / write process, if there is a higher-priority input / output request arriving and the second input / output request can be interrupted again, the read / write process corresponding to the current second input / output request will be interrupted at any time, so as to execute the read / write process corresponding to the input / output request with a higher priority.

[0097] In another implementation, the control unit 121 is further configured to determine that the priority of the second input / output request cannot be interrupted according to the context information in the virtual direct memory access unit 1222 corresponding to the second input / output request;

[0098] In this implementation, if the control unit 121 parses from the context information in the virtual direct memory access unit 1222 corresponding to the second input / output request that the priority of the second input / output request cannot be interrupted, the read / write process corresponding to the first input / output request cannot be executed temporarily.

[0099] Optionally, the control unit 121 is further configured to, after the read / write process corresponding to the second input / output request is completed, determine that the priority of the first input / output request is the highest according to the context information in each virtual direct memory access unit, and then control the input / output processor 123 to execute the read / write process corresponding to the first input / output request based on the context information of the first input / output request.

[0100] In this implementation, after the read / write process corresponding to the second input / output request is completed, the control unit 121 needs to determine the read / write process corresponding to the next input / output request to be executed based on the context information in each virtual direct memory access unit.

[0101] This is because during the execution of the read / write process corresponding to the second input / output request, new input / output requests may be received. Therefore, after the read / write process corresponding to the second input / output request is completed and it is determined that the priority of the first input / output request is the highest, the input / output processor 123 is controlled to execute the read / write process corresponding to the first input / output request based on the context information of the first input / output request.

[0102] The device for improving the read and write performance of a solid-state drive provided by an embodiment of the present application includes: a solid-state drive master controller, and a host drive unit that interacts with the solid-state drive master controller. The solid-state drive master controller includes: a control unit, a plurality of virtual direct memory access units connected to the control unit, and input / output processors respectively connected to the plurality of virtual direct memory access units. The input / output processor is configured to receive a status queue entry corresponding to a first input / output request placed in a status queue by the host drive unit, and determine context information corresponding to the first input / output request based on the status queue entry corresponding to the first input / output request. The input / output processor is further configured to place the context information corresponding to the first input / output request in any idle virtual direct memory access unit, where the context information includes the priority of the first input / output request. The control unit is configured to interrupt the read / write process corresponding to the second input / output request in the input / output processor after determining that the priority of the first input / output request is greater than the priority of the second input / output request currently being executed in the input / output processor. The control unit is further configured to control the input / output processor to execute the read / write process corresponding to the first input / output request based on the context information of the first input / output request. In this technical solution, by setting up virtual direct memory access units, the operations during the read / write process executed by the input / output processor can be switched to the virtual direct memory access unit corresponding to the read / write request with a higher priority, and the execution process of the read / write request with a higher priority can be executed to improve the read / write efficiency and reduce the read / write latency.

[0103] Figure 6 Structural schematic of a device for improving the read and write performance of a solid-state drive provided by an embodiment of the present application Figure 2 As Figure 6 shown, the solid-state drive master controller 12 further includes: a physical direct memory access unit (i.e., physical DMA) 124 connected to the host drive unit 11, and the physical direct memory access unit 124 is connected to the control unit 121;

[0104] Optionally, the input / output processor 123 is further configured to read data from the memory corresponding to the host drive unit 11 through the physical direct memory access unit 124 based on the context information of the first input / output request, and write the data into the physical space.

[0105] In this implementation, when the input / output processor 123 executes the execution process corresponding to the first input / output request, according to the context information in the virtual direct memory access unit 1221, for example, the data source location in the SGL, it reads data from the source location in the memory corresponding to the host drive unit 11 through the physical direct memory access unit 124, and writes the data into the physical space.

[0106] Among them, the physical space may be the storage space determined in the context information.

[0107] It should be understood that since the amount of data required for the first input / output request is different, during the corresponding reading and writing process, it may not be possible to obtain all the data at once, but may be obtained multiple times. Correspondingly, the breakpoint information can also be the information corresponding to the specific data being read or written at the time of interruption, etc.

[0108] Optionally, the input / output processor 123 is further configured to generate a completion queue entry (CQE) corresponding to the input / output request after the reading and writing process corresponding to any input / output request is completed; furthermore, the host driver unit 11 is further configured to release the input / output resources corresponding to the input / output request after receiving the completion queue entry corresponding to the input / output request.

[0109] In this implementation, after the reading and writing process corresponding to any input / output request is completed, the input / output processor 123 generates the completion queue entry corresponding to the input / output request and places the completion queue entry corresponding to the input / output request in the completion queue (CQ) space. The host driver unit 11 reads the completion queue entry in the completion queue space and releases the input / output resources corresponding to the completion queue entry.

[0110] In addition, after the reading and writing process corresponding to the input / output request is completed, the virtual direct memory access unit corresponding to the input / output request also releases the information in the channel to restore the idle state.

[0111] For example, the virtual direct memory access unit 1221 corresponding to the first input / output request is further configured to release the context information corresponding to the first input / output request after the reading and writing process corresponding to the first input / output request ends, so that the virtual direct memory access unit 1221 restores the idle state.

[0112] In the device for improving the read and write performance of the solid-state drive provided by the embodiments of the present application, the solid-state drive controller further includes: a physical direct memory access unit connected to the host driver unit, and the physical direct memory access unit is connected to the control unit; the input / output processor is further configured to read data from the memory corresponding to the host driver unit through the physical direct memory access unit based on the context information of the first input / output request and write the data into the physical space. In this technical solution, the data is read from the memory corresponding to the host driver unit through the physical direct memory access unit based on the context information obtained from the virtual direct memory access unit to implement the processing of the first input / output request.

[0113] Based on the above device embodiment, Figure 7The flowchart shows a method for improving the read and write performance of a solid - state drive provided by an embodiment of the present application. As Figure 7 shown, the method for improving the read and write performance of the solid - state drive is applied to the solid - state drive master controller. The method includes:

[0114] It should be understood that the content in this embodiment is basically given in the above - mentioned embodiment, and only a brief introduction is provided here.

[0115] Step 71: Receive the status - queue entry corresponding to the first input - output request placed in the status queue by the host drive unit, and determine the context information corresponding to the first input - output request based on the status - queue entry corresponding to the first input - output request;

[0116] In this step, the host drive unit fills the SQE of the first input - output request into the status - queue SQ space, and rings the Doorbell to the SSD controller to indicate that there is a new command to be executed. The input - output processor sends a PCIe Read operation to read the SQE corresponding to the first input - output request, and generates the context information corresponding to the first input - output request based on the SQE corresponding to the first input - output request.

[0117] Step 72: Place the context information corresponding to the first input - output request in any idle virtual direct memory access unit. The context information includes the priority of the first input - output request;

[0118] In this step, determine any idle virtual direct memory access unit among multiple virtual direct memory access units, and place the context information corresponding to the first input - output request in the determined idle virtual direct memory access unit.

[0119] Step 73: After determining that the priority of the first input - output request is greater than the priority of the second input - output request currently being executed in the input - output processor, interrupt the read - write process corresponding to the second input - output request in the input - output processor;

[0120] In this step, the priority of the corresponding input - output request in each virtual direct memory access unit can be read. When it is read that the priority of the first input - output request corresponding to the virtual direct memory access unit is greater than the priority of the second input - output request currently being executed in the input - output processor, the read - write process corresponding to the second input - output request in the input - output processor is interrupted.

[0121] Step 74: Control the input - output processor to execute the read - write process corresponding to the first input - output request based on the context information of the first input - output request.

[0122] In this step, there is information related to the read / write process in the context information of the first input / output request. Based on the context information, the input / output processor executes the read / write process corresponding to the first input / output request.

[0123] As an implementation, when the input / output processor executes the execution process corresponding to the first input / output request, according to the context information in the virtual direct memory access unit, for example, the data source location in the SGL, it reads data from the source location in the memory corresponding to the host driver unit through the physical direct memory access unit and writes the data into the physical space.

[0124] Furthermore, after the read / write process corresponding to the second input / output request is interrupted, the breakpoint information corresponding to the second input / output request is recorded into the context information corresponding to the second input / output request; after the read / write process corresponding to the first input / output request is completed, the input / output processor is controlled to continue executing the read / write process corresponding to the second input / output request based on the breakpoint information corresponding to the second input / output request.

[0125] In a possible implementation, when the read / write process corresponding to the second input / output request is interrupted, the virtual direct memory access unit corresponding to the second input / output request records the relevant information at the time of interruption and records the breakpoint information corresponding to the second input / output request into the context information corresponding to the second input / output request.

[0126] After the read / write process corresponding to the first input / output request is completed, the control unit opens the channel of the virtual direct memory access unit corresponding to the second input / output request, and the input / output processor continues to execute the subsequent unfinished read / write process based on the breakpoint information corresponding to the second input / output request.

[0127] It should be understood that the method for improving the read / write performance of the solid-state drive is applied to the solid-state drive controller, that is, other implementations involved in the above device are also included in the method embodiment. Since the content is relatively repetitive, it will not be elaborated here.

[0128] The method for improving the read and write performance of a solid-state drive provided by an embodiment of the present application is applied to the main controller of the solid-state drive. The method includes: receiving a status queue entry corresponding to a first input / output request placed in a status queue by a host drive unit, and determining context information corresponding to the first input / output request based on the status queue entry corresponding to the first input / output request. Then, placing the context information corresponding to the first input / output request in any idle virtual direct memory access unit. The context information includes the priority of the first input / output request. Finally, after determining that the priority of the first input / output request is greater than the priority of a second input / output request currently being executed in an input / output processor, interrupting the read and write process corresponding to the second input / output request in the input / output processor, and then controlling the input / output processor to execute the read and write process corresponding to the first input / output request based on the context information of the first input / output request. In this technical solution, by setting up a virtual direct memory access unit, it is possible to switch the operation of the input / output processor to the virtual direct memory access unit corresponding to a read and write request with a higher priority when executing the read and write process, and execute the execution process of the read and write request with a higher priority to improve the read and write efficiency and reduce the read and write latency.

[0129] Based on the above Figures 2 - 4 request, Figure 8 it is a schematic diagram of the solution steps provided by an embodiment of the present application. As Figure 8 shown, the schematic diagram includes: time (from node t00 to node t56), Host, and SSD main controller.

[0130] For example, at the t02 node, the Host writes the IO A command to the SQ and writes a Doorbell to mark the SQE, and it ends at the t05 node. The SSD controller fetches the SQE corresponding to the IO A command. At the t10 node, the read / write process is implemented by reading the SGL and Data. After the process ends at the t30 node, a CQE is generated and the value CQ is written to indicate the completion of the operation for the IO A command. At the t35 node, the CQE is read, and then the Doorbell is used to release the CQE. At the t06 node, the Host writes the IO B command to the SQ and writes a Doorbell to mark the SQE, and it ends at the t09 node. The SSD controller fetches the SQE corresponding to the IO B command. At the t30 node, the read / write process is implemented by reading the SGL and Data. After the process ends at the t42 node, a CQE is generated and the value CQ is written to indicate the completion of the operation for the IO B command. At the t47 node, the CQE is read, and then the Doorbell is used to release the CQE. At the t10 node, the Host writes the IO C command to the SQ and writes a Doorbell to mark the SQE, and it ends at the t13 node. The SSD controller fetches the SQE corresponding to the IO C command. At the t42 node, the read / write process is implemented by reading the SGL and Data. After the process ends at the t46 node, a CQE is generated and the value CQ is written to indicate the completion of the operation for the IO C command. At the t51 node, the CQE is read, and then the Doorbell is used to release the CQE.

[0131] In the method of this solution, the latency of the low-priority IO A is T49, the latency of the low-priority IO B is T29, and the latency of the high-priority IO C is T17, and the average latency reaches 31.67.

[0132] The traditional solution cannot obtain a better average processing latency and the minimum latency of high-priority commands. However, this solution provides multiple virtual DMA channels that can be automatically switched, and it can achieve a reduction in the average latency based on the PCIe TLP packet granularity.

[0133] For example, at node t02, the Host writes the IO A command to the SQ and writes a Doorbell to mark the SQE, ending at node t05; at node t06, the Host writes the IO B command to the SQ and writes a Doorbell to mark the SQE, ending at node t09; at node t10, the Host writes the IO C command to the SQ and writes a Doorbell to mark the SQE, ending at node t13; the SSD controller fetches the SQE corresponding to the IO A command. At node t10, the read-write process is implemented by reading the SGL and Data. However, since it is detected that the priority of IO B is higher than that of IO A, the read-write process of IO A is interrupted at node t14. Thus, the SGL and Data of IO B are read to implement the read-write process. At node t18, since it is detected that the priority of IO C is higher than that of IO B, the read-write process of IO B is interrupted at node t14. Thus, the SGL and Data of IO C are read to implement the read-write process; at node t22, after the read-write process of IO C is completed, a CQE is generated and written to the value CQ to indicate the completion of the operation for the IO C command. At node t27, the CQE is read, and then the Doorbell is used to release the CQE; at node t22, according to the breakpoint information in the context information corresponding to IO B, the corresponding read-write process is continued. At node t30, after the read-write process of IO B is completed, a CQE is generated and written to the value CQ to indicate the completion of the operation for the IO B command. At node t35, the CQE is read, and then the Doorbell is used to release the CQE; at node t30, according to the breakpoint information in the context information corresponding to IO A, the corresponding read-write process is continued. At node t46, after the read-write process of IO A is completed, a CQE is generated and written to the value CQ to indicate the completion of the operation for the IO A command. At node t51, the CQE is read, and then the Doorbell is used to release the CQE.

[0134] In the traditional solution, software can execute by splitting the IO command into multiple third steps, but this will increase the complexity of software processing and requires the chip to provide stronger CPU Core computing power. Since the splitting granularity cannot be too small, too small will increase the interaction overhead between software and hardware. For example, for a 512KB IO split at a 4KB granularity, at least 129 interactions are required; even when split at a 4KB granularity, high-priority IOs will have a blocking of nearly us and cannot achieve the effect of direct hardware switching in this solution, which can achieve 0-delay blocking.

[0135] The technical implementation and effects of this embodiment are similar to those of the above embodiment and will not be elaborated here.

[0136] The embodiment of the present application also provides a chip, including a digital integrated circuit, and the digital integrated circuit is used to implement the method described in any one of the foregoing embodiments.

[0137] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0138] The integrated modules implemented in the form of software function modules can be stored in a computer-readable storage medium. The above-mentioned software function modules are stored in a storage medium and include several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in various embodiments of the present application.

[0139] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0140] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.

[0141] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in various embodiments of the present application.

[0142] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A device for improving the read and write performance of a solid-state drive, characterized in that, The device includes: a solid-state drive master controller, and a host drive unit that interacts with the solid-state drive master controller; the solid-state drive master controller includes: a control unit, a plurality of virtual direct memory access units connected to the control unit, and input / output processors respectively connected to the plurality of virtual direct memory access units; The input / output processor is configured to receive a status queue entry corresponding to a first input / output request placed in a status queue by the host drive unit, and determine context information corresponding to the first input / output request based on the status queue entry corresponding to the first input / output request; The input / output processor is further configured to place the context information corresponding to the first input / output request in any one of the idle virtual direct memory access units, where the context information includes the priority of the first input / output request; The control unit is configured to interrupt the read / write process corresponding to the second input / output request in the input / output processor after determining that the priority of the first input / output request is greater than the priority of the second input / output request currently being executed in the input / output processor; The control unit is further configured to control the input / output processor to execute the read / write process corresponding to the first input / output request based on the context information of the first input / output request.

2. The device according to claim 1, wherein The virtual direct memory access unit corresponding to the second input / output request is configured to record breakpoint information corresponding to the second input / output request into the context information corresponding to the second input / output request after the read / write process corresponding to the second input / output request is interrupted; The control unit is further configured to control the input / output processor to continue executing the read / write process corresponding to the second input / output request based on the breakpoint information corresponding to the second input / output request after the read / write process corresponding to the first input / output request is completed.

3. The device according to claim 1, characterized in that, The control unit is further configured to determine that the priority of the second input / output request is not interruptible according to the context information in the virtual direct memory access unit corresponding to the second input / output request; The control unit is further configured to, after the read / write process corresponding to the second input / output request is completed, based on the context information in each virtual direct memory access unit, control the input / output processor to execute the read / write process corresponding to the first input / output request based on the context information of the first input / output request after determining that the priority of the first input / output request is the highest.

4. The device according to any one of claims 1 to 3, characterized in that, The solid-state drive master controller further includes: a physical direct memory access unit connected to the host drive unit, and the physical direct memory access unit is connected to the control unit; The input / output processor is further configured to read data from the memory corresponding to the host drive unit through the physical direct memory access unit based on the context information of the first input / output request, and write the data into the physical space.

5. The device according to any one of claims 1 to 3, characterized in that, The input / output processor is further configured to generate a completion queue entry corresponding to the input / output request after the read / write process corresponding to any input / output request is completed; The host drive unit is further configured to release the input / output resources corresponding to the input / output request after receiving the completion queue entry corresponding to the input / output request.

6. The device according to any one of claims 1 to 3, characterized in that, The virtual direct memory access unit corresponding to the first input / output request is further configured to release the context information corresponding to the first input / output request after the read / write process corresponding to the first input / output request is completed.

7. A method for improving the read and write performance of a solid-state drive, characterized in that, A solid-state drive controller applied to the device according to any one of claims 1-6, the method comprising: Receiving a status queue entry corresponding to a first input / output request placed in a status queue by a host drive unit, and determining context information corresponding to the first input / output request based on the status queue entry corresponding to the first input / output request; Placing the context information corresponding to the first input / output request in any idle virtual direct memory access unit, where the context information includes the priority of the first input / output request; After determining that the priority of the first input / output request is greater than the priority of a second input / output request currently being executed in the input / output processor, interrupting the read / write process corresponding to the second input / output request in the input / output processor; Controlling the input / output processor to execute the read / write process corresponding to the first input / output request based on the context information of the first input / output request.

8. The method according to claim 7, wherein The method further comprises: After the read / write process corresponding to the second input / output request is interrupted, recording breakpoint information corresponding to the second input / output request into the context information corresponding to the second input / output request; After the read / write process corresponding to the first input / output request is completed, controlling the input / output processor to continue executing the read / write process corresponding to the second input / output request based on the breakpoint information corresponding to the second input / output request.

9. A chip, characterized in that, Comprising a digital integrated circuit, the digital integrated circuit being configured to execute the method according to claim 7 or 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when executed, are used to implement the method according to claim 7 or 8.