Apparatus and method for scheduling and arbitration QOS improvement through early feedback

By introducing processing circuits into the network scheduler, using early feedback information to adjust the data transmission rate, the performance problems caused by the lack of feedback between resource bottlenecks are solved, and more efficient multi-tenant performance and more stable service quality are achieved.

CN120021222APending Publication Date: 2025-05-20SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411615948.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-08-02
Filing Date
2024-11-13
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

In network schedulers, there is a lack of direct feedback between resource bottlenecks, resulting in backup and performance differences, and feedback at a later stage may slow the system in response to changing workloads and resource bottlenecks.

Method used

By introducing a processing circuit into the device, the function of receiving feedback information from the submission queue and the internal queue of the host device is realized, thereby controlling the transmission of data based on these feedbacks. Specifically, the processing circuit receives feedback information from the submission queue and the internal queue and adjusts the rate of data transmission based on this information.

Benefits of technology

Through early feedback mechanisms, the bandwidth and IOPS performance of multi-tenant is improved, performance variance is reduced, and the system's responsiveness is improved, ensuring more stable service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120021222A_ABST
    Figure CN120021222A_ABST
Patent Text Reader

Abstract

The invention relates to an apparatus and method for scheduling and arbitration QOS improvement through early feedback. An apparatus includes: one or more internal queues; and processing circuitry configured to retrieve data from one or more commit queues of the host device, receive first feedback information from one or more internal queues, control transfer of data from the one or more commit queues to the one or more internal queues based on the first feedback information, and transmit the data from the one or more internal queues to the host device. Second feedback information is received from a memory device coupled to the processing circuitry, and transfer of data from the one or more internal queues to the memory device is controlled based on the second feedback information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Application No. 63 / 600,314, filed on November 17, 2023, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present disclosure relates to scheduling and arbitration quality of service (QoS) improvement through early feedback. Background Art

[0004] In a scheduler such as a network scheduler, an arbitration mechanism retrieves commands from a host device and parses the command. Commands can consume resources in the pipeline and can be stored in one or more internal queues. Commands can be scheduled to NAND storage devices. Rate limiters can be used to limit the performance of tenants of the host device by delaying submissions to the completion queue. However, there is no direct feedback between resource bottlenecks, which can cause backups, resulting in high performance differences. In addition, feedback in later stages in the processing pipeline can slow the system down in response to changing workloads and associated resource bottlenecks. Summary of the invention

[0005] According to one or more embodiments, a device includes: one or more internal queues; and a processing circuit configured to: retrieve data from one or more submission queues of a host device, receive first feedback information from the one or more internal queues, control the transfer of data from the one or more submission queues to the one or more internal queues based on the first feedback information, receive second feedback information from a memory device coupled to the processing circuit, and control the transfer of data from the one or more internal queues to the memory device based on the second feedback information.

[0006] According to one or more embodiments, a method performed by at least one processor includes: retrieving data from one or more submission queues of a host device; receiving first feedback information from one or more internal queues; controlling the transfer of data from the one or more submission queues to the one or more internal queues based on the first feedback information, receiving second feedback information from a memory device coupled to the processor, and controlling the transfer of data from the one or more internal queues to the memory device based on the second feedback information.

[0007] According to one or more embodiments, a non-transitory computer-readable medium having instructions stored therein, which, when executed by a processor, causes the processor to perform a method, the method comprising: retrieving data from one or more submission queues of a host device; receiving first feedback information from one or more internal queues; controlling the transfer of data from the one or more submission queues to the one or more internal queues based on the first feedback information; receiving second feedback information from a memory device coupled to the processor; and controlling the transfer of data from the one or more internal queues to the memory device based on the second feedback information. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0009] Figure 1 is a schematic diagram of a scheduler system according to an embodiment of the present disclosure.

[0010] Figure 2 is a schematic diagram of a scheduler system for implementing intermediate feedback between resource bottlenecks according to an embodiment of the present disclosure.

[0011] Figure 3 is a schematic diagram of a scheduler system according to an embodiment of the present disclosure, the scheduler system implementing a scheduler at a central decision-making point, the scheduler implementing deficit weighted round robin (DWRR) to support weighted fair queuing (WFQ).

[0012] Figure 4 is a schematic diagram of a scheduler system implementing credit-based feedback when multiple paths are involved according to an embodiment of the present disclosure.

[0013] Figure 5 is a schematic diagram of a scheduler system for delaying SQ head pointer update according to an embodiment of the present disclosure.

[0014] Fig. 6A An example table for WFQ according to an embodiment of the present disclosure is shown.

[0015] Figure 6B An example table for DWRR according to an embodiment of the present disclosure is shown.

[0016] Figure 6C An example table for DWRR WFQ according to an embodiment of the present disclosure is shown.

[0017] Fig.6D An example table of a rate limiter according to an embodiment of the present disclosure is shown.

[0018] Figure 7 is a flow chart of an example process for controlling data delivery in a scheduler system based on early feedback according to an embodiment of the present disclosure.

[0019] Figure 8 is a block diagram of an example processor according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0020] The following detailed description of example embodiments refers to the accompanying drawings. The same reference numerals in different drawings may identify the same or similar elements. In the drawings, reference numerals beginning with "S" refer to operations of processes or methods.

[0021] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the embodiments to the disclosed precise form. According to the foregoing disclosure, modifications and variations are possible, or can be obtained from the practice of the embodiments. In addition, one or more features or components of an embodiment may be incorporated into another embodiment (or one or more features of another embodiment) or combined with another embodiment (or one or more features of another embodiment). In addition, in the flowchart and description of the operation provided below, it should be understood that one or more operations may be omitted, one or more operations may be added, one or more operations may be performed simultaneously (at least in part), and the order of one or more operations may be switched.

[0022] It is obvious that the systems and / or methods described herein can be implemented in different forms of hardware or firmware. The actual dedicated control hardware used to implement these systems and / or methods does not limit the implementation.

[0023] Even though particular combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features may be combined in ways not specifically recited in the claims and / or disclosed in the specification. Although each dependent claim listed below may be directly dependent on only one claim, the disclosure of possible implementations includes the combination of each dependent claim with every other claim in the claim set.

[0024] Unless explicitly described as such, the elements, actions, or instructions used herein should not be interpreted as critical or essential. In addition, as used herein, the articles "a" and "an" are intended to include one or more items and can be used interchangeably with "one or more". Figure 1In the case of multiple items, the term "one" or similar language is used. In addition, as used herein, the terms "has," "have," "having," "include," "including," and the like are intended to be open-ended terms. In addition, unless otherwise expressly stated, the phrase "based on" is intended to mean "based at least in part on." In addition, expressions such as "at least one of [A] and [B]" or "at least one of [A] or [B]" are to be understood to include only A, only B, or both A and B.

[0025] Reference throughout the specification to "one embodiment," "an embodiment," or similar language means that a particular feature, structure, or characteristic described in connection with the indicated embodiment is included in at least one embodiment of the present solution. Thus, the phrases "in one embodiment," "in an embodiment," and similar language throughout this specification may, but do not necessarily, all refer to the same embodiment.

[0026] In addition, the described features, advantages and characteristics of the present disclosure may be combined in any suitable manner in one or more embodiments. Based on the description herein, those skilled in the relevant art will recognize that the present disclosure may be practiced without one or more specific features or advantages of a particular embodiment. In other cases, additional features and advantages that may not be present in all embodiments of the present disclosure may be recognized in certain embodiments.

[0027] Embodiments of the present disclosure relate to systems, devices, and methods for improving bandwidth (BW) and input / output operations per second (IOPS) for multiple tenants by providing downstream resource usage feedback to a scheduler. Embodiments of the present disclosure improve multi-tenant performance, QoS, and reduced variance of IOPS / BW specifications.

[0028] Figure 1An example scheduler system 100 including a host 102 is shown. The operation of the scheduler system 100 is described with respect to the operation of the method (e.g., S120, etc.). The host 102 can be any device or device group that provides networking or cloud services to one or more tenants. For example, the host 102 can provide cloud computing services to tenants Tenant 0, Tenant 1, Tenant 2, ... Tenant n. In one or more examples, the host 102 can be a system or one or more servers, which includes one or more processors, and also includes one or more memory devices corresponding to one or more submission queues 102A and one or more completion queues 102B. Tenants can submit one or more commands to the corresponding submission queue. For example, each tenant can be associated with a separate set of one or more submission queues 102A, wherein the tenant submits data or commands to one or more submission queues for processing. Therefore, when data or commands are in the submission queue, the data or commands wait to be processed. The amount of time that data or commands wait in the submission queue can correspond to the QoS provided to a particular tenant. In one or more examples, tenants can subscribe to cloud computing services provided by the host 102, wherein based on the subscription level, tenants can be guaranteed a specific QoS. For example, tenants paying a higher subscription fee may be guaranteed a higher level of QoS than tenants paying a lower subscription fee.

[0029] The output of commands from the submission queues may be determined by an arbitration and command parsing function 104 that performs an arbitration operation (S120) that determines the rate at which data or commands are retrieved from each tenant's submission queue. For example, the arbitration and command parsing function 104 may retrieve or retrieve commands from one or more submission queues based on a set of rules or conditions. For example, the set of rules or conditions may be determined based on the QoS guaranteed to a particular tenant. The arbitration and command parsing function 104 may parse commands into smaller sets of commands or functions for processing. For example, the arbitration and command parsing function 104 may parse commands into subsets of commands. The arbitration and command parsing function 104 may be implemented by processing circuitry. Figure 8 An example of a processing circuit is disclosed in .

[0030] The fetched or retrieved command may consume resources in the pipeline (e.g., resource #1). For example, when a command is fetched from a submission queue, the processing of the command (S122) may consume available processing power. For example, once a command is fetched from a submission queue, one or more calculations that consume resources of a processing circuit may need to be performed. The scheduler system 100 may also include one or more internal queues 106 for storing commands. The one or more internal queues 106 may be implemented by a memory circuit. The memory circuit may be part of the same device that includes the processing circuit that implements the arbitration and command parsing function 104.

[0031] The scheduler system 100 may include a backend memory device, such as NAND 108. Commands in one or more internal queues 106 may be output for command processing and stored in NAND 108. As will be appreciated by one of ordinary skill in the art, NAND 108 may be flash memory. NAND 108 flash memory may be a non-volatile storage technology that does not require power to retain data.

[0032] During command processing, a second set of resources (e.g., resources #2) may be consumed during command processing (S124). For example, when a command is retrieved from one of the one or more internal queues 106, the retrieved command may require storing or retrieving data from NAND 108, which may consume resources such as memory bandwidth (e.g., the rate at which data can be stored or retrieved from memory).

[0033] After the data is transferred between the host memory (e.g., one or more submission queues 102A) and the NAND storage device (or device memory), the command is completed. The command or data can be transferred from the NAND 108 to the direct memory access (DMA) engine 110 (S126). In one or more examples, the DMA engine 110 is configured to access the NAND 108 independently of the processing circuit (such as a processing circuit that implements one or more rate control operations). The DMA engine 110 can be used when the CPU cannot keep up with the rate of data transfer, or when the CPU needs to perform work while waiting for slow I / O data transfer. In one or more examples, the DMA engine 110 can perform a memory-to-memory copy or move of data within the memory. The DMA engine 110 can advantageously offload time-consuming and high-bandwidth memory operations, such as large copies or data transfers, from the CPU.

[0034] The scheduler system 100 may include a rate limiter 112 for limiting the performance of a tenant by delaying submission to one or more completion queues 102B. For example, without the rate limiter 112, when a command is completed (S128), the completed command may be returned to the host 102 for completion queue or submission queue processing (S132). In one or more examples, the rate limiter 112 may delay commands to the submission queue. Delaying the release of new commands to the submission queue may result in improved QoS for one or more tenants. The rate limiter 112 may also control the arbitration command and resolution function 104 to control the rate at which commands are retrieved from one or more submission queues.

[0035] In the scheduler system 100, there is no direct feedback between resource bottlenecks, which can cause temporary backups and thus high performance variance. A scheduler deeply embedded in a pipeline (e.g., a pipeline for processing commands) has more visibility into resource usage, but is limited by the availability of commands in the scheduler's queues (e.g., one or more internal queues 106).

[0036] Feedback at later stages (e.g., the rate limiter 112 posts commands to the completion queue or submission queue (S132)) causes the scheduler system 100 to be slow to respond to changing workloads and associated resource bottlenecks. When there are very large queue depths, completion queue processing and submission queue submissions are not very efficient, or have very high variance, resulting in performance degradation for one or more tenants. In the scheduler system 100, in one or more examples, the two scheduling decision points may be the arbitration and command parsing function 104 and the rate limiter 112, which may implement weighted fair queuing (WFQ). However, in the scheduler system 100, only the rate limiter 112 provides feedback, which occurs at the end of the processing pipeline.

[0037] Figure 2 FIG. 2 shows an example of a scheduler system 200 that implements intermediate feedback between resource bottlenecks according to an embodiment of the present disclosure. Figure 2 As shown, the arbitration and command parsing function 104 may receive first feedback 202 from one or more internal queues 106. The first feedback 202 may be credit feedback generated by a processing circuit system (e.g., a memory controller) of the one or more internal queues 106, which indicates downstream resource usage, such as resources consumed by commands retrieved from the one or more submission queues 102A (e.g., resource #1). For example, the first feedback 202 may indicate the amount of resources consumed by each of the tenants 0...N. Based on the first feedback 202, the arbitration and command parsing function 104 may increase or decrease the rate at which commands are retrieved from the one or more submission queues 102A. For example, if the first feedback 202 indicates that the amount of resources consumed is above a threshold, then the arbitration and command parsing function 104 may decrease the rate at which commands are retrieved from the one or more submission queues 102A. In another example, if the first feedback 202 indicates that the amount of resources consumed is below a threshold, then the arbitration and command parsing function 104 may increase the rate at which commands are retrieved from the one or more submission queues 102A.

[0038] In one or more examples, the arbitration and command resolution function 104 can perform an action for a particular tenant. For example, if the first feedback 202 indicates that the amount of resources consumed by Tenant 0 is above a threshold, the arbitration and command resolution function 104 can reduce the rate at which commands are retrieved from one or more submission queues associated with Tenant 0 based on the first feedback 202. In one or more examples, if the first feedback 202 indicates that the amount of resources consumed by Tenant 0 is below a threshold, and the QoS level that Tenant 0 is currently receiving is below the QoS level guaranteed to Tenant 0, the arbitration and command resolution function 104 can increase the rate at which commands are retrieved from one or more submission queues associated with Tenant 0.

[0039] In one or more examples, the arbitration and command resolution function 104 can perform actions for two or more tenants. For example, if the first feedback 202 indicates that the amount of resources consumed by Tenant 0 is above a threshold, and the amount of resources consumed by Tenant 1 is below a threshold, the arbitration command and resolution function 104 can simultaneously reduce the rate at which commands are retrieved from one or more submission queues associated with Tenant 0, and increase the rate at which commands are retrieved from one or more submission queues associated with Tenant 1. In one or more examples, the amount of the rate reduction for Tenant 0 can be equal to the amount of the rate increase for Tenant 1. In one or more examples, the amount of the rate reduction for Tenant 0 can be different from the amount of the rate increase for Tenant 1.

[0040] In one or more examples, the arbitration and command parsing function 104 can perform the same action for each of the tenants. For example, based on the amount of resources consumed by all tenants, the arbitration and command parsing function 104 can increase or decrease the rate at which commands are retrieved for each tenant. For example, if the first feedback 202 indicates that the total amount of resources consumed by the tenants is above a threshold, the arbitration and command parsing function 104 can reduce the rate at which commands are retrieved from one or more submission queues of each tenant. The amount of rate reduction for each tenant can be different or the same. In one or more examples, if the first feedback 202 indicates that the total amount of resources consumed by the tenants is below a threshold, the arbitration and command parsing function 104 can increase the rate at which commands are retrieved from one or more submission queues of each tenant. The amount of rate increase for each tenant can be different or the same.

[0041] In one or more examples, the scheduler system 200 may include a second feedback 204 provided by a NAND management component that monitors data transferred between one or more internal queues 106 and the NAND 108. In one or more examples, the NAND management component may be a processor, such as the processor 802, which is described below with respect to Figure 8Detailed description is provided in further detail. The second feedback 204 may indicate an amount of resources (e.g., resource #2) consumed by commands retrieved from the one or more internal queues 106. Based on the second feedback 204, the rate at which commands are transmitted from the one or more internal queues 106 to the NAND 108 may be increased or decreased. For example, if the second feedback 204 indicates that the amount of resources consumed is above a threshold, the rate at which commands are retrieved from the one or more internal queues 106 may be reduced. In another example, if the second feedback 204 indicates that the amount of resources consumed is below a threshold, the rate at which commands are retrieved from the one or more internal queues 106 may be increased. In one or more examples, the second feedback information 204 may include multiple feedback loops. For example, the NAND management component may include one or more processors that perform NAND management, command aggregation, advanced NAND scheduling, and the like. Therefore, the second feedback information 204 may provide information related to a feedback loop for NAND management, a feedback loop for command aggregation, a feedback loop for advanced NAND scheduling, and the like.

[0042] In one or more examples, the rate at which commands are retrieved from the one or more internal queues 106 can be adjusted for a particular tenant. For example, if the second feedback 204 indicates that the amount of resources consumed by Tenant 0 (e.g., resource #2) is above a threshold, the rate at which commands are retrieved from the one or more internal queues 106 associated with Tenant 0 can be reduced. In one or more examples, if the second feedback 204 indicates that the amount of resources consumed by Tenant 0 is below a threshold, and the QoS that Tenant 0 is currently receiving is below the QoS level guaranteed to Tenant 0, the rate at which commands are retrieved from the one or more internal queues 106 associated with Tenant 0 can be increased.

[0043] In one or more examples, the rate at which commands are retrieved from the one or more internal queues 106 can be adjusted based on the second feedback 204 for two or more tenants. For example, if the second feedback 204 indicates that the amount of resources consumed by Tenant 0 is above a threshold, and the amount of resources consumed by Tenant 1 is below a threshold, the rate at which commands are retrieved from the one or more internal queues associated with Tenant 0 can be reduced, and the rate at which commands are retrieved from the one or more internal queues 106 associated with Tenant 1 can be increased. In one or more examples, the amount by which the rate is reduced for Tenant 0 can be equal to the amount by which the rate is increased for Tenant 1. In one or more examples, the amount by which the rate is reduced for Tenant 0 can be different from the amount by which the rate is increased for Tenant 1.

[0044] In one or more examples, the rate at which commands are retrieved from the one or more internal queues 106 can be adjusted for each of the tenants. For example, based on the amount of resources consumed by all tenants, the rate at which commands are retrieved from the one or more internal queues 106 can be increased or decreased for each tenant. For example, if the second feedback 204 indicates that the total amount of resources consumed by the tenants (e.g., resource #2) is above a threshold, then for each tenant, the rate at which commands are retrieved from the one or more internal queues 106 can be reduced. The amount by which the rate is reduced for each tenant can be different or the same. In one or more examples, if the second feedback 204 indicates that the total amount of resources consumed by the tenants (e.g., resource #2) is below a threshold, then for each tenant, the rate at which commands are retrieved from the one or more internal queues 106 can be increased. The amount by which the rate is increased for each tenant can be different or the same.

[0045] In one or more examples, the first feedback 202 can be affected by the second feedback 204. For example, based on the second feedback 204, the rate at which commands are retrieved from the one or more internal queues 106 can be increased, thereby resulting in a higher amount of memory being available, which affects the rate at which the arbitration and command parsing function 104 retrieves commands from the submission queue 102A. In another example, based on the second feedback 204, the rate at which commands are retrieved from the one or more internal queues can be decreased, thereby resulting in a lower amount of memory being available, which affects the rate at which the arbitration and command parsing function 104 retrieves commands from the submission queue 102A.

[0046] Figure 3 The scheduler system 300 according to an embodiment of the present disclosure implements a scheduler at a central decision-making point, the scheduler implementing a deficit weighted round robin (DWRR) to support WFQ. The scheduler system 300 is described with respect to the operations of the method (e.g., S320, etc.). Figure 3As shown, DWRR 302 can be implemented at one or more internal queues 106. DWRR 302 can adjust the rate at which commands are retrieved from one or more internal queues 106 based on the second feedback 204. As understood by those of ordinary skill in the art, the DWRR operation can be performed by scanning all non-empty internal queues in sequence. Each queue to which DWRR is applied can be associated with a difference counter and a quantum value indicating the maximum number of bytes that can be retrieved from the corresponding queue. When a non-empty internal command queue i is selected, the difference counter for the queue is incremented by the quantum value. Subsequently, the value of the difference counter can be the maximum number of bytes that can be sent in the round: if the difference counter is greater than the size of the command at the head of the queue, the command can be sent, and the value of the counter is decremented by the command size. Subsequently, the size of the next command is compared with the counter value. Once the queue is empty or the value of the counter is insufficient, the next queue will be checked. If the queue is empty, the value of the difference counter is reset to 0. In the scheduler system 300, DWRR 302 can act as a rate limiter. Therefore, in one or more examples, the rate limiter 112 can be eliminated.

[0047] In one or more exemplary examples, the arbitration and command parsing functions can retrieve commands from one or more submission queues based on the arbitration process S320. In one or more examples, the arbitration process S320 can be a WFQ. Figure 3 As described in , rather than inserting commands into one or more submission queues 102A, retrieving commands from one or more submission queues 102A ( 304 ) may be delayed.

[0048] Figure 4 An example scheduler system 400 implementing credit-based feedback according to an embodiment of the present disclosure is shown. In one or more examples, if multiple paths are involved, credit-based feedback can be implemented. Figure 4As shown, feedback 404 is provided to arbitration and command parsing 404 from one or more internal queues 106. Feedback 404 may provide information for slowing down or speeding up the rate of retrieving commands from one or more submission queues 102A. For example, feedback 404 may include a parameter indicating the amount of reduction or increase in the rate of retrieving commands from one or more submission queues 102A. Therefore, compared with the first feedback 202, feedback 402 does not include information indicating resource usage (such as precise credits). In one or more examples, feedback 402 may specify a rate change for one or more tenants. For example, feedback 402 may specify an increase in the rate for retrieving commands from one or more submission queues associated with tenant 0, while a decrease in the rate for retrieving commands from one or more submission queues associated with tenant 1. In one or more examples, feedback 402 may be specified by an array, wherein each index in the array is associated with a specific tenant. For example, feedback 402 may be specified as: [1 -1 0 0]. In this example, feedback 402 may indicate that the rate of tenant 0 increases by 1, the rate of tenant 1 decreases by 1, and the rates of tenants 2 and 3 remain the same. In one or more examples, the unit used for increase or decrease can correspond to any suitable rate known to one of ordinary skill in the art, such as Gb / s, Mb / s, Kb / s, etc.

[0049] like Figure 4 As shown, early feedback 404 may be provided to the host 102 from the arbitration and command parsing function 104. Based on the early feedback 404, the host may control the rate at which commands are inserted into the one or more submission queues 102A. For example, the feedback 404 may indicate that the rate at which commands are inserted into the one or more submission queues 102A may be increased. In another example, the feedback 404 may indicate that the rate at which commands are inserted into the one or more submission queues 102A may be decreased.

[0050] like Figure 4 As shown, one or more internal queues 106 can receive information from the performance monitor 406, which the DWRR 202 function can use to adjust the rate at which commands are retrieved from the one or more internal queues 106. In one or more examples, the performance monitor 406 can provide information corresponding to the performance of one or more tenants over a period of time. For example, with respect to Tenant 0, the feedback 204 can indicate the resource usage of a particular resource (e.g., resource #2) for one or more commands associated with Tenant 0. Conversely, the performance monitor 406 can provide information about the total number of resources (e.g., resource #1 and resource #2) used by Tenant 0 for a series of commands that includes a greater number of commands than one or more commands reflected in the feedback 204. The performance monitor 406 can generate feedback information after completing one or more commands.

[0051] Figure 5 1 shows a scheduler system 500 according to an embodiment of the present invention, which delays submission queue (SQ) head pointer updates to further refine the rate at which commands are inserted into submission queue 102A. Figure 5 As shown, the delayed SQ head function 502 receives information from the performance monitor 406. In one or more examples, each submission queue in the one or more submission queues includes a head pointer (e.g., an SQ head pointer) and a tail pointer (e.g., an SQ tail pointer). The distance between the head pointer and the tail pointer may represent the number of commands in the corresponding submission queue. When commands are retrieved from the submission queue, the position of the head pointer may be updated to reflect that the number of commands in the submission queue has decreased. When commands are inserted into the submission queue, the position of the tail pointer may be updated to reflect that the number of commands in the submission queue has increased. When the distance between the tail pointer and the head pointer is a maximum size (e.g., the submission queue has a maximum number of allowed commands), the insertion of new commands into the submission queue may be delayed or prevented.

[0052] In one or more examples, the delayed SQ header function 502 can receive information from the performance monitor 406. The information from the performance monitor 406 can include performance information (e.g., the amount of resources consumed) for each tenant in the host 102. Based on the information from the performance monitor 406, the SQ header function 502 can delay the update of the SQ head pointer of the submission queue of the corresponding tenant, thereby causing a delay in inserting commands in the queue. For example, if the performance monitor 406 indicates that the amount of resources consumed by tenant 0 is above a threshold, then when retrieving commands from the submission queue, the update of the position of the head pointer of the submission queue associated with tenant 0 is delayed, thereby delaying the insertion of new commands into the submission queue. The SQ header can be updated based on information from the performance monitor 406 for command completion.

[0053] Figures 6A-6D Shown for the above Figure 1-5An example implementation of a queuing algorithm (e.g., WFQ, DWRR) is implemented in any of the queues associated with tenants (e.g., Tenant 0-Tenant N) discussed in the scheduler system described in . In one or more examples, the queuing algorithm can be implemented based on a virtual token (VT). In one or more examples, the VT can be an integer. For example, a VT of three (3) represents three (3) tokens. When a tenant has one or more tokens available, commands included in a queue associated with the tenant (e.g., a submission queue) can be retrieved. The implementation of the queuing algorithm can also be based on a command size (CMD_SIZE), a weight (e.g., WEIGHT0-WEIGHT3), and a refill parameter (e.g., Refill0-Refill3). In one or more examples, the command size can represent the size of a command inserted or retrieved from a queue. In one or more examples, each tenant can be associated with a corresponding weight (e.g., WEIGHT0-WEIGHT3) that can be related to the priority of the tenant. For example, a first tenant having a higher priority than a second tenant can be associated with a weight so that the VT of the first tenant remains higher or decreases at a lower rate than the second tenant. The weights can be specified as a read weight and a write weight. For example, a read weight may be used when retrieving a command from a queue, and a write weight may be used when inserting a command into a queue. In one or more examples, a refill parameter may be added to the VT to increase the value of the VT. In one or more examples, a refill parameter may be added to the VT at a predetermined time.

[0054] Figures 6A to 6D Various processes for updating the VT are shown. Fig. 6A An example table for WFQ according to an embodiment of the present disclosure is shown. The WFQ algorithm can search for the smallest VT from all active queues. The WFQ algorithm can update the VT based on the command size and the associated weight value (e.g., WEIGHT0-WEIGHT3) for the corresponding tenant. The WFQ algorithm can limit the maximum difference between VTs.

[0055] Figure 6B An example table for DWRR according to an embodiment of the present disclosure is shown. The DWRR algorithm may perform round-robin across all eligible tenants. Figure 6BAs shown, the corresponding VT for each tenant can be updated based on the command size, the corresponding weight for each tenant (e.g., WEIGHT0–WEIGHT3), and the corresponding refill parameter for each tenant (e.g., Refill0–Refill3). DWRR can select all tenants having CMD_SIZE 'n' * weight 'n' < VT 'n'. As understood by those of ordinary skill in the art, DWRR provides implementation flexibility. If CMD_SIZE 'n' is unknown, then the VT is allowed to become negative. The DWRR algorithm can be configured to act like a rate limiter. For example, instead of polling, the DWRR algorithm can implement fixed time slots. Additionally, the DWRR algorithm may be overprovisioned to address backend jitter.

[0056] Figure 6C An example table for DWRR (WFQ) according to an embodiment of the present disclosure is shown. Compared with Fig. 6A the algorithm shown, Figure 6C the algorithm in

[0057] Fig.6D updates the corresponding token for each tenant by subtracting the command size from the corresponding token. The DWRR (WFQ) algorithm can implement polling across all eligible tenants. The DWRR (WFQ) algorithm can select all eligible tenants having VT 'n' > 0 and non-empty queues. The DWRR (WFQ) algorithm can update the VT by subtracting CMD_Size * weight 'n'. If no tenant is selected, the DWRR (WFQ) algorithm can update the VT. The DWRR (WFQ) algorithm can use one of the following two options: (1) reset all VTs to the same configured value; or (2) add the configured value to the VT and saturate the VT to the maximum configured value. In option (2), idle tenants can get additional access, but it may be more complex to implement. Figure 6B the algorithm shown, as Fig.6D shown, the corresponding refill parameter (e.g., Refill0 - Refill3) is added to the VT of the corresponding tenant at the corresponding time (e.g., t0 - t3). The rate limiter can implement polling across all eligible tenants. The rate limiter can select all tenants having VT 'n' > 0 and non-empty queues. The rate limiter can update the VT by subtracting CMD_Size * weight 'n'. The rate limiter can use one or more of the following reload options: (1) a timer-based configured value; (2) a per-tenant timer-based configured value; (3) a per-tenant timer-based configured value with a saturation value.

[0058] Figure 7A flow chart of an example process 700 for controlling the transfer of data in a scheduler system based on early feedback is shown. In one or more examples, the process 700 can begin at operation S702, where data is retrieved from one or more submission queues of a host device. For example, a processing circuit implementing arbitration and command parsing function 104 can retrieve one or more commands from one or more submission queues of host 102.

[0059] The process proceeds to operation S704, where first feedback information is received from one or more internal queues. For example, the processing circuit implementing the arbitration and command parsing function 104 may receive the first feedback information 202 ( Figure 2 ) or the first feedback information 402 ( Figure 4 ) and adjusts the rate at which commands are retrieved from one or more submission queues from host 102 accordingly.

[0060] The process proceeds to operation S706, where data transfer from the one or more submission queues to the one or more internal queues is controlled based on the first feedback information. For example, based on the first feedback information, which may indicate the amount of resources consumed, the processing circuitry implementing the arbitration and command parsing function 104 may speed up or slow down the rate at which commands are retrieved from the one or more submission queues.

[0061] The process proceeds to operation S708, where second feedback information is received from the memory device. For example, the processing circuit may receive second feedback information 204 from NAND 108, where the second feedback information may indicate the amount of resources (eg, memory bandwidth) consumed by commands retrieved from one or more internal queues.

[0062] The process proceeds to operation S710, where data transfer from one or more internal queues to the memory device is controlled based on the second feedback information. For example, based on the second feedback information 202, the processing circuit can increase or decrease the rate at which commands are taken from one or more internal queues 106 and forwarded to NAND 108.

[0063] Figure 8 is a block diagram of example components of one or more devices implementing embodiments of the present disclosure. Figure 8 As shown, device 800 may include a bus 810 , a processor 820 , a memory 830 , a storage component 840 , an input component 850 , an output component 860 , and a communication interface 870 .

[0064] Bus 810 includes components that allow communication between components of device 800. Processor 820 is implemented in hardware, firmware, or a combination of hardware and software. In one or more examples, processor 820 may include a processor for implementing Figure 2 -The embodiment shown in 6 and Figure 7 700. Figure 8 One processor 802 is shown, but as will be appreciated by one of ordinary skill in the art, device 800 may include any number of desired processors.

[0065] The processor 820 is a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or another type of processing component. In some embodiments, the processor 820 includes one or more processors that can be programmed to perform functions. The memory 830 includes a random access memory (RAM), a read-only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic memory, and / or optical memory) that stores information and / or instructions for use by the processor 820. In one or more examples, the memory 830 may correspond to one or more internal queues 106.

[0066] Storage component 840 stores information and / or software related to the operation and use of device 800. For example, storage component 840 may include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid-state disk), a compact disk (CD), a digital versatile disk (DVD), a floppy disk, a cassette, a magnetic tape, and / or another type of non-transitory computer-readable medium and a corresponding drive.

[0067] Input components 850 include components that allow device 800 to receive information, such as via user input (e.g., a touch screen display, a keyboard, a keypad, a mouse, buttons, switches, and / or a microphone). Additionally or alternatively, input components 850 may include sensors for sensing information (e.g., a global positioning system (GPS) component, an accelerometer, a gyroscope, and / or an actuator). Output components 860 include components that provide output information from device 800 (e.g., a display, a speaker, and / or one or more light emitting diodes (LEDs)).

[0068] The communication interface 870 includes transceiver-like components (e.g., transceivers and / or separate receivers and transmitters) that enable the device 800 to communicate with other devices, such as via a wired connection, a wireless connection, or a combination of wired and wireless connections. The communication interface 870 can allow the device 800 to receive information from another device and / or provide information to another device. For example, the communication interface 870 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi interface, a cellular network interface, etc.

[0069] Device 800 may perform one or more processes described herein. Device 800 may perform these processes in response to processor 820 executing software instructions stored by a non-transitory computer-readable medium (e.g., memory 830 and / or storage component 840). Computer-readable media is defined herein as a non-transitory memory device. A memory device includes memory space within a single physical storage device or memory space distributed across multiple physical storage devices.

[0070] The software instructions may be read into the memory 830 and / or storage component 840 from another computer-readable medium or from another device via the communication interface 870. When executed, the software instructions stored in the memory 830 and / or storage component 840 may cause the processor 820 to perform one or more processes described herein. Additionally or alternatively, hardwired circuitry may be used in place of or in combination with software instructions to perform one or more processes described herein. Therefore, the embodiments described herein are not limited to any specific combination of hardware circuitry and software.

[0071] Figure 8 The number and arrangement of components shown in the figure are provided as examples. In practice, the device 800 may include Figure 8 Additional components, fewer components, different components, or differently arranged components may be included in the components shown in FIG. Additionally or alternatively, one set of components (eg, one or more components) of device 800 may perform one or more functions described as being performed by another set of components of device 800.

[0072] Embodiments have been described above, and as shown in the accompanying drawings, embodiments are shown in the form of blocks, which perform one or more functions described. These blocks can be physically implemented by analog and / or digital circuits including one or more of logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, etc., and can also be implemented or driven by software and / or firmware (configured to perform the functions or operations described herein). The circuit can be embodied in one or more semiconductor chips, for example, or embodied on a substrate support such as a printed circuit board. The circuit included in the block can be implemented by dedicated hardware, or by a processor (for example, one or more programmed microprocessors and associated circuits), or by a combination of dedicated hardware that performs some functions of the block and a processor that performs other functions of the block. Each block of the embodiment can be physically divided into two or more interactive and discrete blocks. Similarly, the blocks of the embodiment can be physically combined into more complex blocks.

[0073] In one or more examples, each of the arbitration and command parsing function 104, the DMA engine 110, the rate limiter 112, the DWRR 302, the performance monitor 406, and the delayed SQ header function 502 can be implemented by a processing circuit such as a processor 802. In one or more examples, one or more processors 802 can be used for each of these modules. In one or more examples, the processor 802 can retrieve executable instructions from the storage component 840 to perform the above Figure 1-5 The functionality of these modules is described. In one or more examples, the internal queue 106 can be implemented by a memory 830, where the processor 802 operating as a DWRR 302 communicates with the memory 830 to retrieve data.

[0074] In one or more examples, the device 800 can correspond to a host 102, where the processor 802 can retrieve execution instructions from the storage component 840 to perform processing functions of the host 102, such as inserting and retrieving commands from one or more submission queues 102A. In one or more examples, the one or more submission queues 102A and the completion queues 102B can be implemented by a memory 830, where the processor 802 communicates with the memory 830 to retrieve data from or insert data into the one or more submission queues 102A or completion queues 102B.

[0075] Although the present disclosure has described several non-limiting embodiments, there are changes, permutations, and various substitute equivalents that fall within the scope of the present disclosure. Therefore, it should be understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within its spirit and scope.

[0076] The above disclosure also encompasses the following embodiments:

[0077] (1) A device comprising: one or more internal queues; and a processing circuit configured to: retrieve data from one or more submission queues of a host device, receive first feedback information from the one or more internal queues; control the transfer of data from the one or more submission queues to the one or more internal queues based on the first feedback information, receive second feedback information from a memory device coupled to the processing circuit, and control the transfer of data from the one or more internal queues to the memory device based on the second feedback information.

[0078] (2) An apparatus according to feature (1), wherein the first feedback information includes credit information, and the credit information controls the transfer of data from the one or more submission queues to the one or more internal queues.

[0079] (3) A device according to feature (2), wherein the credit information includes information indicating an amount of resources used by data transmitted from the one or more submission queues to the one or more internal queues, and wherein the processing circuit is further configured to adjust the rate at which commands are retrieved from the one or more submission queues based on the first feedback information.

[0080] (4) An apparatus according to any one of features (1)-(3), wherein the first feedback information indicates a rate change for changing a transfer rate of data from the one or more submission queues to the one or more internal queues.

[0081] (5) An apparatus according to feature (4), wherein the rate change increases the transfer rate of data from the one or more submission queues to the one or more internal queues.

[0082] (6) An apparatus according to feature (4), wherein the rate change reduces the rate at which data is transferred from the one or more submission queues to the one or more internal queues.

[0083] (7) A device according to any one of features (1)-(6), wherein the processing circuit is further configured to implement a deficit weighted round-robin scheduling algorithm based on the second feedback information for transmitting data to and from one or more internal command queues.

[0084] (8) The device of feature (7), wherein the processing circuit is further configured to control the transmission of data from the one or more internal command queues based on the performance monitoring information after the command is completed based on a deficit weighted polling algorithm.

[0085] (9) The device of feature (8), wherein the processing circuit is further configured to delay updating of a head pointer associated with at least one queue from the one or more submission queues based on the first feedback information.

[0086] (10) The device according to any one of features (1)-(9), wherein the second feedback information includes information indicating an amount of resources used by data transmitted from the one or more internal queues to the memory device.

[0087] (11) The device of feature (10), wherein the second feedback information includes information indicating an amount and quantity of memory bandwidth used by data transferred from the one or more internal queues to the memory device.

[0088] (12) An apparatus according to any one of features (1)-(11), wherein the data includes one or more commands.

[0089] (13) A method performed by at least one processor, the method comprising: retrieving data from one or more submission queues of a host device; receiving first feedback information from one or more internal queues; controlling the transfer of data from the one or more submission queues to the one or more internal queues based on the first feedback information, receiving second feedback information from a memory device coupled to the processor, and controlling the transfer of data from the one or more internal queues to the memory device based on the second feedback information.

[0090] (14) A method according to feature (13), wherein the first feedback information includes credit information, and the credit information controls the transfer of data from one or more submission queues to one or more internal queues.

[0091] (15) A method according to feature (14), wherein the credit information includes information indicating an amount of resources used by data transmitted from one or more submission queues to one or more internal queues, and wherein the method also includes adjusting a rate at which commands are retrieved from the one or more submission queues based on the first feedback information.

[0092] (16) A method according to any one of features (13)-(15), wherein the first feedback information indicates a rate change for changing a transfer rate of data from one or more submission queues to one or more internal queues.

[0093] (17) A method according to feature (16), wherein the rate change increases the transfer rate of data from one or more submission queues to one or more internal queues.

[0094] (18) A method according to feature (16), wherein the rate change reduces the rate at which data is transferred from one or more submission queues to one or more internal queues.

[0095] (19) A method according to any one of features (13)-(18), wherein controlling the transfer of data from one or more internal queues to a memory device is also based on a differential weighted round-robin scheduling algorithm, wherein the differential weighted round-robin scheduling algorithm controls the transfer of data to and from one or more internal command queues to the memory device based on second feedback information.

[0096] (20) A non-transitory computer-readable medium having instructions stored therein, which, when executed by a processor, causes the processor to perform a method comprising: retrieving data from one or more submission queues of a host device; receiving first feedback information from one or more internal queues; controlling the transfer of data from the one or more submission queues to the one or more internal queues based on the first feedback information; receiving second feedback information from a memory device coupled to the processor; and controlling the transfer of data from the one or more internal queues to the memory device based on the second feedback information.

Claims

1. A device for scheduling, comprising: One or more internal queues; as well as The processing circuit is configured to: Retrieve data from one or more submission queues of the host device, receiving first feedback information from the one or more internal queues, controlling the transfer of the data from the one or more submission queues to the one or more internal queues based on the first feedback information, receiving second feedback information from a memory device coupled to the processing circuit, and The transferring of the data from the one or more internal queues to the memory device is controlled based on the second feedback information.

2. The device according to claim 1, wherein: The first feedback information includes credit information that controls the transfer of data from the one or more submission queues to the one or more internal queues.

3. The device according to claim 2, wherein the credit information includes information indicating an amount of resources used by the data transferred from the one or more submission queues to the one or more internal queues, and Wherein the processing circuit is further configured to adjust a rate at which data is retrieved from the one or more submission queues based on the first feedback information.

4. The device according to claim 1, wherein: The first feedback information indicates a rate change for changing a transfer rate at which data is transferred from the one or more submission queues to the one or more internal queues.

5. The device according to claim 4, wherein: The rate change increases the transfer rate of the data from the one or more submission queues to the one or more internal queues.

6. The device according to claim 4, wherein: The rate change reduces the transfer rate of the data from the one or more submission queues to the one or more internal queues.

7. The device according to claim 1, wherein: The processing circuit is further configured to implement a deficit weighted round-robin scheduling algorithm based on the second feedback information for the transmission of the data to and from the one or more internal queues.

8. The device according to claim 7, wherein: The processing circuit is further configured to control the transfer of data from the one or more internal queues based on performance monitoring information after command completion based on the deficit weighted round-robin algorithm.

9. The device according to claim 8, wherein: The processing circuitry is further configured to delay updating of a head pointer associated with at least one queue from the one or more submission queues based on the first feedback information.

10. The device according to claim 1, Wherein the second feedback information includes information indicating an amount of resources used by the data transferred from the one or more internal queues to the memory device.

11. The device of claim 10, wherein the second feedback information comprises information indicating an amount and quantity of memory bandwidth used by the data transferred from the one or more internal queues to the memory device.

12. The device according to claim 1, wherein: The data includes one or more commands.

13. A method performed by at least one processor, the method comprising: Retrieving data from one or more submission queues of the host device; receiving first feedback information from one or more internal queues; controlling the transfer of the data from the one or more submission queues to the one or more internal queues based on the first feedback information, receiving second feedback information from a memory device coupled to the processor, and The transferring of the data from the one or more internal queues to the memory device is controlled based on the second feedback information.

14. The method according to claim 13, wherein: The first feedback information includes credit information that controls the transfer of data from the one or more submission queues to the one or more internal queues.

15. The method according to claim 14, wherein the credit information includes information indicating an amount of resources used by the data transferred from the one or more submission queues to the one or more internal queues, and The method further includes adjusting a rate at which data is retrieved from the one or more submission queues based on the first feedback information.

16. The method according to claim 13, wherein: The first feedback information indicates a rate change for changing a transfer rate at which data is transferred from the one or more submission queues to the one or more internal queues.

17. The method according to claim 16, wherein: The rate change increases the transfer rate of the data from the one or more submission queues to the one or more internal queues.

18. The method according to claim 16, wherein: The rate change reduces the transfer rate of the data from the one or more submission queues to the one or more internal queues.

19. The method according to claim 13, wherein: Controlling the transfer of the data from the one or more internal queues to the memory device is also based on a quota weight round-robin scheduling algorithm, which controls the transfer of the data to and from the one or more internal queues to the memory device based on the second feedback information.

20. A non-transitory computer readable medium having stored therein instructions which, when executed by a processor, cause the processor to perform a method comprising: Retrieving data from one or more submission queues of the host device; receiving first feedback information from one or more internal queues; controlling the transfer of the data from the one or more submission queues to the one or more internal queues based on the first feedback information; receiving second feedback information from a memory device coupled to the processor; as well as The transferring of the data from the one or more internal queues to the memory device is controlled based on the second feedback information.