I / O time delay fault injection method and device, server and storage medium

By temporarily existing Outstanding Tracker queue in the SPDK process and submitting I/O requests when the set delay arrives, the shortcomings of I/O delay failure simulation in the storage system are solved, and a comprehensive verification of the reliability of the storage system is achieved.

CN120256180APending Publication Date: 2025-07-04NEW H3C TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510377094.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

When existing storage systems verify the I/O delay fault response capabilities of NVMe SSDs, they lack effective simulation methods, making it difficult to fully verify the reliability of the system.

Method used

When the SPDK process receives an I/O request in the user state, it temporarily stores it in the Outstanding Tracker queue and submits it to the submission queue when the set delay time arrives, delay processing is implemented and I/O delay failure scenario is constructed.

Benefits of technology

Accurately simulates I/O delay failures, which can fully verify the reliability of the storage system. By configuring the delay time and the threshold of the number of I/O requests, it reproduces complex and changeable real delay scenarios, improving the comprehensiveness and accuracy of verification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256180A_ABST
    Figure CN120256180A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an I / O time delay fault injection method and device, a server and a storage medium. In the application, the I / O request received by the SPDK process is not directly put into the submission queue for SSD processing, but is temporarily stored in the Outstanding Tracker queue, and the I / O request is put into the submission queue for SSD processing only when the issuing moment (i.e., when the set delay time is up) is reached, so that the I / O request is submitted in a delayed manner, and the delayed processing is realized based on the delayed submission; and the I / O time delay fault is constructed based on delay processing, so that the I / O time delay fault scene is accurately simulated, and the reliability of the storage system is fully verified based on the constructed I / O time delay fault.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and in particular, to an I / O latency fault injection method, apparatus, server, and storage medium. Background Art

[0002] Since the Storage Performance Development Kit (SPDK) can accelerate the read and write performance of solid state drives (SSDs) that adopt the Non-Volatile Memory Express Solid State Disk (NVMe) standard, the SPDK is widely applied to storage systems.

[0003] In a storage system based on the SPDK, it is necessary to check the response capabilities of the storage system for I / O latency faults such as hot plugging of NVMe SSDs, input / output (I / O) timeouts, and slow I / Os, so as to fully verify the reliability of the storage system.

[0004] Therefore, there is an urgent need for a method to construct the above I / O latency faults to better verify the reliability of the storage system. Summary of the Invention

[0005] In view of this, embodiments of this application provide an I / O latency fault injection method, apparatus, server, and storage medium, which can accurately simulate I / O latency fault scenarios, thereby better verifying the reliability of the storage system.

[0006] Embodiments of this application provide an I / O latency fault injection method, which is applied to an electronic device. The method includes:

[0007] When an input / output (I / O) request is received through a storage performance development kit process SPDK thread in the user space, according to the configured latency time and the I / O request quantity threshold of the process, the I / O request is placed into an Outstanding Tracker queue;

[0008] According to the configured latency time of the process and the reception time of the received I / O request, determine the issuing moment for issuing the I / O request from the Outstanding Tracker queue to a submission queue different from the Outstanding Tracker queue, and when the issuing moment arrives, issue the I / O request from the Outstanding Tracker queue to the submission queue, so as to process the I / O request in the submission queue and after the processing, issue the I / O request from the request queue to the completed queue;

[0009] The duration from when the I / O request is received to when it is sent to the completed queue through this process is sent to the upper-layer service, so that the upper-layer service can obtain the latency fault of the executed I / O request based on the duration.

[0010] As an embodiment, the latency time is determined according to the latency time indicated by the latency type of the latency fault to be injected; the latency time indicated by any latency type is the latency time specified in the detection rule for detecting this latency type.

[0011] As an embodiment, the I / O request quantity threshold is determined according to the maximum count value configured for the process and the percentage indicated by the latency type of the latency fault to be injected; the percentage indicated by any latency type is the percentage specified in the detection rule for detecting this latency type.

[0012] As an embodiment, putting the I / O request into the Outstanding Tracker queue according to the latency time configured for this process and the I / O request quantity threshold includes:

[0013] When the latency time configured for this process is not the set value, determine whether the current count of the counter configured for this process is greater than the I / O request quantity threshold;

[0014] If not, determine to put the I / O request into the Outstanding Tracker queue;

[0015] If so, determine to send the I / O request to the submission queue.

[0016] As an embodiment, this method further includes:

[0017] Every time this process receives an I / O request, the current count corresponding to this process is incremented by the set value, and when the current count of this process reaches the maximum count value corresponding to this process, it is reset to the initial value.

[0018] As an embodiment, determining the sending time for sending the I / O request from the Outstanding Tracker queue to a submission queue different from the Outstanding Tracker queue according to the latency time configured for this process and the reception time of the received I / O request includes:

[0019] Obtain the sum of the latency time and the reception time; wherein, if the sum of the times is greater than or equal to the current time, determine that the sending time has arrived; otherwise, determine that the sending time has not arrived.

[0020] As an embodiment, after the I / O request is sent from the Outstanding Tracker queue to the submission queue, the method further includes:

[0021] Set a tag for the I / O request in the Outstanding Tracker queue to indicate that it has been sent.

[0022] As an embodiment, the method further includes:

[0023] Receive a Remote Procedure Call (RPC) instruction, and when the delay time indicated by the instruction is not the set value and the percentage is not the set value, enable the I / O latency fault injection function of the process; after the I / O latency fault injection function is enabled, the process is configured with a latency time and an I / O request quantity threshold.

[0024] As an embodiment, the method further includes:

[0025] When an I / O latency fault injection cancellation event is detected, or it is determined that the delay time indicated by the received PCR is the set value, send each unsent I / O request in the Outstanding Tracker queue to the submission queue.

[0026] An embodiment of the present application further provides an I / O latency fault injection device, which is applied to an electronic device, and the device includes:

[0027] A receiving module, configured to, when receiving an input / output (I / O) request through a Storage Performance Development Kit (SPDK) thread in user space, put the I / O request into an Outstanding Tracker queue according to the latency time and the I / O request quantity threshold configured for the process;

[0028] A sending module, configured to determine a sending time for sending the I / O request from the Outstanding Tracker queue to a submission queue different from the Outstanding Tracker queue according to the latency time configured for the process and the receiving time of the received I / O request, and when the sending time arrives, send the I / O request from the Outstanding Tracker queue to the submission queue to process the I / O request in the submission queue and send the I / O request from the request queue to a completed queue after processing;

[0029] A transmitting module, configured to send the duration from when the I / O request is received to when it is sent to the completed queue by the process to an upper-layer service, so that the upper-layer service can obtain an I / O request execution latency fault based on the duration.

[0030] As an example, the latency time is determined according to the latency time indicated by the latency type of the latency fault to be injected; the latency time indicated by any latency type is the latency time specified in the detection rule for detecting this latency type.

[0031] As an example, the I / O request quantity threshold is determined according to the maximum count value configured for the process and the percentage indicated by the latency type of the latency fault to be injected; the percentage indicated by any latency type is the percentage specified in the detection rule for detecting this latency type.

[0032] As an example, putting the I / O request into the Outstanding Tracker queue according to the latency time configured for the process and the I / O request quantity threshold includes:

[0033] When the latency time configured for the process is not the set value, determine whether the current count of the counter configured for the process is greater than the I / O request quantity threshold;

[0034] If not, determine to put the I / O request into the Outstanding Tracker queue;

[0035] If so, determine to send the I / O request to the submission queue.

[0036] As an example, the receiving module is further configured to:

[0037] For each I / O request received by the process, the current count corresponding to the process is incremented by a set value, and when the current count of the process reaches the maximum count value corresponding to the process, it is reset to the initial value.

[0038] As an example, determining the sending moment for sending the I / O request from the Outstanding Tracker queue to a submission queue different from the Outstanding Tracker queue according to the latency time configured for the process and the receiving time of the received I / O request includes:

[0039] Obtain the sum of the latency time and the receiving time; wherein, if the sum of the times is greater than or equal to the current time, determine that the sending moment has arrived; otherwise, determine that the sending moment has not arrived.

[0040] As an example, the sending module is further configured to: after sending the I / O request from the Outstanding Tracker queue to the submission queue, set a label for indicating that it has been sent for the I / O request in the Outstanding Tracker queue.

[0041] As an embodiment, the apparatus further includes: an enabling module, configured to receive a Remote Procedure Call (RPC) instruction, and enable the I / O latency fault injection function of a process when the latency time indicated by the instruction is not a set value; after the I / O latency fault injection function is enabled, the process is configured with a latency time and a threshold of the number of I / O requests.

[0042] As an embodiment, the apparatus further includes: an execution module, configured to, when detecting an I / O latency fault injection cancellation event, or determining that the latency time indicated by the received PCR is a set value, issue each outstanding I / O request in the Outstanding Tracker queue to the submission queue.

[0043] An embodiment of the present application further provides an electronic device, including: a processor and a computer-readable storage medium for storing computer program instructions, where the computer program instructions, when run by the computer-readable storage medium, cause the processor to execute the steps of the above method.

[0044] An embodiment of the present application further provides a machine-readable storage medium, which stores computer program instructions, and when the computer program instructions are executed, can implement the steps of the above method.

[0045] It can be seen from the above technical solutions that, in this embodiment, when the SPDK process in the user mode receives an I / O request, according to the configured latency time and the threshold of the number of I / O requests of the process, the I / O request is placed in the Outstanding Tracker queue, and then according to the configured latency time of the process and the reception time of the received I / O request, the issue time for issuing the I / O request from the Outstanding Tracker queue to the submission queue is determined, and when the issue time arrives, the I / O request is issued from the Outstanding Tracker queue to the submission queue, so that the SSD obtains the I / O request from the submission queue for processing, and finally the duration from receiving the I / O request to issuing it to the completed queue is sent to the upper-layer service, and the upper-layer service can obtain the latency fault according to the duration corresponding to the I / O request.

[0046] This way of not directly putting the I / O request into the submission queue for the SSD to process, but temporarily storing it in the Outstanding Tracker queue and putting the I / O request into the submission queue for the SSD to process only when the issue time arrives (i.e., when the set latency time arrives) realizes delaying the submission of the I / O request, realizes delayed processing based on the delayed submission, constructs the I / O latency fault based on the delayed processing, so as to accurately simulate the I / O latency fault scenario, and fully verify the reliability of the storage system based on the constructed I / O latency fault.

[0047] Further, in the above method, different types of I / O delay faults can be constructed by configuring the delay time and the I / O request quantity threshold, so as to reproduce complex and changeable real delay scenarios, and further fully verify the reliability of the storage system. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 The storage architecture diagram provided by the embodiment of the present application;

[0049] Figure 2 The flowchart of the method provided by the embodiment of the present application;

[0050] Figure 3 The flowchart of determining whether to put the I / O request into the Outstanding Tracker queue provided by the embodiment of the present application;

[0051] Figure 4 The flowchart of sending the I / O request from the Outstanding Tracker queue to the submission queue provided by the embodiment of the present application;

[0052] Figure 5 The structural schematic diagram of the device provided by the embodiment of the present application;

[0053] Figure 6 The structural schematic diagram of the electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] In order to enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application, and to make the above-mentioned objects, features, and advantages of the embodiments of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0055] Before introducing the method provided by the embodiment of the present application, the storage architecture provided by the embodiment of the present application will be described first:

[0056] See Figure 1 , Figure 1 the storage architecture diagram provided by the embodiment of the present application. This storage architecture is the implementation environment of the method provided by the embodiment of the present application, as Figure 1As shown in the figure, the storage architecture involves upper-layer services, SPDK, and NVMe SSDs (which can also be referred to as SSD hard drives). The upper-layer services issue I / O requests. Each process in the user-space SPDK receives the I / O request and processes the I / O request in the manner provided in the embodiments of the present application (the specific process will be described in detail in the embodiments hereinafter and will not be elaborated here). The time taken from receiving the I / O request to completing its processing (for the sake of easy description, hereinafter referred to as the completion time) is sent to the upper-layer services. The upper-layer services obtain the latency faults of the executed I / O requests based on the completion users corresponding to the received I / O requests and the set detection logic.

[0057] Here, SPDK is a user-space, polled-mode, asynchronous, lockless NVMe SSD driver. SPDK directly accesses NVMe SSDs in the user space, avoiding the switch between the kernel space and the user space and reducing the storage overhead. SPDK processes the completed I / O requests of NVMe SSDs in a polling manner. This alternative to the traditional kernel interrupt mode can reduce latency and improve storage performance.

[0058] It should be noted that the above storage architecture can be deployed on an electronic device (such as a server) or on distributed electronic devices, which can be set according to specific storage requirements and are not specifically limited in the embodiments of the present application.

[0059] Combined with the above storage architecture, the method provided in the embodiments of the present application will be described below:

[0060] Please refer to Figure 2 , Figure 2 which is the flowchart of the method provided in the embodiments of the present application. As Figure 2 shown, the process may include the following steps:

[0061] S201, when an input / output I / O request is received by the storage performance development kit process SPDK thread in the user space, the I / O request is placed in the Outstanding Tracker queue for uncompleted task tracking according to the configured latency time and the I / O request quantity threshold of the process.

[0062] As an example, before performing the above step S201, the method further includes: enabling the I / O latency fault injection function of the above process. Optionally, the spdk_for_each_channel() interface is called to receive a Remote Procedure Call (RPC) command. If the latency time indicated by the RPC instruction is not a set value (such as 0) and the percentage is not a set value (such as 0), the I / O latency fault injection function of the process is enabled. During the enabling process of the latency fault injection function, the SPDK process is configured with a latency time and an I / O request quantity threshold. As for how to specifically configure the latency time and the I / O request quantity threshold of the SPDK process, it will be described in a specific example later and will not be elaborated here.

[0063] When the PDK process receives an I / O request, if it is determined to put the I / O request into the Outstanding Tracker queue according to the latency time and the I / O request quantity threshold configured for the process, the I / O request is put into the Outstanding Tracker queue. As for the specific implementation of determining whether to put the I / O request into the Outstanding Tracker queue according to the latency time and the I / O request quantity threshold configured for the process, it will be described later and will not be elaborated here.

[0064] S202, according to the latency time configured for the process and the reception time of the received I / O request, determine the issuance moment of the I / O request from the Outstanding Tracker queue to a submission queue different from the Outstanding Tracker queue, and when the issuance moment arrives, issue the I / O request from the Outstanding Tracker queue to the submission queue to process the I / O request in the submission queue and after processing, issue the I / O request from the request queue to the completed queue.

[0065] Here, the implementation of step S202 requires the use of the polling mechanism of SPDK. First, the polling mechanism of SPDK will be introduced. The polling mechanism of SPDK is used to poll the Complete Queue (CQ queue) to read the I / O requests that have been processed by the current NVMe SSD but have not yet returned the completion time to the upper-layer service. The polling is periodic polling, that is, the CQ is queried every set time period, such as 0.002 μs. Based on this, at each detection time point, the I / O requests in the Outstanding Tracker queue are traversed to determine whether to send them to the submission queue, and the CQ queue is queried after traversing the Outstanding Tracker queue. As for the above judgment on whether the I / O request has reached the sending time, it will be elaborated in the following with specific embodiments and will not be repeated here.

[0066] S203, send the time taken for the I / O request to be sent from reception to the completed queue by this process to the upper-layer service, so that the upper-layer service can obtain the latency fault of the I / O request being executed based on the time.

[0067] In this embodiment, the upper-layer service detects the latency fault of the I / O request being executed based on the time taken for the I / O request to be sent from reception to the completed queue and using the set detection rules.

[0068] So far, the Figure 2 shown process is completed.

[0069] Through Figure 2 As can be seen from the shown process, in this embodiment, when the SPDK process in the user state receives an I / O request, according to the configured latency time and the I / O request quantity threshold of this process, the I / O request is put into the Outstanding Tracker queue, and then according to the configured latency time of this process and the reception time of the received I / O request, the sending time for sending the I / O request from the Outstanding Tracker queue to the submission queue is determined, and when the sending time arrives, the I / O request is sent from the Outstanding Tracker queue to the submission queue, so that the SSD can obtain the I / O request from the submission queue for processing, and finally the time taken for the I / O request to be sent from reception to the completed queue is sent to the upper-layer service, and the upper-layer service can obtain the latency fault according to the time corresponding to the I / O request.

[0070] Instead of directly putting the I / O request into the submission queue for the SSD to process, it is temporarily stored in the OutstandingTracker queue and only put into the submission queue for the SSD to process when the sending time arrives (i.e., when the set delay time arrives). This method realizes the delayed submission of the I / O request, achieves delayed processing based on the delayed submission, constructs the I / O delay fault based on the delayed processing, thus accurately simulating the I / O delay fault scenario to fully verify the reliability of the storage system based on the constructed I / O delay fault.

[0071] Furthermore, in the above method, different types of I / O delay faults with different delays can be constructed by configuring the delay time and the I / O request quantity threshold, reproducing complex and variable real delay scenarios to further fully verify the reliability of the storage system.

[0072] The following elaborates in detail on how to configure the delay time and the I / O request quantity threshold of the SPDK process:

[0073] The delay time configured for the above SPDK process is determined according to the delay time indicated by the delay type of the delay fault to be injected. The delay time indicated by any delay type is the delay time specified in the detection rule for detecting this delay type. Here, the delay types include slow I / O, I / O timeout, hot plug, etc. Different delay types have different detection rules, and each detection rule for a delay type stipulates the delay time. When the upper-layer service detects that the current delay time of the detected I / O request (the difference between the current completion time and the set completion time) exceeds the delay time specified by this rule, it is determined that the I / O request has a delay fault of this delay type.

[0074] For example, if the delay type of the delay fault to be injected is I / O timeout, and the detection rule corresponding to I / O timeout stipulates that: if the current delay time of the I / O request exceeds 2 ms, it is determined that there is a delay fault of I / O timeout. At this time, the delay time carried by the RPC instruction is 2 ms, and the delay time configured for this SPDK process is also 2 ms.

[0075] The threshold of the number of I / O requests configured for the above SPDK process is determined based on the maximum count value configured for the process and the percentage indicated by the latency type of the latency fault to be injected. The percentage indicated by any latency type is the percentage specified in the detection rule for detecting that latency type. Different latency types have different detection rules, and the detection rule for each latency type specifies not only the latency time but also the percentage. Here, within a specified time period, if the ratio of the number of I / O requests whose current latency time exceeds the latency time specified in the rule to the total number of I / O requests received by the SPDK process within the specified time period exceeds the percentage specified in the rule, it is confirmed that a latency fault of that latency type has occurred.

[0076] For example, the latency type of the latency fault to be injected is slow I / O. The detection rule corresponding to slow I / O specifies that within 60 s, the percentage of the number of I / O requests whose current latency time exceeds 2 ms to the total number of I / O requests received by the SPDK process within 60 s exceeds 80%. At this time, the latency time and 80% carried by the RPC instruction are used. The threshold of the number of I / O requests configured for the SPDK process is the product of 80% and 100 times (the maximum count value configured for the process is 100 times), that is, the threshold of the number of I / O requests is 80 times.

[0077] It should be noted that multiple SPDK processes run in parallel in the storage system. Each SPDK process has a corresponding input / output channel (I / O channel). When receiving an RPC instruction, each SPDK process receives the latency time and percentage through the corresponding I / O channel, and configures the latency time and the threshold of the number of I / O requests in the above manner.

[0078] Through the above configuration of the latency time and percentage, I / O latency faults of different latency types can be constructed according to the indication of the PCR instruction. In this way, through the indication of the latency time and percentage, the IO abnormal conditions in different real scenarios can be accurately simulated, so as to fully verify the reliability of the storage system.

[0079] The above has elaborated in detail on the configuration of the latency time and the threshold of the number of I / O requests for the SPDK process.

[0080] The following elaborates in detail on the judgment of putting the I / O request into the outstanding Tracker queue:

[0081] Please refer to Figure 3 , Figure 3 which is the flowchart for judging whether to put the I / O request into the outstanding Tracker queue provided by the embodiment of the present application. As Figure 3As shown, the process may include the following steps:

[0082] S301, when the configured delay time of the process is not the set value, determine whether the current count of the counter configured for the process is greater than the I / O request quantity threshold.

[0083] If the execution result of step S301 is no, then execute the following step S302; if the execution result of step S301 is yes, then execute the following step S303.

[0084] S302, determine to put the I / O request into the Outstanding Tracker queue.

[0085] S303, determine to send the I / O request to the submission queue.

[0086] In this embodiment, every time the SPDK process receives an I / O request, the corresponding current count of the process is incremented by a set value (such as 1), and when the current count of the process reaches the maximum count value corresponding to the process (here, the maximum count value is the upper limit of the configured counter, such as 100), it is reset to the initial value.

[0087] For example, suppose the I / O request quantity threshold is 80 times. When the SPDK process receives an I / O request, if the count of the I / O request is 66 and 66 is less than 80, then put the I / O request into the Outstanding Tracker queue. Specifically, save the I / O request and the time when the I / O request arrives at the SPDK process (submit time, that is, receive time) in the structure of the Outstanding Tracker. If the count of the I / O request is 95 and 95 is greater than 80, then directly send the I / O request to the submission queue for the NVMe SSD to obtain and process the I / O request from the submission queue.

[0088] By the above method, a part of the I / O requests are directly sent to the submission queue for processing, and another part of the I / O requests are delayed and submitted to the submission queue (that is, injecting an I / O delay fault), thereby constructing an I / O delay fault of the delay type corresponding to the percentage carried in the PCR instruction.

[0089] The above has elaborated in detail on judging to put the I / O request into the Outstanding Tracker queue.

[0090] The following elaborates in detail on determining to send the I / O request from the Outstanding Tracker queue to the submission queue:

[0091] Please refer to Figure 4 ,Figure 4 This is a flowchart for the embodiment of this application to issue the I / O request from the Outstanding Tracker queue to the submission queue. It should be noted that the following steps are performed for each unissued I / O request in the Outstanding Tracker queue at intervals of a set time period (i.e., at each polling time point):

[0092] S401, Obtain the sum of the latency time configured for the process and the reception time corresponding to the I / O request.

[0093] S402, Determine whether the obtained sum of times is greater than or equal to the polling time point.

[0094] If the execution result of the above step S402 is yes, then execute the following step S403; if the execution result of the above step S402 is no, then execute the following step S404.

[0095] S403, Determine that the issue time of the I / O request has arrived, and issue the I / O request from the Outstanding Tracker queue to the submission queue.

[0096] S404, Determine that the issue time of the I / O request has not arrived, and keep the I / O request in the Outstanding Tracker queue.

[0097] In the above manner, the delayed submission of the I / O request is achieved, so as to implement delayed processing based on the delayed submission, and construct the I / O latency fault based on the delayed processing.

[0098] The above has elaborated in detail the determination of issuing the I / O request from the Outstanding Tracker queue to the submission queue.

[0099] As an embodiment, after performing the above step S403, the method further includes: setting a tag for the I / O request in the Outstanding Tracker queue to indicate that it has been issued, so as to distinguish the I / O requests that have been issued and those that have not been issued in the Outstanding Tracker queue, and prevent duplicate submissions.

[0100] As an embodiment, the method further includes: when an I / O latency fault injection cancellation event is detected (for example, an explicit disable enable instruction), or it is determined that the latency time indicated by the received PCR is a set value (for example, 0), at this time, it indicates that the I / O latency fault injection function is cancelled. At this time, each outstanding I / O request in the Outstanding Tracker queue is issued to the submission queue. In the above manner, the switching between enabling and stopping I / O latency fault injection can be achieved, and enabling and stopping can be switched at any time according to actual needs, which greatly improves the flexibility and convenience of reliability testing.

[0101] The method provided by the embodiments of the present application has been described above. Next, the device provided by the embodiments of the present application will be described:

[0102] See Figure 5 , Figure 5 which is the structural diagram of the device provided by the embodiments of the present application. As Figure 5 shown, the device includes: a receiving module 501, a issuing module 502, and a sending module 503.

[0103] The receiving module 501 is configured to, when an input / output I / O request is received through the storage performance development kit process SPDK thread in the user space, put the I / O request into the outstanding task tracking Outstanding Tracker queue according to the latency time configured for the process and the I / O request quantity threshold;

[0104] The issuing module 502 is configured to determine the issuing time for issuing the I / O request from the Outstanding Tracker queue to a submission queue different from the Outstanding Tracker queue according to the latency time configured for the process and the receiving time of the received I / O request, and when the issuing time arrives, issue the I / O request from the Outstanding Tracker queue to the submission queue to process the I / O request in the submission queue and issue the I / O request from the request queue to the completed queue after processing;

[0105] The sending module 503 is configured to send the duration from when the I / O request is received to when it is issued to the completed queue through the process to the upper-layer service, so that the upper-layer service can obtain the latency fault of the executed I / O request based on the duration.

[0106] As an embodiment, the latency time is determined according to the latency time indicated by the latency type of the latency fault to be injected; the latency time indicated by any latency type is the latency time specified in the detection rule for detecting the latency type.

[0107] As an example, the I / O request quantity threshold is determined based on the maximum count value configured for the process and the percentage indicated by the latency type of the latency fault to be injected; the percentage indicated by any latency type is the percentage specified in the detection rule for detecting that latency type.

[0108] As an example, based on the latency time configured for the process and the I / O request quantity threshold, putting the I / O request into the Outstanding Tracker queue includes:

[0109] When the latency time configured for the process is not the set value, determining whether the current count of the counter configured for the process is greater than the I / O request quantity threshold;

[0110] If not, determining to put the I / O request into the Outstanding Tracker queue;

[0111] If so, determining to send the I / O request to the submission queue.

[0112] As an example, the receiving module is further configured to:

[0113] Each time the process receives an I / O request, the current count corresponding to the process is incremented by the set value, and when the current count of the process reaches the maximum count value corresponding to the process, it is reset to the initial value.

[0114] As an example, if determining the sending moment for sending the I / O request from the Outstanding Tracker queue to a submission queue different from the Outstanding Tracker queue based on the latency time configured for the process and the receiving time of the received I / O request includes:

[0115] Obtaining the sum of the latency time and the receiving time; wherein, if the sum of the times is greater than or equal to the current time, determining that the sending moment has arrived; otherwise, determining that the sending moment has not arrived.

[0116] As an example, the sending module is further configured to: after sending the I / O request from the Outstanding Tracker queue to the submission queue, set a tag for indicating that it has been sent for the I / O request in the Outstanding Tracker queue.

[0117] As an example, the apparatus further includes: an enabling module, configured to receive a Remote Procedure Call (RPC) instruction, and when the latency time indicated by the instruction is not the set value, enabling the I / O latency fault injection function of the process; after the I / O latency fault injection function is enabled, the process is configured with a latency time and an I / O request quantity threshold.

[0118] As an embodiment, the apparatus further includes an execution module, configured to, when detecting an I / O latency fault injection cancellation event or determining that the latency time indicated by the PCR is a set value, issue each outstanding I / O request in the Outstanding Tracker queue to the submission queue.

[0119] So far, Figure 5 the structural description of the shown apparatus is completed.

[0120] Refer to Figure 6 , Figure 6 which is the structural diagram of the electronic device provided by the embodiment of the present application. As Figure 6 shown, the hardware structure may include a processor and a machine-readable storage medium, where the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is configured to execute the machine-executable instructions to implement the method disclosed in the above examples of the present application.

[0121] Based on the same application concept as the above method, the embodiment of the present application further provides a machine-readable storage medium, on which a number of computer instructions are stored, and when the computer instructions are executed by a processor, the method disclosed in the above examples of the present application can be implemented.

[0122] Exemplarily, the above machine-readable storage medium may be any electronic, magnetic, optical or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, the machine-readable storage medium may be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid state drives, any type of storage disk (such as optical disks, DVDs, etc.), or similar storage media, or a combination thereof.

[0123] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. An I / O delay fault injection method, characterized in that, The method is applied to an electronic device, and the method includes: When an input / output I / O request is received by a storage performance development kit process SPDK thread in user mode, according to the configured latency time and the I / O request quantity threshold of the process, the I / O request is placed into an Outstanding Tracker queue; According to the configured latency time of the process and the reception time of the received I / O request, determine the dispatch time for dispatching the I / O request from the Outstanding Tracker queue to a submission queue different from the Outstanding Tracker queue, and when the dispatch time arrives, dispatch the I / O request from the Outstanding Tracker queue to the submission queue to process the I / O request in the submission queue and after processing, dispatch the I / O request from the request queue to the completed queue; Send the duration from when the I / O request is received to when it is dispatched to the completed queue by the process to the upper-layer service, so that the upper-layer service can obtain the latency fault of the executed I / O request based on the duration.

2. The method according to claim 1, wherein The latency time is determined according to the latency time indicated by the latency type of the latency fault to be injected; the latency time indicated by any latency type is the latency time specified in the detection rule for detecting this latency type.

3. The method according to claim 1, characterized in that, The I / O request quantity threshold is determined according to the configured maximum count value of the process and the percentage indicated by the latency type of the latency fault to be injected; The percentage indicated by any latency type is the percentage specified in the detection rule for detecting this latency type.

4. The method according to claim 1, wherein The step of placing the I / O request into the Outstanding Tracker queue according to the configured latency time of the process and the I / O request quantity threshold includes: When the configured latency time of the process is not a set value, determine whether the current count of the configured counter of the process is greater than the I / O request quantity threshold; If not, determine to place the I / O request into the Outstanding Tracker queue; If so, determine to dispatch the I / O request to the submission queue.

5. The method according to claim 1, wherein The method further includes: When the process receives each I / O request, the corresponding current count of the process is incremented by a set value, and when the current count of the process reaches the corresponding maximum count value of the process, it is reset to the initial value.

6. The method according to claim 1, characterized in that, The step of determining the dispatch time for dispatching the I / O request from the Outstanding Tracker queue to a submission queue different from the Outstanding Tracker queue according to the configured latency time of the process and the reception time of the received I / O request includes: Obtain the sum of the latency time and the reception time; wherein, if the sum is greater than or equal to the current time, determine that the dispatch time has arrived; otherwise, determine that the dispatch time has not arrived.

7. The method according to claim 6, characterized in that, After the I / O request is dispatched from the OutstandingTracker queue to the submission queue, the method further includes: Setting a tag for the I / O request in the Outstanding Tracker queue to indicate that it has been dispatched.

8. The method according to claim 1, characterized in that, Before this method, it also includes: Receiving a Remote Procedure Call (RPC) instruction, and when the delay time indicated by the instruction is not a set value and the percentage is not a set value, enabling the I / O latency fault injection function of the process; after the I / O latency fault injection function is enabled, the process is configured with a latency time and an I / O request quantity threshold.

9. The method according to claim 1, characterized in that The method also includes: When an I / O latency fault injection cancellation event is detected, or when it is determined that the delay time indicated by the received PCR is a set value, dispatching each undispatched I / O request in the Outstanding Tracker queue to the submission queue.

10. An I / O latency fault injection device, characterized in that, The device is applied to an electronic device, and the device includes: A receiving module, configured to, when receiving an input / output (I / O) request through a Storage Performance Development Kit (SPDK) thread in user mode, if it is determined, based on the latency time and the I / O request quantity threshold configured for the process, to place the I / O request in the Outstanding Tracker queue for uncompleted tasks, then place the I / O request in the Outstanding Tracker queue; A dispatching module, configured to determine the dispatch time for dispatching the I / O request from the Outstanding Tracker queue to a submission queue different from the Outstanding Tracker queue based on the latency time configured for the process and the reception time of the received I / O request, and when the dispatch time arrives, dispatch the I / O request from the Outstanding Tracker queue to the submission queue to process the I / O request in the submission queue and, after processing, dispatch the I / O request from the request queue to the completed queue; A sending module, configured to send, through the process, the duration from when the I / O request is received to when it is dispatched to the completed queue to an upper-layer service, so that the upper-layer service can obtain the latency fault of the I / O request being executed based on the duration.

11. The device according to claim 10, characterized in that, The latency time is determined based on the latency time indicated by the latency type of the latency fault to be injected; the latency time indicated by any latency type is the latency time specified in the detection rule for detecting that latency type; And / or The I / O request quantity threshold is determined based on the maximum count value configured for the process and the percentage indicated by the latency type of the latency fault to be injected; the percentage indicated by any latency type is the percentage specified in the detection rule for detecting that latency type; And / or The determination of placing the I / O request in the Outstanding Tracker queue based on the latency time and the I / O request quantity threshold configured for the process includes: When the configured delay time of the process is not the set value, determine whether the current count of the counter configured for the process is greater than the I / O request quantity threshold; If not, determine to put the I / O request into the Outstanding Tracker queue; If so, determine to send the I / O request to the submission queue; And / or, The receiving module is further configured to: Each time an I / O request is received by the process, the current count corresponding to the process is increased by a set value, and when the current count of the process reaches the maximum count value corresponding to the process, it is reset to the initial value; And / or, The determining of the sending moment for sending the I / O request from the Outstanding Tracker queue to a submission queue different from the Outstanding Tracker queue according to the configured delay time of the process and the receiving time of the received I / O request includes: Obtain the sum of the delay time and the receiving time; wherein, if the sum is greater than or equal to the current time, determine that the sending moment has arrived; otherwise, determine that the sending moment has not arrived; And / or, The sending module is further configured to: after sending the I / O request from the Outstanding Tracker queue to the submission queue, set a tag for indicating that it has been sent for the I / O request in the Outstanding Tracker queue; And / or, The device further includes: An enabling module, configured to receive a remote procedure call (RPC) instruction, and when the delay time indicated by the instruction is not the set value, enable the I / O delay fault injection function of the process; after the I / O delay fault injection function is enabled, the process is configured with a delay time and an I / O request quantity threshold; And / or, The device further includes: An execution module, configured to, when detecting an I / O delay fault injection cancellation event, or determining that the delay time indicated by the received PCR is the set value, send each un-sent I / O request in the Outstanding Tracker queue to the submission queue.

12. An electronic device, characterized in that, The electronic device includes: A processor; and A computer-readable storage medium, in which computer program instructions are stored, and when the computer program instructions are run by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 9.

13. A computer-readable storage medium, characterized in that, Computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are run by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 9.