Debugging method, electronic equipment and readable storage medium
By enhancing the CQE structure in the NVMe protocol, carrying error details and combining it with logs, the problem of overly broad error status code classification in existing technologies is solved, enabling rapid location of error causes and improving debugging efficiency.
Patent Information
- Application Number
- CN202511028063.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-10-28
AI Technical Summary
The existing NVMe storage protocol has overly broad error status code classifications, making it difficult to quickly pinpoint the specific cause of errors during debugging, which is particularly inefficient in complex enterprise-level storage systems.
By enhancing the CQE structure in the NVMe protocol, adding invalid byte identifiers, error subtype codes, and context parameters to carry error details, and combining this with error logs, the host can quickly locate the cause of errors.
It improved debugging efficiency, reduced the time spent on manual analysis, accurately located the cause of errors, and shortened the debugging time.
Smart Images

Figure CN120849253A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer application technology, and in particular to debugging methods, electronic devices, and readable storage media. Background Technology
[0002] In NVMe (Non-Volatile Memory Express), error status codes can be carried with command completion. These error status codes can be used to pinpoint the cause of errors during debugging. However, the existing error status codes have relatively broad error classifications, making it difficult to directly locate the specific cause of the error. This forces developers to spend a significant amount of time troubleshooting each error individually, especially when debugging complex enterprise-level storage systems, resulting in low problem localization efficiency and further impacting debugging overall efficiency.
[0003] In conclusion, how to effectively solve problems such as storage debugging efficiency is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0004] The purpose of this application is to provide a debugging method, electronic device, and readable storage medium that optimizes the completion instructions so that they can carry not only basic error information but also detailed error information, enabling the host to quickly locate the specific cause of the error, eliminating the need for manual error analysis, improving the location efficiency, and thus accelerating the debugging process.
[0005] To solve the above-mentioned technical problems, this application provides the following technical solution:
[0006] A debugging method applied to a storage controller, comprising:
[0007] After receiving the instruction read notification sent by the host, it reads the commit instruction to be executed from the commit queue;
[0008] If an error is detected during the execution of the submission instruction, basic error information and error details are obtained.
[0009] The basic error information is added to the core control field of the completion instruction, and the error details are added to the error details field of the completion instruction in the form of invalid byte identifier, error subtype code and context parameter.
[0010] The completion instruction is placed in the completion queue, and a command read notification is sent to the host so that the host can locate the cause of the error by combining the basic error information and the error details information.
[0011] Preferably, in the case of detecting an error, it further includes:
[0012] Retrieve the original value of the error command field, the status of the relevant registers when the error was triggered, and detailed information about the error command;
[0013] The original value of the error command field, the state of the relevant registers when the error is triggered, and the detailed information of the error command are determined as error detail information;
[0014] The error details are recorded in an error log; the error log is available for the host to retrieve so that the host can locate the cause of the error based on the error details.
[0015] A debugging method, applied to a host computer, includes:
[0016] The commit command to be executed is placed in the commit queue, and a command read notification is sent to the storage controller;
[0017] After receiving the instruction read notification sent by the storage controller, the system reads the completion instruction corresponding to the submission instruction from the completion queue.
[0018] Basic error information is parsed from the core control field of the completion instruction, and error detail information is parsed from the error detail field of the completion instruction; the error detail information exists in the error detail field in the form of invalid byte identifier, error subtype code, and context parameter;
[0019] By combining the basic error information and the detailed error information, the cause of the error can be located, and if the target cause is located, the target cause can be processed to correct the error.
[0020] Preferably, if the target cause cannot be located by combining the basic error information and the detailed error information, the method further includes:
[0021] Execute a log retrieval command to obtain the error information logs recorded by the storage controller during the execution of the commit instruction;
[0022] The error details recorded in the error log are used to locate the cause of the error, and if the target cause is located, the target cause is processed to correct the error.
[0023] Preferably, locating the cause of the error using the error details recorded in the error information log includes:
[0024] Parse the error message log to obtain the error details; the error details include the original value of the error command field, the state of the relevant registers when the error was triggered, and detailed error command information.
[0025] The cause of the error is located by using the original value of the error command field, the state of the relevant registers when the error is triggered, and the detailed information of the error command.
[0026] Preferably, the cause of the error is located using the original value of the error command field, the state of the relevant registers when the error is triggered, and the detailed information of the error command, including:
[0027] By using the original value of the error command field corresponding to each error, the state of the relevant registers when the error is triggered, and the detailed information of the error command, the error causes corresponding to multiple different errors can be located respectively.
[0028] Preferably, locating the cause of the error by combining the basic error information and the detailed error information includes:
[0029] Execute a log retrieval command to obtain the error information log recorded by the storage controller during the execution of the commit instruction;
[0030] Obtain detailed error information from the error log;
[0031] By combining the basic error information, the detailed error information, and the specific error information, the cause of the error can be located.
[0032] Preferably, the error details information is parsed from the error details field of the completion instruction, including:
[0033] Read the invalid byte identifier, the error subtype code, and the context parameter from the error details field;
[0034] The read content is parsed to obtain the error details.
[0035] A debugging device, applied to a storage controller, comprising:
[0036] The instruction reading module is used to receive an instruction reading notification sent by the host and then read the submission instructions to be executed from the submission queue.
[0037] The instruction execution module is used to obtain basic error information and error details information if an error is detected during the execution of the submission instruction.
[0038] The instruction packaging module is used to add the basic error information to the core control field of the completion instruction, and to add the error details information to the error details field of the completion instruction in the form of invalid byte identifier, error subtype code and context parameter;
[0039] The instruction transmission module is used to put the completion instruction into the completion queue and send a command read notification to the host so that the host can locate the cause of the error by combining the basic error information and the error details information.
[0040] A debugging device, applied to a host computer, includes:
[0041] The instruction submission module is used to put the submission instructions to be executed into the submission queue and send an instruction read notification to the storage controller.
[0042] The instruction acquisition module is used to receive the instruction read notification sent by the storage controller and then read the completion instruction corresponding to the submission instruction from the completion queue.
[0043] The information extraction module is used to parse basic error information from the core control field of the completion instruction and to parse error detail information from the error detail field of the completion instruction; the error detail information exists in the error detail field in the form of invalid byte identifier, error subtype code, and context parameters;
[0044] The error correction module is used to locate the cause of the error by combining the basic error information and the detailed error information, and to process the target cause to correct the error if the target cause is located.
[0045] An electronic device, comprising:
[0046] Memory, used to store computer programs;
[0047] A processor is used to implement the steps of the above-described debugging method when executing the computer program.
[0048] A readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described debugging method.
[0049] Applying the method provided in this application embodiment, after receiving the instruction read notification sent by the host, the submit instruction to be executed is read from the submit queue; during the execution of the submit instruction, if an error is detected, basic error information and error details information are obtained; the basic error information is added to the core control field of the completion instruction, and the error details information is added to the error details field of the completion instruction in the form of invalid byte identifier, error subtype code, and context parameter; the completion instruction is placed in the completion queue, and a command read notification is sent to the host so that the host can locate the cause of the error by combining the basic error information and the error details information.
[0050] In this application, when the storage controller receives a read instruction notification from the host, it can read the commit instruction to be executed from the commit queue. During the execution of the commit instruction, error detection is performed. When an error is detected, not only basic error information but also detailed error information is obtained. The basic error information is then added to the core control field of the completion instruction, while the detailed error information is added to the error details field of the completion instruction in the form of an invalid byte identifier, error subtype code, and context parameters. The completion instruction is then placed in the completion queue, and a read instruction notification is sent to the host. In this way, the host can locate the specific cause of the error by combining the basic error information and the detailed error information. Because the completion instruction carries the detailed error information, the specific cause of the error can be effectively located based on this information, eliminating the need for manual analysis to pinpoint the error cause.
[0051] The technical effect of this application is that by optimizing the completion instruction, it can carry not only basic error information but also detailed error information, enabling the host to quickly locate the specific cause of the error, eliminating the need for manual error analysis, improving the location efficiency, and thus accelerating the debugging efficiency.
[0052] Accordingly, embodiments of this application also provide debugging devices, electronic devices, and readable storage media corresponding to the above-described debugging methods, which have the aforementioned technical effects, and will not be elaborated further here. Attached Figure Description
[0053] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0054] Figure 1 This is a flowchart illustrating a debugging method applicable to a storage controller, as described in an embodiment of this application.
[0055] Figure 2 This is a schematic diagram of an error log information in an embodiment of this application;
[0056] Figure 3 This is a flowchart illustrating a debugging method applicable to a host computer, as described in an embodiment of this application.
[0057] Figure 4 This is a schematic diagram illustrating an implementation of a debugging method in an embodiment of this application;
[0058] Figure 5 This is a schematic diagram of the structure of a debugging device according to an embodiment of this application;
[0059] Figure 6 This is a schematic diagram of another debugging device in an embodiment of this application;
[0060] Figure 7 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application;
[0061] Figure 8 This is a schematic diagram of the specific structure of an electronic device in an embodiment of this application. Detailed Implementation
[0062] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely some embodiments of the present application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0063] To facilitate understanding, the technical terms used in this article will be explained in detail below.
[0064] NVMe: Non-Volatile Memory Express, a storage protocol.
[0065] SQ: Submission Queue, the queue of submission commands in the NVMe protocol.
[0066] CQ: Completion Queue, the queue for completing instructions in the NVMe protocol.
[0067] SQE: Submission Queue Entry, a submission command in the NVMe protocol.
[0068] CQE: Completion Queue Entry, the completion command in the NVMe protocol.
[0069] PCIe: Peripheral Component Interconnect Express, a high-speed serial computer expansion bus standard.
[0070] Dword: Double Word represents four bytes, corresponding to the first four bytes of the SQE instruction.
[0071] Doorbell: Used to notify NVMe devices or hosts that new commands or data have been added to the queue.
[0072] PRP1 and PRP2 are two fields in the command data structure used to pass data buffer addresses between host memory and the NVMe controller.
[0073] Data Pointer: Includes PRP1 and PRP2, both of which are used to indicate the location of data in the host memory.
[0074] Controller: In NVMe technology, the controller is responsible for handling all NVMe protocol commands, including PCIe device enumeration and configuration for PCIe SSDs, NVMe controller identification and initialization, NVMe queue setup and initialization, etc.
[0075] The Get Log Page command is a control command in the NVMe protocol used to retrieve log page data from the controller. These log pages record critical information about the device.
[0076] The Get Feature command: A control command in the NVMe protocol used to read the current settings of a specific function on a controller or namespace.
[0077] The Set Feature command: An NVMe control command used to modify a specific feature attribute of a controller or namespace.
[0078] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating a debugging method applicable to a storage controller, as described in this application. The method includes the following steps:
[0079] S101. After receiving the instruction read notification sent by the host, read the commit instruction to be executed from the commit queue.
[0080] The host is the device that needs to communicate with the storage controller via NVMe so that it can use the storage devices controlled by the storage controller.
[0081] The Submission Queue (SQ) is a queue for NVMe protocol submission commands.
[0082] The Submission Queue Entry (SQE) is a command submitted by the host in the NVMe protocol that requires execution by the controller.
[0083] Once the controller receives a command read notification (such as Doorbell) from the host, it can read the SQE command from the SQ queue in the host memory via PCIe messages and execute the command internally.
[0084] S102. If an error is detected during the execution of the submission command, obtain the basic error information and error details.
[0085] During the process of the controller executing the submit command, it will check for errors. If an error is detected, it will obtain basic error information and error details.
[0086] In this embodiment, the submission command can be executed first, and then the execution process can be monitored to detect errors (such as command failure). Alternatively, the submission command can be checked first to detect errors within it (such as invalid or erroneous instructions or data). Error detection can be performed first, followed by monitoring of the execution process to obtain basic error information and detailed error information. The specific methods for obtaining basic information and detailed error information can be referenced in relation to the definitions and detection methods of different error contents, and will not be elaborated upon here.
[0087] Among them, the basic error information is the error-related information that the CQE in the NVMe protocol usually carries, such as invalid fields in the command, errors occurring during data transmission, etc.
[0088] Error details are more comprehensive than those included in a regular CQE and contain more specific error information. For example, error details may specify an invalid CID or an invalid NS ID in the SQE.
[0089] S103. Add the basic error information to the core control field of the completion instruction, and add the error details information to the error details field of the completion instruction in the form of invalid byte identifier, error subtype code and context parameter.
[0090] CQE is a fixed structure for the controller to report the command execution results to the host. It is 16 bytes long and is divided into four 32-bit double words (DW0~DW3) as Dword0, Dword1, Dword2 and Dword3 respectively.
[0091] Dword0: The core control field, which is associated with the command (SQE) in the submission queue through CID and informs the execution result through the Status Field. It is the core basis for the host to determine the status of the command.
[0092] Dword1: Supplementary information field. For example, after a read command is completed, this field records the actual number of bytes read, which helps the host verify data integrity.
[0093] Dword2 and Dword3: These are mainly reserved for protocol extensions, ensuring that the overall CQE structure does not need to be modified when adding new command types in the future.
[0094] In the NVMe protocol, the host sends commands to the controller via the Submit Queue (SQ), and the controller returns the command execution status (CQE) via the Completion Queue (CQ). However, the standard CQE error status codes (as shown in Table 1, such as Invalid Field in Command and Data Transfer Error) provide rather broad error classifications, failing to clearly identify the specific cause of the error. This forces developers to spend a significant amount of time troubleshooting each error individually, especially in complex enterprise-level storage systems, resulting in inefficient problem localization.
[0095] Table 1 shows examples of CQE error status codes and associated problems.
[0096]
[0097] Table 1 lists several CQE error status codes and related issues that cannot be clearly identified. This is for illustrative purposes only and does not cover all similar issues in the NVMe protocol. The following sections will use the "Invalid Field in Command" command as an example to illustrate the improved CQE.
[0098] In this embodiment, the CQE will be enhanced. The enhanced CQE structure is as follows: Dword 1 is the Reserved field. The CQE error information is further refined by adding an error details field to the CQE. The information is shown in Table 2 (the values marked in the header of Table 2 are the number of bits). An example of Dword 1 for the enhanced CQE of the Invalid Field in Command (0x02) command is shown in Table 3.
[0099] Table 2 shows a schematic diagram of the enhanced CQE structure.
[0100]
[0101] Among them, Invalid Byte Identifier: This marks the starting position of invalid bytes in the SQE. Sub Status Code: This refines the error based on the original error code. Context Parameters: This records the key parameters that caused the error (such as parameters in the SQE that led to the CQE error, out-of-bounds LBA ranges, invalid NS_IDs, etc.).
[0102] Table 3 is an example table of the Invalid Field in Command enhanced CQE Dword 1 table.
[0103]
[0104] That is, in this embodiment, the basic error information can be added to the core control field of the completion instruction according to the structure of enhanced CQE, and the error detail information can be added to the error detail field of the completion instruction in the form of invalid byte identifier, error subtype code and context parameter.
[0105] S104. Place the completion instruction into the completion queue and send a command read notification to the host so that the host can locate the cause of the error by combining the basic error information and the error details information.
[0106] Once the completion instruction is generated, it can be placed in the completion queue, and a command read notification can be sent to the host.
[0107] Once the host receives the command read notification, it can retrieve the completion command from the completion queue and locate the specific cause of the error by combining the basic error information and the error details.
[0108] Applying the method provided in this application embodiment, after receiving the instruction read notification sent by the host, the submit instruction to be executed is read from the submit queue; during the execution of the submit instruction, if an error is detected, basic error information and error details information are obtained; the basic error information is added to the core control field of the completion instruction, and the error details information is added to the error details field of the completion instruction in the form of invalid byte identifier, error subtype code, and context parameter; the completion instruction is placed in the completion queue, and a command read notification is sent to the host so that the host can locate the cause of the error by combining the basic error information and the error details information.
[0109] In this application, when the storage controller receives a read instruction notification from the host, it can read the commit instruction to be executed from the commit queue. During the execution of the commit instruction, error detection is performed. When an error is detected, not only basic error information but also detailed error information is obtained. The basic error information is then added to the core control field of the completion instruction, while the detailed error information is added to the error details field of the completion instruction in the form of an invalid byte identifier, error subtype code, and context parameters. The completion instruction is then placed in the completion queue, and a read instruction notification is sent to the host. In this way, the host can locate the specific cause of the error by combining the basic error information and the detailed error information. Because the completion instruction carries the detailed error information, the specific cause of the error can be effectively located based on this information, eliminating the need for manual analysis to pinpoint the error cause.
[0110] It should be noted that, based on the above embodiments, the embodiments of this application also provide corresponding improvement schemes. In the preferred / improved embodiments, the same or corresponding steps as in the above embodiments can be referred to each other, and the corresponding beneficial effects can also be referred to each other; however, these will not be elaborated upon in the preferred / improved embodiments herein.
[0111] In one specific embodiment of this application, the controller may further provide the host with detailed error information via logs to pinpoint the cause of the error. Specifically, when an error is detected, it may also include:
[0112] Retrieve the original value of the error command field, the status of the relevant registers when the error was triggered, and detailed information about the error command;
[0113] The original value of the error command field, the state of the relevant registers when the error was triggered, and the detailed information of the error command are determined as the error details;
[0114] Error details are recorded in the error log; the error log is available for the host to retrieve so that the host can locate the cause of the error based on the error details.
[0115] For ease of description, the above steps will be combined below.
[0116] Considering that in practical applications, CQE will only provide one type of error information, and due to the limitation of CQE's byte size, more error details cannot be carried, in this embodiment, when an error is detected, the controller can obtain error details including but not limited to the original value of the error command field (recording the SQE field value that caused the error), the status of the relevant registers when the error is triggered (such as SQ_Head, CQ_Tail, etc.), and the starting byte, field width, and field value of the error command.
[0117] These error details are recorded in an error log and made available to the host. This allows the host to pinpoint the specific cause of the error based on the error details.
[0118] Specifically, the host can retrieve error log information using the extended Get Log Page command. That is, when the host issues an SQE, the controller returns a CQE that only reports one error. If the SQE contains multiple errors that cannot be reported at once, to ensure that a modification command succeeds, the Get Log Page command can be used to check if there are any errors other than the reported errors.
[0119] like Figure 2 As shown (Byte in the figure represents bytes), when recording error logs, if there are multiple field errors, multiple errors can be recorded in sequence.
[0120] In one specific embodiment of this application, before debugging, the host and controller can determine the debugging mode. The debugging modes are coarse mode, precise mode, and in-depth mode. In coarse mode, error information is directly fed back to the host using the NVMe protocol. In precise mode, error information is fed back using an enhanced CQE structure. In in-depth mode, in addition to the precise mode, the controller also records detailed error information in the error log to help the host locate the error. In practical applications, coarse mode can be used in the initial debugging stage to narrow down the scope of precise debugging; then, precise mode is entered for precise debugging to further narrow down the scope; finally, in-depth mode is entered to perform in-depth error correction on the fewest complex errors. This maximizes debugging efficiency.
[0121] Please refer to Figure 3 , Figure 3 This is a flowchart illustrating a debugging method applicable to a host computer, as described in this application. The method includes the following steps:
[0122] S201. Place the commit instruction to be executed into the commit queue and send an instruction read notification to the storage controller.
[0123] The host places the commit instructions required for debugging into the commit queue and sends an instruction read notification to the controller.
[0124] After receiving the instruction read notification, the controller can retrieve the submission instruction from the submission queue and execute it. For details on the execution and processing procedures, please refer to [link / reference needed]. Figure 1 The method steps of the illustrated embodiment will not be described in detail here.
[0125] S202. After receiving the instruction read notification sent by the storage controller, read the completion instruction corresponding to the submission instruction from the completion queue.
[0126] Once the host receives the instruction read notification sent by the storage controller, it can read the completion instruction corresponding to the submission instruction from the completion queue.
[0127] It should be noted that the completion instruction can be specifically a completion instruction for the enhanced CQE result fed back by the controller. The specific structure of the completion instruction can be found in the description above.
[0128] S203. Parse the basic error information from the core control field of the completion instruction, and parse the error details information from the error details field of the completion instruction.
[0129] The error details information is presented in the error details field as an invalid byte identifier, an error subtype code, and context parameters.
[0130] Specifically, error details are parsed from the error details field of the completion instruction, including:
[0131] Read the invalid byte identifier, error subtype code, and context parameters from the error details field;
[0132] The read content is parsed to obtain error details.
[0133] The fields can be parsed one by one according to their definitions to obtain basic error information and detailed error information. The basic error information is shown in Table 1, and the detailed error information is shown in Table 3.
[0134] S204. Combine the basic error information and error details to locate the cause of the error, and if the target cause is located, handle the target cause to correct the error.
[0135] Basic error information indicates the basic circumstances of the error, while detailed error information indicates the specific details of the error. Therefore, combining the two can effectively pinpoint the cause of the error.
[0136] For example, if an error occurs due to an invalid CID, the basic error information will indicate the presence of the erroneous CID, and the detailed error information will specify the invalid CID. In this way, the cause of the error can be precisely determined as the invalid CID, and the error can be corrected by processing that CID.
[0137] Applying the method provided in this application embodiment, the submission instruction to be executed is placed in the submission queue, and an instruction read notification is sent to the storage controller; after receiving the instruction read notification sent by the storage controller, the completion instruction corresponding to the submission instruction is read from the completion queue; basic error information is parsed from the core control field of the completion instruction, and error detail information is parsed from the error detail field of the completion instruction; the error detail information exists in the error detail field in the form of invalid byte identifier, error subtype code, and context parameter; the error cause is located by combining the basic error information and the error detail information, and the target cause is processed to correct the error if the target cause is located.
[0138] In this application, when the storage controller receives a read instruction notification from the host, it can read the commit instruction to be executed from the commit queue. During the execution of the commit instruction, error detection is performed. When an error is detected, not only basic error information but also detailed error information is obtained. Then, the basic error information is added to the core control field of the completion instruction, and the detailed error information is added to the error details field of the completion instruction in the form of invalid byte identifier, error subtype code, and context parameters. Then, the completion instruction is placed in the completion queue, and a read instruction notification is sent to the host. In this way, the host can locate the specific cause of the error by combining the basic error information and the detailed error information. Since the completion instruction carries the detailed error information, the host can effectively locate the specific cause of the error based on this detailed error information, saving manual analysis and location of the error cause.
[0139] It should be noted that, based on the above embodiments, the embodiments of this application also provide corresponding improvement schemes. In the preferred / improved embodiments, the same or corresponding steps as in the above embodiments can be referred to each other, and the corresponding beneficial effects can also be referred to each other; however, these will not be elaborated upon in the preferred / improved embodiments herein.
[0140] In one specific embodiment of this application, when the target cause cannot be located by combining the basic error information and the detailed error information, the method further includes:
[0141] Execute the log retrieval command to obtain the error information logs recorded by the storage controller during the execution of the commit command;
[0142] Use the error details recorded in the error log to locate the cause of the error, and if the target cause is located, handle the target cause to correct the error.
[0143] This includes using the error details recorded in the error log to locate the cause of the error, including:
[0144] Parse the error log to obtain detailed error information, including the original value of the error command field, the state of the relevant registers when the error was triggered, and detailed information about the error command.
[0145] The cause of the error can be located by using the original value of the error command field, the state of the relevant registers when the error was triggered, and the detailed information of the error command.
[0146] Specifically, the cause of the error is located by utilizing the original value of the error command field, the state of the relevant registers at the time of the error triggering, and detailed information about the error command, including:
[0147] By using the original value of the error command field corresponding to each error, the state of the relevant registers when the error was triggered, and the detailed information of the error command, the error causes corresponding to multiple different errors can be located.
[0148] In other words, if the cause of the error cannot be located based on the CQE, the error information log recorded by the storage controller during the execution of the commit command can be obtained by executing the log retrieval command. By utilizing the detailed error information recorded in the error information log, the cause of the error can be located and corrected.
[0149] This error details information can record more specific error information than the basic error information and error details information, making it easier for the host to locate complex errors.
[0150] Specifically, when the host issues an SQE, the controller returns a CQE that only reports one error. If the SQE contains multiple errors that cannot be reported at once, to ensure that a modification command is successful, the Get Log Page command can be used to check if there are any errors other than the reported errors.
[0151] In one specific embodiment of this application, locating the cause of an error by combining basic error information and detailed error information includes:
[0152] Execute the log retrieval command to obtain the error information logs recorded during the execution of the commit instruction by the storage controller;
[0153] Obtain detailed error information from the error log;
[0154] By combining the basic error information, detailed error information, and error nuances, the cause of the error can be located.
[0155] In other words, when locating errors, it is also possible to obtain the basic error information, detailed error information, and error nuance information corresponding to the error in advance. The content and acquisition methods of this error information can be referred to the description above, and will not be repeated here.
[0156] Using multiple error messages to pinpoint the cause of an error can make the location more accurate and increase the probability of successful location.
[0157] For example, when the PRP2 of the Data Pointer field (specifying the location of data storage) in the SQE points to a linked list of pointers (addresses where several data items are stored), the CQE will typically only return an invalid field error, which is very vague. In this embodiment, the enhanced CQE will return a Data Pointer field error, but the error is still unclear because the Data Pointer contains pointers to several addresses where data is stored, and it is not yet certain which address pointer is faulty. At this point, using the Get Log Page command can obtain more detailed information, such as which specific address pointer is faulty, thus quickly locating the specific address pointer that has encountered the error.
[0158] In practical applications, for storage systems, during the debugging process, the host can apply, for example... Figure 2 The debugging method shown can be applied to the storage controller as follows: Figure 3 The debugging methods described above are used to complete the debugging of the storage system. The following example illustrates how to implement these debugging methods in a storage system.
[0159] As can be seen from the above, the embodiments provided in this application can implement a hierarchical error feedback mechanism, mainly including:
[0160] Base layer: Maintain traditional CQE error codes to be compatible with older protocols.
[0161] Enhancement layer: Provides detailed information through Reserved and extended logs, and the host detects whether the controller supports enhanced error reporting via the Get Feature command.
[0162] Please refer to Figure 4 , Figure 4 This is a schematic diagram illustrating an implementation of a debugging method provided in an embodiment of this application. The steps for implementing this debugging method in a storage system are as follows:
[0163] Step 1: After initialization, the host uses the Get Feature command (log retrieval command) to query whether the controller supports the enhanced error mechanism.
[0164] Step 2: The controller returns the query results.
[0165] The host and controller can agree in advance on the query results that support the enhanced error mechanism.
[0166] Step 3: If the controller does not support the enhanced error mechanism, then the interaction will proceed according to the NVMe protocol. If the controller supports the enhanced error mechanism, the host needs to enable it, which can be done using the Set Feature command (a management command in NVMe used to configure features).
[0167] Step 4: The controller returns the setting results.
[0168] Step 5: The host creates an SQE according to business requirements, submits commands to the SQ, and waits for the controller to return a CQE.
[0169] Step 6: If an error occurs during controller execution, the error will be detected and the error context will be captured, a CQE (including enhanced error information) will be generated, and the CQE will be returned to the host.
[0170] Step 7: After receiving the CQE, the host analyzes the cause of the error based on Status Field, Sub SC, Invalid Byte, and ContextParameters.
[0171] Step 8: If the host can analyze the cause of the error based on the information in the CQE, then correct the error and resubmit the SQE to the SQ.
[0172] Step 9: If the host cannot analyze the cause of the error based on the information in the CQE or needs to record the specific information when the error occurred, it can read the error details recorded by the controller using the Get Log Page command.
[0173] Step 10: The controller returns detailed error information.
[0174] Step 11: The host reads or records the obtained error details to analyze the cause of the error, corrects the error, and then resubmits the SQE to the SQ.
[0175] Step 12: The controller executes correctly and returns the correct result.
[0176] As can be seen, the debugging method provided in this application has the following technical effects:
[0177] Precise error location: By reusing the CQE Dword 1 field and adding Invalid Byte, Sub SC, and related Context Parameters, the cause of the error can be quickly located, greatly reducing debugging time.
[0178] Reduce maintenance costs: Combined with dynamic logs, it supports offline analysis of complex error scenarios (such as intermittent DMA errors).
[0179] Compatibility and flexibility: The layered design ensures compatibility with older protocol devices while allowing enhancements to be enabled on demand.
[0180] Corresponding to the above method embodiments, this application embodiment also provides a debugging device applicable to controllers. The debugging device described below and the debugging method described above can be referred to each other.
[0181] See Figure 5 As shown, the device includes the following modules:
[0182] The instruction reading module 501 is used to read the submission instruction to be executed from the submission queue after receiving the instruction reading notification sent by the host.
[0183] The instruction execution module 502 is used to obtain basic error information and error details if an error is detected during the execution of the submission instruction.
[0184] The instruction packaging module 503 is used to add basic error information to the core control field of the completion instruction, and to add error detail information to the error detail field of the completion instruction in the form of invalid byte identifier, error subtype code and context parameter;
[0185] The instruction transmission module 504 is used to put the completion instruction into the completion queue and send a command read notification to the host so that the host can locate the cause of the error by combining the basic error information and the error details information.
[0186] Corresponding to the above method embodiments, this application also provides a debugging device applicable to a host computer. The debugging device described below and the debugging method described above can be referred to in correspondence.
[0187] See Figure 6 As shown, the device includes the following modules:
[0188] The instruction submission module 601 is used to put the submission instructions to be executed into the submission queue and send an instruction read notification to the storage controller.
[0189] The instruction acquisition module 602 is used to receive the instruction read notification sent by the storage controller and then read the completion instruction corresponding to the submission instruction from the completion queue.
[0190] The information extraction module 603 is used to parse basic error information from the core control field of the completion instruction and to parse error detail information from the error detail field of the completion instruction; the error detail information exists in the error detail field in the form of invalid byte identifier, error subtype code and context parameter;
[0191] Error correction module 604 is used to locate the cause of the error by combining the basic error information and the error details information, and to process the target cause to correct the error if the target cause is located.
[0192] Corresponding to the above method embodiments, this application also provides an electronic device. The electronic device described below and the debugging method described above can be referred to in correspondence.
[0193] See Figure 7 As shown, the electronic device includes:
[0194] Memory 332 is used to store computer programs;
[0195] The processor 322 is used to implement the steps of the debugging method in the above method embodiment when executing a computer program.
[0196] This electronic device can be either a host or a storage controller. When the electronic device is a host, please refer to [reference needed]. Figure 8 , Figure 8 This is a schematic diagram of a specific structure of an electronic device provided in this embodiment. The electronic device can vary significantly due to differences in configuration or performance. It may include one or more central processing units (CPUs) (e.g., one or more processors) and a memory 332. The memory 332 stores one or more computer programs 342 or data 344. The memory 332 can be temporary or permanent storage. The program stored in the memory 332 may include one or more modules (not shown in the diagram), each module may include a series of instruction operations on the data processing device. Furthermore, the processor 322 may be configured to communicate with the memory 332 and execute the series of instruction operations stored in the memory 332 on the electronic device 301.
[0197] Electronic device 301 may also include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or one or more operating systems 341.
[0198] The steps in the debugging method described above can be implemented by the structure of the electronic device.
[0199] Corresponding to the above method embodiments, this application also provides a readable storage medium. The readable storage medium described below and the debugging method described above can be referred to in correspondence.
[0200] A readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the debugging method described in the above method embodiments.
[0201] The readable storage medium can specifically be a USB flash drive, external hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, or any other readable storage medium capable of storing program code.
[0202] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0203] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0204] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0205] Finally, it should be noted that in this document, relationships such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "include," "contain," or any other variations are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus.
[0206] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A debugging method, characterized in that, Applied to storage controllers, including: After receiving the instruction read notification sent by the host, it reads the commit instruction to be executed from the commit queue; If an error is detected during the execution of the submission instruction, basic error information and error details are obtained. The basic error information is added to the core control field of the completion instruction, and the error details are added to the error details field of the completion instruction in the form of invalid byte identifier, error subtype code and context parameter. The completion instruction is placed in the completion queue, and a command read notification is sent to the host so that the host can locate the cause of the error by combining the basic error information and the error details information.
2. The method according to claim 1, characterized in that, In cases where an error is detected, the following are also included: Retrieve the original value of the error command field, the status of the relevant registers when the error was triggered, and detailed information about the error command; The original value of the error command field, the state of the relevant registers when the error is triggered, and the detailed information of the error command are determined as error detail information; The error details are recorded in an error log; the error log is available for the host to retrieve so that the host can locate the cause of the error based on the error details.
3. A debugging method, characterized in that, Applied to the host, including: The commit command to be executed is placed in the commit queue, and a command read notification is sent to the storage controller; After receiving the instruction read notification sent by the storage controller, the system reads the completion instruction corresponding to the submission instruction from the completion queue. Basic error information is parsed from the core control field of the completion instruction, and error detail information is parsed from the error detail field of the completion instruction; the error detail information exists in the error detail field in the form of invalid byte identifier, error subtype code, and context parameter; By combining the basic error information and the detailed error information, the cause of the error can be located, and if the target cause is located, the target cause can be processed to correct the error.
4. The method according to claim 3, characterized in that, If the target cause cannot be located by combining the basic error information and the detailed error information, the following methods are also included: Execute a log retrieval command to obtain the error information logs recorded by the storage controller during the execution of the commit instruction; The error details recorded in the error log are used to locate the cause of the error, and if the target cause is located, the target cause is processed to correct the error.
5. The method according to claim 4, characterized in that, Locating the cause of the error using the error details recorded in the error log includes: Parse the error message log to obtain the error details; the error details include the original value of the error command field, the state of the relevant registers when the error was triggered, and detailed error command information. The cause of the error is located by using the original value of the error command field, the state of the relevant registers when the error is triggered, and the detailed information of the error command.
6. The method according to claim 5, characterized in that, Using the original value of the error command field, the state of the relevant registers when the error was triggered, and the detailed information of the error command, the cause of the error can be located, including: By using the original value of the error command field corresponding to each error, the state of the relevant registers when the error is triggered, and the detailed information of the error command, the error causes corresponding to multiple different errors can be located respectively.
7. The method according to claim 3, characterized in that, By combining the basic error information and the detailed error information, the cause of the error can be located, including: Execute a log retrieval command to obtain the error information log recorded by the storage controller during the execution of the commit instruction; Obtain detailed error information from the error log; By combining the basic error information, the detailed error information, and the specific error information, the cause of the error can be located.
8. The method according to any one of claims 3 to 7, characterized in that, Error details are parsed from the error details field of the completion instruction, including: Read the invalid byte identifier, the error subtype code, and the context parameter from the error details field; The read content is parsed to obtain the error details.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for implementing the steps of the debugging method as described in any one of claims 1 to 8 when executing the computer program.
10. A readable storage medium, characterized in that, The readable storage medium stores a computer program that, when executed by a processor, implements the steps of the debugging method as described in any one of claims 1 to 8.
Citation Information
Cited By
Storage device control method and electronic device
CN122337409A