Method, apparatus, device, and medium for stress testing NVMe hard disk dropout recovery

By performing Dcreboot stress test on the server and combining the self-recovery mechanism of NVME hard disk, the problem of not being able to wake up the NVME hard disk is solved, and the stable wake-up and data protection of the hard disk are achieved, reducing the risk of customer use.

CN116166489BActive Publication Date: 2025-05-27INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310175783.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-28
Publication Date
2025-05-27
Estimated Expiration
2043-02-28

AI Technical Summary

Technical Problem

The prior art cannot wake up the NVME hard drives that have been removed through external intervention, resulting in risk of customer use.

Method used

It provides a stress test method for recovery of NVME hard disk drives. By controlling the server to power on and start running, performing Dcreboot stress test, detecting whether the NVME hard disk is dropped, and awakening the NVME hard disk by obtaining hard disk location information, identifying the model, obtaining preset interactive modes, and waking up the hard disk self-recovery mechanism.

Benefits of technology

It realizes that the NVME hard disk that has been woken up by external intervention during the test. It is suitable for NVME hard disks that are completely powered or have no communication. It keeps the hard disk from being lost and prevents data loss, ensuring that the server runs stably in extreme environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116166489B_ABST
    Figure CN116166489B_ABST
Patent Text Reader

Abstract

The present application relates to a method, device, equipment and medium for stress testing the power-off recovery of an NVMe hard disk. The method includes: controlling a server equipped with an NVMe hard disk to power on and start running; after the server finishes booting, performing a Dcreboot stress test under the server system; detecting whether there is an NVMe hard disk power-off in the server; if there is an NVMe hard disk power-off, performing an operation to wake up the power-off NVMe hard disk; in response to being unable to wake up the power-off NVMe hard disk, determining whether to perform an NVMe hard disk initialization repair operation, and if so, performing an NVMe hard disk initialization repair operation on the power-off NVMe hard disk; in response to being unable to complete the NVMe hard disk initialization repair, returning to the step of performing an operation to wake up the power-off NVMe hard disk. This method is applicable to triggering the self-wake-up repair mechanism of the NVMe hard disk even when the NVMe hard disk is completely powered off or has no communication, so as to prevent data loss on the hard disk.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of servers, and in particular to a method, device, computer device, and storage medium for stress testing NVMe hard disk drop recovery. Background Art

[0002] AMD servers have been the servers with the most rapid growth in recent years. With their ultra-high performance and extremely high cost performance, they are highly favored by customers and consumers in the international market. For the previous ROME series CPUs and the current servers using the milan series CPU architecture, their high-speed processing performance has been consistently pursued by many customers. As the basic carrier of information servers, servers play an important role in the massive storage of information in particular; for information storage, hard disks, one of the basic components of servers, are required, and it is very important for servers to monitor and use hard disks; the monitoring operation of server hard disks is particularly important. For developers, in order to avoid the situation of runaway hard disk monitoring, it is necessary to use extreme usage methods such as DCreboot stress testing during the test to simulate whether there are abnormalities in a harsh usage environment. If there are abnormal monitoring problems, the firmware needs to take corresponding measures. Currently, the measure to wake up a dropped NVMe hard disk is to perform self-wake-up based on the control of the hard disk itself. This method is not applicable to NVMe hard disks that are completely powered off or have no communication, resulting in the inability to trigger the wake-up mechanism of NVMe hard disks and posing risks for customers. Summary of the Invention

[0003] Based on this, in view of the above technical problems, it is necessary to provide a method, device, computer device, and storage medium for stress testing NVMe hard disk drop recovery that can wake up a dropped NVMe hard disk through external intervention during the test, and solve the technical problem that the current measure cannot wake up a dropped NVMe hard disk through external intervention, resulting in risks for customers.

[0004] On the one hand, a method for stress testing NVMe hard disk drop recovery is provided. The method includes:

[0005] Control the server equipped with an NVMe hard disk to power on and start running.

[0006] In response to the completion of the server startup, perform DCreboot stress testing under the server system.

[0007] Detect whether there is an NVMe hard disk drop in the server; if there is an NVMe hard disk drop, perform an operation to wake up the dropped NVMe hard disk.

[0008] In response to the inability to wake up the powered-off NVMe hard disk, it is determined whether to perform an NVMe hard disk initialization repair operation. If so, the NVMe hard disk initialization repair operation is performed on the powered-off NVMe hard disk. Among them, the steps of waking up the NVMe hard disk for the powered-off NVMe hard disk include: obtaining the location information of the powered-off NVMe hard disk; identifying the model information of the powered-off NVMe disk according to the location information of the powered-off NVMe hard disk; obtaining a preset interaction method according to the model information of the powered-off NVMe disk; the preset interaction method includes a direct interaction method or a command interaction method; waking up the self-recovery mechanism of the powered-off NVMe disk according to the preset interaction method.

[0009] In response to the inability to complete the NVMe hard disk initialization repair, the steps of waking up the NVMe hard disk for the powered-off NVMe hard disk are returned.

[0010] In one embodiment, the step of obtaining the location information of the powered-off NVMe hard disk includes:

[0011] Detect whether the registers of all NVMe disks in the server can be read normally. If so, it is determined that the NVMe disk has normal interactive communication. If not, it is determined that the NVMe disk has abnormal interactive communication.

[0012] Obtain the running status in the interaction content of each NVMe disk that can have normal interactive communication, obtain the initial status before the Dcreboot stress test, compare the running status of each NVMe disk with the initial status. If the status is the same, it is determined that the NVMe disk is running normally. If the status is different, it is determined that the NVMe disk is running abnormally.

[0013] The firmware BIOS obtains the preset standard value of the register of each NVMe disk that is running normally, and obtains the real-time value of the register of each NVMe disk, and compares the real-time value of the register of each NVMe disk with its preset standard value. If they are the same, it is determined that the register value of the NVMe disk is normal. If they are different, it is determined that the register value of the NVMe disk is abnormal.

[0014] Define the NVMe disk with abnormal interactive communication or abnormal operation or abnormal register value as the powered-off NVMe hard disk, and obtain the location information of the powered-off NVMe hard disk according to the location information of all NVMe disks in the server.

[0015] In one embodiment, the step of waking up the self-recovery mechanism of the powered-off NVMe disk according to the preset interaction method includes:

[0016] In response to the existence of a powered-off NVMe hard disk, communicate with the register of the powered-off NVMe hard disk to write the corresponding preset standard value, and determine whether the powered-off NVMe hard disk is awakened.

[0017] In response to the inability to wake up the power-off NVMe hard disk, communicate with the register of the power-off NVMe hard disk again to write the corresponding preset standard value and determine whether the power-off NVMe hard disk is woken up.

[0018] In one embodiment, the step of determining whether to perform the NVMe hard disk initialization repair operation in response to the inability to wake up the power-off NVMe hard disk includes:

[0019] Obtain the execution times of recovering the power-off NVMe disk according to the preset interaction method;

[0020] Determine whether the power-off NVMe hard disk is woken up within the execution times threshold;

[0021] If the execution times of recovering the power-off NVMe disk is equal to the execution times threshold and the power-off NVMe hard disk is not woken up, it is determined that the NVMe hard disk initialization repair operation needs to be performed.

[0022] In one embodiment, the step of performing the NVMe hard disk initialization repair operation on the power-off NVMe hard disk includes:

[0023] In response to the inability to wake up the power-off NVMe hard disk, determine whether to initialize the power-off NVMe hard disk through the interaction of a complex programmable logic device (CPLD) and a baseboard management controller (BMC);

[0024] If so, complete the NVMe hard disk initialization repair operation on the power-off NVMe hard disk by means of power-on and power-off restart or power-on hold.

[0025] In one embodiment, after the step of detecting whether there is an NVMe hard disk power-off in the server, it further includes:

[0026] In response to the absence of an NVMe hard disk power-off, end;

[0027] In response to being able to wake up the power-off NVMe hard disk, end.

[0028] In one embodiment, after the step of detecting whether there is an NVMe hard disk power-off in the server, it further includes:

[0029] In response to not needing to perform the NVMe hard disk initialization repair operation, end;

[0030] In response to being able to complete the NVMe hard disk initialization repair, end.

[0031] On the other hand, a pressure test NVMe hard disk power-off recovery device is provided, and the device includes:

[0032] The power-on control module is used to control the server equipped with an NVME hard disk to power on and run.

[0033] The stress test module is used to perform a Dcreboot stress test under the server system in response to the completion of the server's power-on.

[0034] The NVME hard disk wake-up module is used to detect whether there is an NVME hard disk disconnection in the server; if there is an NVME hard disk disconnection, perform an NVME hard disk wake-up operation on the disconnected NVME hard disk.

[0035] The NVME hard disk initialization and repair module is used to, in response to the inability to wake up the disconnected NVME hard disk, determine whether to perform an NVME hard disk initialization and repair operation, and if so, perform an NVME hard disk initialization and repair operation on the disconnected NVME hard disk.

[0036] The multiple repair control module is used to, in response to the inability to complete the NVME hard disk initialization and repair, return to the step of performing an NVME hard disk wake-up operation on the disconnected NVME hard disk.

[0037] On the other hand, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:

[0038] Control the server equipped with an NVME hard disk to power on and run.

[0039] In response to the completion of the server's power-on, perform a Dcreboot stress test under the server system.

[0040] Detect whether there is an NVME hard disk disconnection in the server; if there is an NVME hard disk disconnection, perform an NVME hard disk wake-up operation on the disconnected NVME hard disk.

[0041] In response to the inability to wake up the disconnected NVME hard disk, determine whether to perform an NVME hard disk initialization and repair operation, and if so, perform an NVME hard disk initialization and repair operation on the disconnected NVME hard disk. Among them, the step of performing an NVME hard disk wake-up operation on the disconnected NVME hard disk includes: obtaining the location information of the disconnected NVME hard disk; identifying the model information of the disconnected NVME disk according to the location information of the disconnected NVME hard disk; obtaining a preset interaction method according to the model information of the disconnected NVME disk; the preset interaction method includes a direct interaction method or a command interaction method; wake up the self-recovery mechanism of the disconnected NVME disk according to the preset interaction method.

[0042] In response to the failure to complete the NVME hard disk initialization repair, return the operation steps for waking up the dropped NVME hard disk for the dropped NVME hard disk.

[0043] In another aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0044] Control the server equipped with an NVME hard disk to power on and run;

[0045] In response to the completion of the server startup, perform a Dcreboot stress test under the server system;

[0046] Detect whether there is an NVME hard disk drop in the server; if there is an NVME hard disk drop, perform an operation to wake up the dropped NVME hard disk;

[0047] In response to the failure to wake up the dropped NVME hard disk, determine whether to perform an NVME hard disk initialization repair operation. If so, perform an NVME hard disk initialization repair operation on the dropped NVME hard disk; wherein, the operation steps for waking up the dropped NVME hard disk for the dropped NVME hard disk include: obtaining the location information of the dropped NVME hard disk; identifying the model information of the dropped NVME disk according to the location information of the dropped NVME hard disk; obtaining a preset interaction method according to the model information of the dropped NVME disk; the preset interaction method includes a direct interaction method or a command interaction method; waking up the self-recovery mechanism of the dropped NVME disk according to the preset interaction method;

[0048] In response to the failure to complete the NVME hard disk initialization repair, return the operation steps for waking up the dropped NVME hard disk for the dropped NVME hard disk.

[0049] The above stress test NVME hard disk drop recovery method, device, computer device and storage medium wake up the dropped NVME hard disk through an external intervention method during the test, and execute the corresponding wake-up strategy according to the drop location. It is applicable to the situation where the NVME hard disk is completely powered off or has no communication, and can also trigger the self-wake-up repair mechanism of the NVME hard disk, so as to keep the hard disk from being lost and keep the data on the hard disk from being lost, so as not to affect the customer due to data loss. Through the efficient monitoring mechanism of the firmware BIOS, it can ensure that the server maintains an efficient and stable operating state when operating in an extreme usage environment, and helps the server service to obtain a better usage experience during use. Description of the Drawings

[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0051] Figure 1 It is a schematic flowchart of the method for recovering a dropped NVMe hard disk in a pressure test in one embodiment;

[0052] Figure 2 It is a schematic flowchart of the method for recovering a dropped NVMe hard disk in a pressure test in another embodiment;

[0053] Figure 3 It is a schematic flowchart of the operation steps for waking up a dropped NVMe hard disk in one embodiment;

[0054] Figure 4 It is a schematic flowchart of the steps for obtaining the location information of a dropped NVMe hard disk in one embodiment;

[0055] Figure 5 It is a schematic flowchart of the steps for waking up the self - recovery mechanism of a dropped NVMe disk according to a preset interaction method in one embodiment;

[0056] Figure 6 It is a schematic flowchart of the steps for, in one embodiment, in response to being unable to wake up a dropped NVMe hard disk, determining whether to perform an NVMe hard disk initialization repair operation, and if so, performing the NVMe hard disk initialization repair operation on the dropped NVMe hard disk;

[0057] Figure 7 It is a structural block diagram of a pressure test NVMe hard disk drop - recovery device in one embodiment;

[0058] Figure 8 It is an internal structure diagram of a computer device in one embodiment. Detailed implementation manners

[0059] In order to make the purpose, technical solutions and advantages of the present application clearer, the following further details the present application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0060] Embodiment 1

[0061] As described in the background art, the current measure to wake up the power-off NVMe hard disk is to perform self-waking based on the control of the hard disk itself. This method is not applicable to the NVMe hard disk that is completely powered off or has no communication, resulting in the failure to trigger the wake-up mechanism of the NVMe hard disk and posing risks to customer use.

[0062] To solve the above problems, in Embodiment 1 of the present invention, a method for stress testing the NVMe hard disk disk-off recovery is creatively proposed, which is a design method based on the AMD server BIOS to avoid abnormal disk-off of the NVMe hard disk during DCreboot stress testing. This method mainly solves the problem of abnormal disk-off of NVMe hard disks from different manufacturers on AMD milan series servers during the DCreboot stress testing of the server. The abnormal disk-off behavior is mainly manifested as the inability to recognize the NVMe hard disk or the problem of abnormal power-off and loss of the NVMe hard disk when the system performs a powercycle test or a reboot operation (lower reboot operation). For the abnormal loss behavior of the NVMe hard disk during the system reboot operation, and for the disk-off behaviors of NVMe hard disks from different manufacturers, there are also different manifestation forms, and corresponding countermeasures need to be made for the disk-off states in different forms. For the problem of disk-off, it is necessary to identify the model information of the NVMe disk according to the last disk-off position to determine which hard disks from different manufacturers have problems, and give protection and repair strategies for how to save the hard disk data without loss for different NVMe manufacturers; if the NVMe hard disk can adopt a soft repair strategy, the soft repair strategy can be tried, and if the NVMe hard disk requires a hard repair strategy, the hard repair strategy can be tried; while maintaining the strategy of not losing the hard disk, the data of the hard disk is also not lost to avoid affecting the customer due to data loss. By combining the two repair methods of soft repair and hard repair, the disk-off behavior of the NVMe hard disk on the server is guaranteed to occur.

[0063] Firstly, the implementation of the interaction scheme for the NVMe hard disk is carried out. For the specification design requirements of the NVMe hard disk, the design interaction is higher than that of general hard disks, in accordance with the NVMe hard disk interaction design specification. The firmware BIOS makes preparations in advance for the operations of the relevant registers for BIOS disk-off recognition. For the registers for recognizing the disk-off of the NVMe hard disk, at least two or more methods are reserved or made, directly operating on the register to read the value of the register. Judging whether the NVMe hard disk is initialized is also an operation method, and it is judged through the interaction between the CPLD and the firmware BMC.

[0064] Secondly, during the DCreboot process, the repair operation process for disk loss and disk disappearance. One is the soft repair process. The main method is to act and judge according to the previously designed interaction scheme to repair the lost NVME hard disk slot information, which can be considered to wake up the NVME hard disk recovery mechanism. The other is the hard disk hard repair operation. The hard repair operation generally completes the repair operation by interacting with the CPLD to perform a hard restart operation on the NVME hard disk, or by pulling the power to protect and repair the hard disk disk disappearance problem.

[0065] Finally, for the specific operation of identifying the problem of abnormal disk disappearance of the NVME hard disk. First, it is necessary to determine whether the disk loss behavior of the NVME hard disk occurs during the DCreboot process. In addition to the firmware BIOS's own judgment mechanism, it is also necessary to obtain the operation status command interface of the firmware BMC power-on and power-off to determine whether it is a DCreboot operation. After determining the operation status, it is necessary to judge whether it is an NVME hard disk, and at the same time, it is also necessary to judge NVME hard disks of different manufacturers in order to adopt different interactive repair methods for different manufacturers. For the identified loss of the NVME hard disk, first, the soft repair method is generally to directly operate on the recovery register of the NVME hard disk to wake it up through direct interaction or command interaction, etc., and try to repeatedly initialize the abnormal NVME hard disk. If it cannot be woken up after reaching the limited number of times, hard repair operation intervention is required, which generally completes the repair operation through power-on and power-off operations such as pulling the power and restarting / powering on and maintaining.

[0066] Combined with Figure 1 , the specific implementation design method is as follows:

[0067] 1. The AMD server equipped with the NVME hard disk is powered on and runs.

[0068] 2. After the server is powered on, perform the dcreboot operation under the server system.

[0069] 3. Check whether there is a problem of NVME disk loss.

[0070] 4. If there is a disk loss problem, first perform soft repair. After determining the operation status, judge whether it is an NVME hard disk, and at the same time, it is also necessary to judge NVME hard disks of different manufacturers in order to adopt different interactive repair methods for different manufacturers. For the identified loss of the NVME hard disk, first, the soft repair method is generally to directly operate on the recovery register of the NVME hard disk to wake it up through direct interaction or command interaction, etc., and try to repeatedly initialize the abnormal NVME hard disk. If it cannot be woken up after reaching the limited number of times, judge whether it can be woken up.

[0071] 5. If the soft repair fails to wake up, hard repair needs to be started. Generally, the NVME hard disk initialization repair operation is completed through power-on and power-off operations such as power-on and power-off restart / power-on hold.

[0072] 6. If the above repairs cannot be completed, the repair processes in steps 3 and 4 need to be repeated.

[0073] Through the efficient monitoring mechanism of the firmware BIOS, it can ensure that the AMD server maintains an efficient and stable operating state when operating in an extreme environment, and helps the AMD CPU platform server service obtain a better user experience during use.

[0074] Embodiment 2

[0075] This embodiment includes all the technical features of Embodiment 1. As Figure 2 shown, in Embodiment 2, a method for recovering from NVME hard disk drop during stress testing is provided, including the following steps:

[0076] Step S1, control the server equipped with the NVME hard disk to power on and start running;

[0077] Step S2, in response to the completion of the server startup, perform a Dcreboot stress test under the server system;

[0078] Step S3, detect whether there is an NVME hard disk drop in the server; if there is an NVME hard disk drop, perform an operation to wake up the dropped NVME hard disk;

[0079] Step S4, in response to the inability to wake up the dropped NVME hard disk, determine whether to perform an NVME hard disk initialization repair operation. If so, perform an NVME hard disk initialization repair operation on the dropped NVME hard disk;

[0080] Step S5, in response to the inability to complete the NVME hard disk initialization repair, return to the step of performing the operation to wake up the dropped NVME hard disk.

[0081] Among them, the DCreboot stress test can be understood to include a DC test and a reboot test.

[0082] As Figure 3 shown, in this embodiment, the step of performing an operation to wake up the dropped NVME hard disk includes:

[0083] Step S31, obtain the location information of the dropped NVME hard disk;

[0084] Step S32, identify the model information of the dropped NVME disk according to the location information of the dropped NVME hard disk;

[0085] Step S33: Obtain a preset interaction method according to the model information of the NVME disk that has fallen off; the preset interaction method includes a direct interaction method or a command interaction method.

[0086] Step S34: Wake up the self - recovery mechanism of the NVME disk that has fallen off according to the preset interaction method.

[0087] As Figure 4 shown, in this embodiment, the step of obtaining the location information of the NVME hard disk that has fallen off includes:

[0088] Step S311: Detect whether the registers of all NVME disks in the server can be read normally. If so, it is determined that the NVME disk has normal interactive communication; if not, it is determined that the NVME disk has abnormal interactive communication.

[0089] Step S312: Obtain the running state in the interaction content of each NVME disk that can have normal interactive communication, obtain the initial state before the Dcreboot stress test, compare the running state of each NVME disk with the initial state. If the states are the same, it is determined that the NVME disk is running normally; if the states are different, it is determined that the NVME disk is running abnormally.

[0090] Step S313: The firmware BIOS obtains the preset standard values of the registers of each NVME disk that is running normally, and obtains the real - time values of the registers of each NVME disk. Compare the real - time values of the registers of each NVME disk with their preset standard values. If they are the same, it is determined that the register values of the NVME disk are normal; if they are different, it is determined that the register values of the NVME disk are abnormal.

[0091] Step S314: Define the NVME disk with abnormal interactive communication, abnormal operation, or abnormal register values as the NVME hard disk that has fallen off, and obtain the location information of the NVME hard disk that has fallen off according to the location information of all NVME disks in the server.

[0092] As Figure 5 shown, in this embodiment, the step of waking up the self - recovery mechanism of the NVME disk that has fallen off according to the preset interaction method includes:

[0093] Step S341: In response to the existence of an NVME hard disk that has fallen off, communicate with the register of the NVME hard disk that has fallen off to write the corresponding preset standard value, and determine whether the NVME hard disk that has fallen off is awakened.

[0094] Step S342: In response to the inability to wake up the NVME hard disk that has fallen off, communicate with the register of the NVME hard disk that has fallen off again to write the corresponding preset standard value and determine whether the NVME hard disk that has fallen off is awakened.

[0095] AsFigure 6 As shown, in this embodiment, the step of determining whether to perform the NVME hard disk initialization repair operation in response to the inability to wake up the dropped NVME hard disk includes:

[0096] Step S41, obtaining the number of execution times for recovering the dropped NVME disk according to a preset interaction method;

[0097] Step S42, determining whether the dropped NVME hard disk is woken up within the execution times threshold;

[0098] Step S43, if when the number of execution times for recovering the dropped NVME disk is equal to the execution times threshold, and in response to the failure to wake up the dropped NVME hard disk, it is determined that the NVME hard disk initialization repair operation needs to be performed.

[0099] As Figure 6 shown, in this embodiment, the step of performing the NVME hard disk initialization repair operation on the dropped NVME hard disk includes:

[0100] Step S44, in response to the inability to wake up the dropped NVME hard disk, determining whether to initialize the dropped NVME hard disk through the interaction between the complex programmable logic device (CPLD) and the baseboard management controller (BMC);

[0101] Step S45, if so, completing the NVME hard disk initialization repair operation on the dropped NVME hard disk by means of power-on and power-off restart or power-on hold.

[0102] Please refer to Figure 1 , in this embodiment, after the step of detecting whether there is a dropped NVME hard disk in the server, it further includes:

[0103] In response to the absence of a dropped NVME hard disk, end;

[0104] In response to being able to wake up the dropped NVME hard disk, end;

[0105] In response to the case where the NVME hard disk initialization repair operation is not required, end;

[0106] In response to being able to complete the NVME hard disk initialization repair, end.

[0107] In the above method for recovering a dropped NVME hard disk during a stress test, during the test, the dropped NVME hard disk is awakened through an external intervention method, and the corresponding awakening strategy is executed according to the position where the hard disk drops. When it is applicable to an NVME hard disk that is completely powered off or has no communication, it can also trigger the self-awakening repair mechanism of the NVME hard disk, so as to keep the hard disk from being lost and at the same time keep the data on the hard disk from being lost, so as to avoid affecting the customer due to data loss. Through the efficient monitoring mechanism of the firmware BIOS, it can ensure that the server maintains an efficient and stable operating state when operating in an extreme usage environment, and helps the server service to obtain a better usage experience during the use process.

[0108] It should be understood that although Figures 2 - 6 the steps in the flowchart of Figures 2 - 6 are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,

[0109] In one embodiment, as Figure 7 shown, a stress test NVME hard disk drop recovery device 10 is provided, including: a power-on control module 1, a stress test module 2, a wake-up NVME hard disk module 3, an initialization repair NVME hard disk module 4, and a multiple repair control module 5.

[0110] The power-on control module 1 is used to control the server equipped with an NVME hard disk to power on and run.

[0111] The stress test module 2 is used to perform a Dcreboot stress test under the server system in response to the completion of the server startup.

[0112] The wake-up NVME hard disk module 3 is used to detect whether there is an NVME hard disk drop in the server; if there is an NVME hard disk drop, perform a wake-up NVME hard disk operation on the dropped NVME hard disk.

[0113] The initialization repair NVME hard disk module 4 is used to determine whether to perform an NVME hard disk initialization repair operation in response to the inability to wake up the dropped NVME hard disk. If so, perform an NVME hard disk initialization repair operation on the dropped NVME hard disk.

[0114] The multiple repair control module 5 is configured to return the operation steps for waking up the dropped NVME hard disk in response to the failure to complete the initialization repair of the NVME hard disk.

[0115] In this embodiment, the operation steps for waking up the dropped NVME hard disk include:

[0116] Obtain the location information of the dropped NVME hard disk;

[0117] Identify the model information of the dropped NVME disk according to the location information of the dropped NVME hard disk;

[0118] Obtain the preset interaction method according to the model information of the dropped NVME disk; the preset interaction method includes a direct interaction method or a command interaction method;

[0119] Wake up the self-recovery mechanism of the dropped NVME disk according to the preset interaction method.

[0120] In this embodiment, the step of obtaining the location information of the dropped NVME hard disk includes:

[0121] Detect whether the registers of all NVME disks in the server can be read normally. If so, it is determined that the NVME disk has normal interactive communication. Otherwise, it is determined that the NVME disk has abnormal interactive communication;

[0122] Obtain the running status in the interaction content of each NVME disk that can have normal interactive communication, obtain the initial status before the Dcreboot stress test, compare the running status of each NVME disk with the initial status. If the status is the same, it is determined that the NVME disk is running normally. If the status is different, it is determined that the NVME disk is running abnormally;

[0123] Obtain the preset standard value of the register of each NVME disk that is running normally, and obtain the real-time value of the register of each NVME disk. Compare the real-time value of the register of each NVME disk with its preset standard value. If they are the same, it is determined that the register value of the NVME disk is normal. If they are different, it is determined that the register value of the NVME disk is abnormal;

[0124] Define the NVME disk with abnormal interactive communication, abnormal operation, or abnormal register value as the dropped NVME hard disk, and obtain the location information of the dropped NVME hard disk according to the location information of all NVME disks in the server.

[0125] In this embodiment, the step of waking up the self-recovery mechanism of the dropped NVME disk according to the preset interaction method includes:

[0126] In response to the presence of a non - volatile memory express (NVME) hard disk with disk drop, communicate with the registers of the NVME hard disk with disk drop to write corresponding preset standard values, and determine whether the NVME hard disk with disk drop is awakened;

[0127] If the NVME hard disk with disk drop cannot be awakened, communicate with the registers of the NVME hard disk with disk drop again to write corresponding preset standard values and determine whether the NVME hard disk with disk drop is awakened.

[0128] In this embodiment, the step of determining whether to perform the NVME hard disk initialization repair operation in response to the inability to awaken the NVME hard disk with disk drop includes:

[0129] Obtain the number of execution times for recovering the NVME disk with disk drop according to a preset interaction method;

[0130] Determine whether the NVME hard disk with disk drop is awakened within the execution - time threshold;

[0131] If the number of execution times for recovering the NVME disk with disk drop is equal to the execution - time threshold and the NVME hard disk with disk drop is not awakened, it is determined that the NVME hard disk initialization repair operation needs to be performed.

[0132] In this embodiment, the step of performing the NVME hard disk initialization repair operation on the NVME hard disk with disk drop includes:

[0133] If the NVME hard disk with disk drop cannot be awakened, interact through a complex programmable logic device (CPLD) and a baseboard management controller (BMC) to determine whether to initialize the NVME hard disk with disk drop;

[0134] If so, complete the NVME hard disk initialization repair operation on the NVME hard disk with disk drop by means of power - on and power - off restart or power - on hold.

[0135] In this embodiment, after the step of detecting whether there is an NVME hard disk with disk drop in the detection server, it further includes:

[0136] If there is no NVME hard disk with disk drop, end;

[0137] If the NVME hard disk with disk drop can be awakened, end;

[0138] If the NVME hard disk initialization repair operation is not required, end;

[0139] If the NVME hard disk initialization repair can be completed, end.

[0140] In the above NVME hard disk drop recovery device for stress testing, during the testing process, the dropped NVME hard disk is awakened through an external intervention method, and the corresponding awakening strategy is executed according to the drop position. It is applicable to the situation where the NVME hard disk is completely powered off or has no communication, and can also trigger the self-awakening repair mechanism of the NVME hard disk, so as to keep the hard disk from being lost and the data on the hard disk from being lost, so as not to affect the customer due to data loss. Through the efficient monitoring mechanism of the firmware BIOS, it can ensure that the server maintains an efficient and stable operating state when operating in an extreme usage environment, and helps the server service to obtain a better usage experience during the usage process.

[0141] For the specific limitations of the NVME hard disk drop recovery device for stress testing, reference can be made to the limitations of the NVME hard disk drop recovery method for stress testing in the above text, which will not be elaborated here. Each module in the above NVME hard disk drop recovery device for stress testing can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in the processor in the computer device in hardware form or independent of the processor, or stored in the memory in the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above modules.

[0142] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 8 shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store NVME hard disk drop recovery data for stress testing. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes a method for recovering NVME hard disk drop during stress testing.

[0143] Those skilled in the art can understand that Figure 8 the structure shown in

[0144] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0145] Control the power-on and startup of a server equipped with an NVMe hard disk to run;

[0146] In response to the server's successful startup, perform a Dcreboot stress test under the server system;

[0147] Detect whether there is an NVMe hard disk disconnection in the server; if there is an NVMe hard disk disconnection, perform an operation to wake up the disconnected NVMe hard disk;

[0148] In response to the inability to wake up the disconnected NVMe hard disk, determine whether to perform an NVMe hard disk initialization repair operation. If so, perform an NVMe hard disk initialization repair operation on the disconnected NVMe hard disk; wherein, the steps of the operation to wake up the disconnected NVMe hard disk include: obtaining the location information of the disconnected NVMe hard disk; identifying the model information of the disconnected NVMe disk according to the location information of the disconnected NVMe hard disk; obtaining a preset interaction method according to the model information of the disconnected NVMe disk; the preset interaction method includes a direct interaction method or a command interaction method; waking up the self-recovery mechanism of the disconnected NVMe disk according to the preset interaction method;

[0149] In response to the inability to complete the NVMe hard disk initialization repair, return to the steps of the operation to wake up the disconnected NVMe hard disk.

[0150] For the specific limitations on the steps implemented when the processor executes the computer program, reference can be made to the limitations on the method for stress testing NVMe hard disk disconnection recovery in the above text, which will not be elaborated here.

[0151] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0152] Control the power-on and startup of a server equipped with an NVMe hard disk to run;

[0153] In response to the server's successful startup, perform a Dcreboot stress test under the server system;

[0154] Detect whether there is an NVMe hard disk disconnection in the server; if there is an NVMe hard disk disconnection, perform an operation to wake up the disconnected NVMe hard disk;

[0155] In response to the inability to wake up a powered-off NVMe hard disk, it is determined whether to perform an NVMe hard disk initialization repair operation. If so, the NVMe hard disk initialization repair operation is performed on the powered-off NVMe hard disk. Among them, the steps of waking up the powered-off NVMe hard disk include: obtaining the location information of the powered-off NVMe hard disk; identifying the model information of the powered-off NVMe disk according to the location information of the powered-off NVMe hard disk; obtaining a preset interaction method according to the model information of the powered-off NVMe disk; the preset interaction method includes a direct interaction method or a command interaction method; waking up the self-recovery mechanism of the powered-off NVMe disk according to the preset interaction method.

[0156] In response to the inability to complete the NVMe hard disk initialization repair, the steps of waking up the powered-off NVMe hard disk are returned.

[0157] For the specific limitations on the steps implemented when the computer program is executed by the processor, reference can be made to the limitations on the method for recovering from an NVMe hard disk power-off during stress testing in the above text, which will not be elaborated here.

[0158] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0159] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0160] The embodiments described above merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A method for recovering from NVMe hard disk disconnection during stress testing, characterized in that, it includes: Controlling the server equipped with an NVMe hard disk to power on and start running; In response to the completion of the server startup, perform a Dc reboot stress test under the server system; Detect whether there is an NVMe hard disk disconnection in the server; if there is an NVMe hard disk disconnection, perform an operation to wake up the disconnected NVMe hard disk; In response to the inability to wake up the disconnected NVMe hard disk, determine whether to perform an NVMe hard disk initialization repair operation. If so, perform an NVMe hard disk initialization repair operation on the disconnected NVMe hard disk; wherein, the step of performing an operation to wake up the disconnected NVMe hard disk includes: obtaining the location information of the disconnected NVMe hard disk; identifying the model information of the disconnected NVMe disk according to the location information of the disconnected NVMe hard disk; obtaining a preset interaction method according to the model information of the disconnected NVMe disk; the preset interaction method includes a direct interaction method or a command interaction method; waking up the self-recovery mechanism of the disconnected NVMe disk according to the preset interaction method; In response to the inability to complete the NVMe hard disk initialization repair, return to the step of performing an operation to wake up the disconnected NVMe hard disk; The step of obtaining the location information of the disconnected NVMe hard disk includes: Detect whether the registers of all NVMe disks in the server can be read normally. If so, determine that the NVMe disk has normal interaction communication. Otherwise, determine that the NVMe disk has abnormal interaction communication; Obtain the running status in the interaction content of each NVMe disk that can have normal interaction communication, obtain the initial status before the Dc reboot stress test, compare the running status of each NVMe disk with the initial status. If the status is the same, determine that the NVMe disk is running normally. If the status is different, determine that the NVMe disk is running abnormally; The firmware BIOS obtains the preset standard value of the register of each NVMe disk that is running normally, and obtains the real-time value of the register of each NVMe disk. Compare the real-time value of the register of each NVMe disk with its preset standard value. If they are the same, determine that the register value of the NVMe disk is normal. If they are different, determine that the register value of the NVMe disk is abnormal; Define the NVMe disk with abnormal interaction communication or abnormal running or abnormal register value as the disconnected NVMe hard disk, and obtain the location information of the disconnected NVMe hard disk according to the location information of all NVMe disks in the server.

2. The method for recovering from NVMe hard disk disconnection during stress testing according to claim 1, characterized in that, The step of waking up the self-recovery mechanism of the disconnected NVMe disk according to the preset interaction method includes: In response to the existence of a disconnected NVMe hard disk, communicate with the register of the disconnected NVMe hard disk and write the corresponding preset standard value, and determine whether the disconnected NVMe hard disk is awakened; In response to the inability to wake up the disconnected NVMe hard disk, communicate with the register of the disconnected NVMe hard disk again and write the corresponding preset standard value and determine whether the disconnected NVMe hard disk is awakened.

3. The method for recovering a dropped NVME hard disk during a stress test according to claim 2, characterized in that, the step of determining whether to perform an NVME hard disk initialization repair operation in response to the inability to wake up the dropped NVME hard disk includes: obtaining the number of execution times for recovering the dropped NVME disk according to a preset interaction method; judging whether the dropped NVME hard disk is woken up within the execution times threshold; if when the number of execution times for recovering the dropped NVME disk is equal to the execution times threshold, and in response to the failure to wake up the dropped NVME hard disk, it is determined that an NVME hard disk initialization repair operation needs to be performed.

4. The method for recovering a dropped NVME hard disk during a stress test according to claim 3, characterized in that, the step of performing an NVME hard disk initialization repair operation on the dropped NVME hard disk includes: in response to the inability to wake up the dropped NVME hard disk, determining whether to initialize the dropped NVME hard disk through the interaction of a complex programmable logic device and a baseboard management controller; if so, completing the NVME hard disk initialization repair operation on the dropped NVME hard disk by means of power-on and power-off restart or power-on hold.

5. The method for recovering a dropped NVME hard disk during a stress test according to claim 1, characterized in that, after the step of detecting whether there is a dropped NVME hard disk in the server, it further includes: in response to the absence of a dropped NVME hard disk, ending; in response to the ability to wake up the dropped NVME hard disk, ending.

6. The method for recovering a dropped NVME hard disk during a stress test according to claim 1, characterized in that, after the step of detecting whether there is a dropped NVME hard disk in the server, it further includes: in response to the need not to perform an NVME hard disk initialization repair operation, ending; in response to the ability to complete the NVME hard disk initialization repair, ending.

7. A stress test NVME hard disk drop recovery device for implementing the stress test NVME hard disk drop recovery method according to any one of claims 1-6, characterized in that, the device includes: a power-on control module for controlling the power-on and startup operation of a server equipped with an NVME hard disk; a stress test module for performing a Dc reboot stress test under the server system in response to the completion of the server startup; an NVME hard disk wake-up module for detecting whether there is a dropped NVME hard disk in the server; if there is a dropped NVME hard disk, performing an NVME hard disk wake-up operation on the dropped NVME hard disk; An initialization and repair module for NVMe hard drives is used to determine whether to perform an NVMe hard drive initialization and repair operation in response to an NVMe hard drive that cannot be woken up from a dropped disk state. If so, an NVMe hard drive initialization and repair operation is performed on the dropped disk NVMe hard drive. Among them, the steps of waking up the dropped disk NVMe hard drive operation include: obtaining the location information of the dropped disk NVMe hard drive; identifying the model information of the dropped disk NVMe hard drive according to the location information of the dropped disk NVMe hard drive; obtaining a preset interaction method according to the model information of the dropped disk NVMe hard drive; the preset interaction method includes a direct interaction method or a command interaction method; waking up the self-recovery mechanism of the dropped disk NVMe hard drive according to the preset interaction method. A multiple repair control module is used to return to the steps of waking up the NVMe hard drive operation on the dropped disk NVMe hard drive in response to the inability to complete the NVMe hard drive initialization and repair.

8. A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, when the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Low-power-consumption mode wakeup recovery method and device for solid state disk and computer equipment

    CN111625284A

  • Test method and device for quickly judging link rate of solid state disk and computer equipment

    CN113835944A