Fault detection method, device and equipment, readable storage medium and program product

By simulating the read and write processing of storage devices and combining error prompts and read and write timeout detection, the accuracy and versatility issues of storage device fault detection are solved, comprehensive identification of hardware and software faults is achieved, and business continuity and stability are guaranteed.

CN120723508APending Publication Date: 2025-09-30TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410367661.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-27
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing technologies have low accuracy and poor versatility in storage device fault detection, making it difficult to fully identify hardware and software faults and limiting their use scenarios.

Method used

By simulating the read and write processing of the storage device, combining the error prompt information and the read and write timeout detection results, the fault detection results of the storage device are determined, including detecting error prompt information in the read and write processing or calculating the read and write timeout based on the current and recorded timestamps, and comprehensively judging the fault status of the storage device.

Benefits of technology

It improves the accuracy and versatility of storage device fault detection, can identify hardware and software faults, is applicable to various scenarios, and ensures business continuity and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723508A_ABST
    Figure CN120723508A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a fault detection method, device and equipment, a readable storage medium and a program product, which can be applied to the fields or scenes of cloud technology, artificial intelligence, vehicle-mounted and the like, and the method comprises the following steps: simulating execution of first read-write processing on a storage device; if error reporting prompt information is detected in the process of executing the first read-write processing, determining that a fault detection result of the storage device is a first fault detection result; the error prompt information is used for prompting that the storage device has input or output errors, and the first fault detection result indicates that the storage device has faults; if the error reporting prompt information is not detected in the process of executing the first read-write processing, determining a read-write timeout detection result of the first read-write processing according to the current timestamp and the recorded timestamp; and determining a fault detection result of the storage device according to the read-write timeout detection result. According to the embodiment of the invention, the accuracy and universality of fault detection of the storage device can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a fault detection method, a fault detection apparatus, a computer device, a computer-readable storage medium, and a computer program product. Background Art

[0002] Storage device failure is a common system failure. Once a storage device failure occurs, it will seriously affect the availability of the storage system, and further affect the upper-layer services and applications that rely on the storage system. Therefore, it is necessary to perform fault detection on the storage device.

[0003] Currently, fault detection is achieved by collecting and analyzing storage device operating status data and using neural network models to analyze and predict this data. However, this operating status data primarily reflects the health of the storage device at the hardware level, which means that only hardware faults within the storage device itself can be detected. This results in partial fault detection and low accuracy. Furthermore, these methods are limited in many scenarios, such as storage device model restrictions, resulting in limited versatility. Therefore, improving the accuracy and versatility of storage device fault detection is an urgent issue that needs to be addressed. Summary of the Invention

[0004] The present application provides a fault detection method, device and equipment, readable storage medium, and program product, which can improve the accuracy and versatility of storage device fault detection.

[0005] In one aspect, the present application provides a fault detection method, the method comprising:

[0006] simulating a first read / write process for the storage device;

[0007] If error prompt information is detected during the execution of the first read / write process, determining that the fault detection result of the storage device is a first fault detection result; the error prompt information is used to prompt that an input or output error exists in the storage device, and the first fault detection result indicates that a fault exists in the storage device;

[0008] If the error prompt information is not detected during the execution of the first read-write process, the read-write timeout detection result of the first read-write process is determined based on the current timestamp and the recorded timestamp; the current timestamp is the timestamp during the execution of the first read-write process, the recorded timestamp is the timestamp when the second read-write process is successfully executed, and the second read-write process is the read-write process that was successfully executed and did not time out before the first read-write process;

[0009] The failure detection result of the storage device is determined according to the read / write timeout detection result of the first read / write process; the read / write timeout detection result is used to indicate whether the first read / write process has timed out.

[0010] On the other hand, the present application provides a fault detection device, comprising:

[0011] A simulation execution module, configured to simulate execution of a first read / write process on a storage device;

[0012] a detection module configured to, if error prompt information is detected during the execution of the first read / write process, determine that a fault detection result of the storage device is a first fault detection result; the error prompt information is used to indicate that an input or output error exists in the storage device, and the first fault detection result indicates that a fault exists in the storage device;

[0013] The detection module is further configured to determine, if the error prompt information is not detected during the execution of the first read-write process, a read-write timeout detection result of the first read-write process based on a current timestamp and a recorded timestamp; the current timestamp is a timestamp during the execution of the first read-write process, the recorded timestamp is a timestamp when the second read-write process is successfully executed, and the second read-write process is a read-write process that was successfully executed and did not time out before the first read-write process;

[0014] The detection module is further configured to determine a fault detection result of the storage device based on a read / write timeout detection result of the first read / write process; the read / write timeout detection result is configured to indicate whether the first read / write process has timed out.

[0015] In one possible implementation, when the detection module is used to determine the read / write timeout detection result of the first read / write process according to the current timestamp and the timestamp of the record, it is specifically used to:

[0016] Determine the time difference between the current timestamp and the timestamp of the above record;

[0017] If the time difference is less than or equal to the time difference threshold, and the current timestamp is the timestamp when the first read-write process is successfully executed, then determining that the read-write timeout detection result of the first read-write process is a first read-write timeout detection result; the first read-write timeout detection result indicates that the first read-write process has not timed out;

[0018] If the time difference is greater than or equal to the time difference threshold, and the first read / write process is not successfully executed, the read / write timeout detection result of the first read / write process is determined to be a second read / write timeout detection result; the second read / write timeout detection result indicates that the first read / write process has timed out. In one possible implementation, when the detection module is used to determine the fault detection result of the storage device based on the read / write timeout detection result of the first read / write process, it is specifically used to:

[0019] If the read / write timeout detection result of the first read / write process indicates that the first read / write process has timed out, determining that the fault detection result of the storage device is the first fault detection result; the first read / write process timeout means that the first read / write process has not been successfully executed within the target time;

[0020] If the read-write timeout detection result of the above-mentioned first read-write process indicates that the above-mentioned first read-write process has not timed out, then the fault detection result of the above-mentioned storage device is determined to be the second fault detection result; the above-mentioned second fault detection result indicates that there is no fault in the above-mentioned storage device, and the above-mentioned first read-write process has not timed out means that the above-mentioned first read-write process is successfully executed within the above-mentioned target time.

[0021] In one possible implementation, the simulation execution module is further configured to:

[0022] Creating a simulated read / write program; the simulated read / write program is used to generate simulated read / write tasks according to a set time interval, where the set time interval is less than the time difference threshold;

[0023] The simulation execution module, when used to simulate the execution of the first read / write process on the storage device, is specifically used to:

[0024] According to the target simulated read / write task generated by the simulated read / write program, a first read / write process is executed on the storage device.

[0025] In one possible implementation, when the simulation execution module is used to simulate the execution of the first read / write process on the storage device, it is specifically configured to:

[0026] simulating a data read process for the first data to be executed on the storage device, and simulating a data write process for the second data to be executed on the storage device;

[0027] The above detection module is also used for:

[0028] Determine a read time difference threshold based on the data-related information of the first data, and determine a write time difference threshold based on the data-related information of the second data;

[0029] The time difference threshold is determined according to the read time difference threshold and the write time difference threshold.

[0030] In one possible implementation, the storage device is configured on a target service server among a plurality of service servers included in a server cluster, where the target service server is a master server or a slave server in the server cluster; wherein the fault detection device further includes a sending module, which is further configured to:

[0031] When the fault detection result of the storage device is the first fault detection result, sending the fault detection result to the management server of the server cluster, so that the management server performs disaster recovery processing on the target business server;

[0032] When the target business server is the master server, the disaster recovery process includes replacing the target business server; when the target business server is the slave server, the disaster recovery process includes deleting the target business server.

[0033] In a possible implementation, the sending module is further configured to:

[0034] If the fault detection result of the storage device is the first fault detection result, generating fault prompt information;

[0035] The fault prompt information is sent to the prompt object; the fault prompt information is used to prompt the prompt object that there is an input or output error in the storage device, or the fault prompt information is used to prompt the prompt object that the first read / write process has timed out.

[0036] Accordingly, the present application provides a computer device comprising: a processor, a storage device and a communication interface, wherein the processor, the communication interface and the storage device are interconnected, wherein the storage device stores a computer program, and the processor is used to call the computer program to implement the above-mentioned fault detection method.

[0037] Accordingly, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and the program instructions are executed by a processor to implement the fault detection method as described above.

[0038] Accordingly, the present application provides a computer program product, which includes a computer program. The computer program is executed by a processor to implement the above-mentioned fault detection method.

[0039] The embodiment of the present application simulates the execution of a first read-write process on a storage device to achieve early detection of potential storage device failures without affecting actual business, thereby ensuring business continuity and stability. If an error message indicating an input or output error is detected during the execution of the first read-write process, the fault detection result of the storage device is determined based on the read-write timeout detection result of the first read-write process. The above method combines the error message and the read-write timeout detection result to determine the fault detection result of the storage device, which can effectively identify hardware failures, software failures, etc. that occur in the storage device, making fault detection more comprehensive and applicable to various fault detection scenarios, thereby improving the accuracy and versatility of fault detection. In addition, the read-write timeout detection result is determined based on the current timestamp in the current read-write process and the timestamp corresponding to the last successful and non-timed read-write process of the current read-write process. This enables the read-write timeout detection result to reflect whether the storage device has continuously completed normal read and write operations within an acceptable time range, ensuring the accuracy of the read-write timeout detection, and effectively identifying problems of poor read and write performance such as slow disk reading and slow disk writing that may occur in the storage device. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0041] Figure 1 is a schematic diagram of the architecture of a detection system provided by an exemplary embodiment of the present application;

[0042] Figure 2 is a flowchart of a fault detection method provided by an exemplary embodiment of the present application;

[0043] Figure 3 is a flowchart of another fault detection method provided by an exemplary embodiment of the present application;

[0044] Figure 4A This is a schematic diagram of the architecture of a fault detection solution provided by an exemplary embodiment of the present application;

[0045] Figure 4B is a schematic diagram of a fault detection process provided by an exemplary embodiment of the present application;

[0046] Figure 5 is a structural diagram of a fault detection device provided by an exemplary embodiment of the present application;

[0047] Figure 6It is a structural diagram of a computer device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0048] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0049] The present application will be illustrated by the following examples.

[0050] See also Figure 1 , this figure is a schematic diagram of the architecture of a detection system provided by an exemplary embodiment of the present application. The detection system may specifically include a management server 101, a business server 102, and a fault detection device 103. The management server 101, the business server 102, and the fault detection device 103 are connected via a network, for example, via a local area network, a wide area network, a mobile Internet, etc. Figure 1 In the illustrated architecture, each service server 102 is configured with a corresponding fault detection device 103 , that is, there is a one-to-one correspondence between the service servers 102 and the fault detection devices 103 .

[0051] Each business server 102 includes a storage device, which may be a disk, a solid-state hard disk, an optical disk, a USB flash drive, a floppy disk, etc. Taking any fault detection device 103 as an example, the fault detection device 103 is connected to the corresponding business server 102 to perform fault detection (such as disk fault detection) on the corresponding business server 102. The management server 101 is connected to the fault detection device 103 to obtain the fault detection result obtained by the fault detection device 103. The management server 101 can perform device management work on the business server 102 based on the obtained fault detection result.

[0052] In one possible implementation, in a scenario where each business server 102 is configured with a corresponding fault detection device 103, the fault detection device 103 can be configured in the corresponding business server 102, such as the fault detection device 103 can be a detection device or module in the corresponding business server 102. In addition, the fault detection device 103 can also be configured outside the corresponding business server 102 and connected to the corresponding business server 102 through a network, such as the fault detection device 103 can be a terminal device or server independent of the corresponding business server 102.

[0053] In one possible implementation, in a scenario where only one fault detection device 103 is configured for multiple business servers 102, the fault detection device 103 can be configured in the management server 101, such as the fault detection device 103 can be a detection device or module in the management server 101. In addition, the fault detection device 103 can also be configured outside the management server 101 and multiple business servers 102. This fault detection device 103 is connected to the management server 101 and each business server 102 through a network, such as the fault detection device 103 can be a terminal device or server independent of the management server 101 and multiple business servers 102.

[0054] The above-mentioned terminal device is also referred to as a terminal, user equipment (UE), access terminal, subscriber unit, mobile device, user terminal, wireless communication device, user agent, or user device. The terminal device may be, but is not limited to, a smart home appliance, a handheld device with wireless communication capabilities (such as a smartphone or tablet), a computing device (such as a personal computer (PC)), an in-vehicle terminal, an aircraft, an intelligent voice interaction device, a wearable device, or other smart device.

[0055] The above-mentioned servers (such as management servers, business servers) can be independent physical servers, or they can be server clusters or distributed systems composed of multiple physical servers. They can also be cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), as well as big data and artificial intelligence platforms.

[0056] In one possible implementation, the fault detection device 103 can simulate executing a first read / write process on a storage device in the business server 102. If the fault detection device 103 detects an error message during the execution of the first read / write process, the fault detection result of the storage device is determined to be a first fault detection result. The error message is used to indicate that an input or output error has occurred in the storage device, and the first fault detection result indicates that the storage device has a fault. If the fault detection device 103 does not detect an error message during the execution of the first read / write process, the fault detection device 103 determines a read / write timeout detection result of the first read / write process based on a current timestamp and a recorded timestamp. The current timestamp is the timestamp during the execution of the first read / write process, and the recorded timestamp is the timestamp of the last successful read / write process that did not time out. The fault detection device 103 then determines a fault detection result of the storage device based on the read / write timeout detection result of the first read / write process. The read / write timeout detection result indicates whether the first read / write process has timed out. In addition, the fault detection device 103 can return the fault detection result to the management server 101 for subsequent analysis and use by the management server 101.

[0057] It is understood that the system architecture diagram described in the embodiment of the present application is for the purpose of more clearly illustrating the technical solution of the embodiment of the present application, and does not constitute a limitation on the technical solution provided by the embodiment of the present application. Figure 1 The number of management servers 101, business servers 102, and fault detection devices 103 is merely illustrative. Depending on the business implementation needs, any number of management servers 101, business servers 102, and fault detection devices 103 can be configured. Moreover, with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems. In subsequent embodiments, the management server will refer to the above-mentioned management server 101, the business server will refer to the above-mentioned business server 102, and the fault detection device will refer to the above-mentioned fault detection device 103, and they will not be repeated in subsequent embodiments.

[0058] See also Figure 2 , which is a flowchart of a fault detection method provided by an exemplary embodiment of the present application. Taking the method applied to a fault detection device as an example for explanation, the method may include the following steps:

[0059] S201 , simulating a first read / write process on a storage device.

[0060] In one possible implementation, the storage device may refer to a hardware device for storing data, such as a magnetic disk, a solid-state drive, an optical disk, a USB flash drive, etc. Of course, the storage device may also be a computer device.

[0061] In a possible implementation, executing the first read / write process on the storage device may refer to performing one or both of a data reading process and a data writing process on the storage device.

[0062] In an embodiment of the present application, the fault detection device simulates performing a first read and write process on the storage device. Here, "simulation" refers to simulating the normal read and write behavior of the storage device, such as by writing a specific program or software to imitate or simulate the process of the storage device in the computer system reading and / or writing data under normal working conditions. By simulating the read and write processing on the storage device, it is helpful to promptly discover and resolve storage device faults. Compared with troubleshooting and resolving storage device faults after they occur in an actual business environment, this can avoid business interruption or data loss caused by storage device faults in an actual business environment, and can discover and resolve potential storage device faults in advance without affecting actual business, thereby reducing the risks and losses caused by storage device failures.

[0063] S202: If error prompt information is detected during the execution of the first read / write process, determine that the fault detection result of the storage device is a first fault detection result.

[0064] In one possible implementation, the error prompt information is used to prompt the storage device that there is an input or output error (that is, the data cannot be read or written normally). For example, the error prompt information can be used to prompt the storage device that an error has occurred during one or both of the data reading and data writing processes. The error prompt information here can refer to an error message returned by the system or software to the fault detection device when an input or output error is detected in the storage device. The error prompt information can generally include one or more of the following information: a description of the error, the cause, and possible solutions to help users quickly locate and solve the problem.

[0065] For example, the error message may be an input / output (I / O) error message generated by the system. If the error message is detected, it indicates that an input or output error has actually occurred in the storage device. Therefore, judging that the storage device has failed based on the error message is accurate.

[0066] In one possible implementation, the first fault detection result is used to indicate that a fault exists in the storage device.

[0067] In an embodiment of the present application, the fault detection device detects an error prompt message during the execution of the first read and write processing, indicating that the storage device cannot read or write data normally. Then, it can be determined that the fault detection result of the storage device is the first fault detection result, that is, there is a fault in the storage device, thereby achieving rapid identification of the storage device fault. Since the error prompt message is generated when an input or output error actually occurs in the storage device, the above method ensures the accuracy of fault detection.

[0068] S203: If no error prompt information is detected during the execution of the first read / write process, determine a read / write timeout detection result of the first read / write process according to the current timestamp and the recorded timestamp.

[0069] In one possible implementation, the current timestamp is the timestamp during the execution of the first read-write process, the recorded timestamp is the timestamp when the second read-write process is successfully executed, and the second read-write process is the last read-write process of the first read-write process that was successfully executed and did not time out.

[0070] In one possible implementation, successful execution may mean that: no error message is detected during the execution of a read / write process (such as a first read / write process) on the storage device, and the data is successfully read and / or written. In other words, the following two conditions must be met to determine that a read / write process is successfully executed: during the execution of a read / write process on the storage device, the system does not detect input / output (I / O) error message and the data is successfully read and / or written.

[0071] In the embodiment of the present application, the current timestamp is a timestamp during the execution of the first read / write process. Since the first read / write process is a time period, the current timestamp here can be regarded as a timestamp dynamically selected within this time period.

[0072] The fault detection device can determine whether the first read / write process has timed out based on the current timestamp and the recorded timestamp, that is, based on the timestamp associated with the first read / write process and the timestamp when the second read / write process was successfully executed, and then obtain the read / write timeout detection result of the first read / write process, and then accurately obtain the fault detection result of the storage device based on the read / write timeout detection result. Because the read / write timeout detection result is determined based on the current timestamp during the current read / write process and the timestamp corresponding to the previous read / write process that was successfully executed and did not time out, the read / write timeout detection result can reflect whether the storage device has continuously completed normal read / write operations within an acceptable time range, ensuring the accuracy of read / write timeout detection and effectively identifying problems with poor read / write performance such as slow disk reading and writing that may occur in the storage device.

[0073] S204 : Determine a fault detection result of the storage device according to a read / write timeout detection result of the first read / write process.

[0074] In one possible implementation, read / write timeout detection is a fault detection method for determining the read / write performance of a storage device. It is used to detect whether a read / write process on the storage device has timed out. Generally speaking, if a read / write process times out, it indicates a fault with the storage device (e.g., low read / write performance, slow disk read / write speed, etc.). If a read / write process does not time out, it indicates that the storage device is functioning normally.

[0075] In one possible implementation, the fault detection device may perform a read / write timeout detection on the storage device. The read / write timeout detection can detect problems such as slow disk reading and writing in the storage device. The read / write timeout detection results are obtained through the read / write timeout detection, and the read / write timeout detection results can be used to indicate whether the first read / write process has timed out. It should be noted that the read / write timeout detection can be started at a preset start time, and the preset start time can be the start time of the first read / write process.

[0076] In the embodiment of the present application, the fault detection device does not detect an error message during the execution of the first read-write process, which means that there is no input or output error in the storage device. Then, the fault detection device determines the fault detection result of the storage device based on the read-write timeout detection result of the first read-write process. The above method combines the two dimensions of error message and read-write timeout detection result to determine the fault detection result of the storage device. It can effectively identify hardware faults, software faults, etc. that occur in the storage device, making fault detection more comprehensive, improving the accuracy of fault detection, reducing misjudgments and missed judgments, and the above fault detection method can separate fault detection from actual business, thereby ensuring the continuity and stability of the business. At the same time, the above method does not need to consider the storage device model, operating system, and hardware platform type, and is applicable to various fault detection scenarios with high versatility.

[0077] Based on the above embodiments, the beneficial effects of the present application are as follows: the embodiments of the present application simulate the execution of the first read-write process on the storage device, so as to realize the early detection of potential storage device failures without affecting the actual business, thereby ensuring the continuity and stability of the business. If an error message indicating an input or output error is detected during the execution of the first read-write process, the fault detection result of the storage device is determined to be the first fault detection result, and the first fault detection result indicates that there is a fault in the storage device; if no error message is detected during the execution of the first read-write process, the fault detection result of the storage device is further determined based on the read-write timeout detection result of the first read-write process. The above method combines the two dimensions of error message and read-write timeout detection result to determine the fault detection result of the storage device, and can effectively identify hardware faults, software faults, etc. that occur in the storage device, making the fault detection more comprehensive and improving the accuracy of the fault detection. Moreover, the above method is applicable to various fault detection scenarios, ensuring high versatility.

[0078] See also Figure 3 , which is a flow chart of another fault detection method provided by an exemplary embodiment of the present application, and is described by taking the method applied to a fault detection device as an example, the method may include the following steps:

[0079] S301 , simulating a first read / write process on a storage device.

[0080] In one possible implementation, in order to perform continuous fault detection on the storage device, the fault detection device may create a simulated read / write program; the simulated read / write program is used to generate simulated read / write tasks at set time intervals.

[0081] For example, the time interval can be flexibly set according to business conditions. For example, if the time interval is set to 1 minute, the simulated read and write program can generate a simulated read and write task every 1 minute. The time interval can also be set to 1 second, 30 seconds, 3 minutes, etc., which is not limited in the embodiments of the present application. The simulated read and write task can simulate the normal read and write behavior of the storage device and perform read and write processing on the storage device (such as performing a first read and write processing and a second read and write processing).

[0082] It should be noted that, since the embodiment of the present application determines whether the current read-write process has timed out by comparing the time difference between the current read-write process and the last read-write process that was successfully executed and did not time out, and the time difference threshold, it is necessary to ensure that each simulated read-write task generated by the simulated read-write program is the same, thereby ensuring the accuracy and reliability of fault detection. If the simulated read-write tasks generated each time are different, the time required to execute these tasks will also be different, which will directly affect the accuracy of timeout detection. For example, simulated read-write tasks with a large amount of data may cause the normal read-write processing to take longer, and this is not due to a decrease in storage device performance.

[0083] It should be noted that the set time interval needs to be smaller than the time difference threshold, such as setting the time interval to 1 minute and the time difference threshold to 3 minutes. This is because setting a smaller set time interval allows the storage device to be tested more frequently, which helps to detect potential problems in a timely manner. In actual operation, the system may cause fluctuations in the time required for read and write processing due to various reasons (such as system load, network latency, etc.), and setting a relatively large time difference threshold can provide a buffer for these normal fluctuations, avoid false alarms caused by short-term performance fluctuations, and ensure the accuracy of fault detection. Such a design helps to achieve continuous and stable fault detection of storage devices.

[0084] Based on this, the above-mentioned simulation of executing the first read / write process on the storage device (i.e., step S301) can be implemented as follows: executing the first read / write process on the storage device according to the target simulated read / write task generated by the simulated read / write program. The target simulated read / write task is one of the read / write tasks generated by the simulated read / write program, and the target simulated read / write task is used to execute the first read / write process on the storage device.

[0085] S302: Determine whether error prompt information is detected during the execution of the first read / write process on the storage device.

[0086] When the fault detection device detects error information during the first read / write process on the storage device, the fault detection device may perform the following steps S303:

[0087] S303: Determine that the fault detection result of the storage device is a first fault detection result.

[0088] The specific implementation of steps S302-S303 refers to the relevant description in the above embodiment and will not be repeated here.

[0089] When the fault detection device does not detect error prompt information during the first read / write process on the storage device, the fault detection device may perform the following steps S304-S308:

[0090] S304: Obtain the timestamp of the record.

[0091] In the embodiment of the present application, the recorded timestamp is the timestamp when the second read-write process is successfully executed, and the second read-write process is the last read-write process of the first read-write process that was successfully executed and did not time out.

[0092] S305: Determine a read / write timeout detection result of the first read / write process according to the current timestamp and the recorded timestamp.

[0093] In the embodiment of the present application, the current timestamp is a timestamp during the execution of the first read / write process. Since the first read / write process is a time period, the current timestamp here can be regarded as a timestamp dynamically selected within this time period.

[0094] The fault detection device can determine whether the first read / write process has timed out based on the current timestamp and the recorded timestamp, that is, based on the timestamp related to the first read / write process and the timestamp when the second read / write process is successfully executed, and then obtain the read / write timeout detection result of the first read / write process, so as to accurately detect whether the storage device has problems such as slow disk reading and slow disk writing, and then accurately obtain the fault detection result of the storage device based on the read / write timeout detection result.

[0095] In a possible implementation, step S305 may be implemented as follows:

[0096] (1) Determine the time difference between the current timestamp and the recorded timestamp.

[0097] (2) If the time difference is less than or equal to the time difference threshold, and the current timestamp is the timestamp when the first read-write process is successfully executed, the read-write timeout detection result of the first read-write process is determined to be the first read-write timeout detection result; the first read-write timeout detection result indicates that the first read-write process has not timed out.

[0098] In the embodiments of the present application, by setting a time difference threshold, a clear evaluation standard is set for determining whether the storage device has timed out, which helps to quickly determine whether the storage device meets performance requirements (for example, determining whether a read or write timeout has occurred in the storage device). Timeout detection is performed based on the time difference between the current timestamp and the recorded timestamp. It is possible to determine whether two normal read and write processes have been performed within a preset time (a normal read and write process is a read and write process that is successfully executed and does not time out), thereby obtaining a timeout detection result.

[0099] The time difference threshold can be flexibly set according to the business situation, for example, the time difference threshold can be set to 1 minute, 3 minutes, etc.

[0100] The following is a more detailed explanation of the current timestamp and the recorded timestamp:

[0101] For example, assuming that the time interval is set to 1 minute (i.e., the simulated read-write program generates a simulated read-write task every 1 minute), the time difference threshold is 3 minutes, the start time of the second read-write process is 10:0:0, and the recorded timestamp is 10:0:10, indicating that the second read-write process is successfully executed at 10:0:10, and the start time of the first read-write process is 10:01:0. The current timestamp is the timestamp during the execution of the first read-write process, which can refer to: the timestamp determined during the execution of the first read-write process according to the detection frequency interval. For example, if the detection frequency interval is 1 second, then the first current timestamp is 10:01:01, the second current timestamp is 10:01:02, and so on. It should be noted that the detection frequency interval can be flexibly set according to business conditions, such as 2 seconds, 1 minute, etc.

[0102] The fault detection device will perform timeout judgments in sequence starting from the first current timestamp. The specific method can be as follows: first, a timeout judgment is performed on the first current timestamp. Since the time difference between the first current timestamp and the recorded timestamp (the time difference is 1 second at this time) is less than or equal to the time difference threshold (3 minutes), it is determined whether the first read and write processing is successfully executed at the first current timestamp. If it is not successfully executed, then a timeout judgment is performed on the second current timestamp, and so on.

[0103] Assume that when performing the timeout judgment of the 10th current timestamp, since the time difference between the 10th current timestamp and the recorded timestamp (the time difference is 10 seconds at this time) satisfies the time difference threshold that is less than or equal to the time difference threshold, it is judged whether the first read and write processing is successfully executed at the 10th current timestamp. If the first read and write processing is successfully executed, the read and write timeout detection result of the first read and write processing is determined to be the first read and write timeout detection result.

[0104] (3) If the time difference is greater than or equal to the time difference threshold and the first read / write process is not successfully executed, the read / write timeout detection result of the first read / write process is determined to be the second read / write timeout detection result; the second read / write timeout detection result indicates that the first read / write process has timed out.

[0105] Exemplarily, based on the setting in the above step (2), it is assumed that when the timeout judgment of the 180th current timestamp is performed, since the time difference between the 180th current timestamp and the recorded timestamp satisfies or is greater than or equal to the time difference threshold (in this case, the time difference is equal to the time difference threshold, and the time difference threshold is 3 minutes), it is judged whether the first read and write processing is successfully executed at the 180th current timestamp. If the first read and write processing is not successfully executed, it means that the data reading and writing are not completed within the specified time (3 minutes), then the read and write timeout detection result of the first read and write processing is determined to be the second read and write timeout detection result.

[0106] It should be noted that the fault detection device may determine the read / write timeout detection result of the first read / write process as the second read / write timeout detection result if the time difference is greater than the time difference threshold and the first read / write process is unsuccessful. At the same time, it is necessary to ensure that the degree to which the time difference exceeds the time difference threshold is small, which can ensure the accuracy of timeout detection.

[0107] Exemplarily, based on the setting in the above step (2), it is assumed that when the timeout judgment of the 181st current timestamp is performed, since the time difference between the 181st current timestamp and the recorded timestamp satisfies the time difference threshold (in this case, the time difference is equal to the time difference threshold + 1 second, and the time difference threshold is 3 minutes), it is judged whether the first read and write processing is successfully executed at the 181st current timestamp. If the first read and write processing is not successfully executed, it means that the data reading and writing are not completed within the specified time. Thereafter, the timeout judgment of the current timestamp after the 181st current timestamp will no longer be performed.

[0108] In a possible implementation, the above step (3) can also be limited to: if the time difference is equal to the time difference threshold and the first read-write processing is not successfully executed, then the read-write timeout detection result of the first read-write processing is determined to be the second read-write timeout detection result, which will not be repeated here.

[0109] In the above steps (1)-(3), the fault detection device can calculate the time difference between the current timestamp and the recorded timestamp, and then compare the time difference with the time difference threshold, and determine the read / write timeout detection result of the first read / write process based on the comparison result and the execution status of the first read / write process, so as to ensure the accuracy of the read / write timeout detection.

[0110] In one possible implementation, the fault detection device may further perform the following steps: after simulating the successful execution of the first read / write process on the storage device without timing out, the current timestamp is stored. That is, after the first read / write process is successfully executed, the timestamp when the first read / write process is successfully executed is stored, so that when the storage device is subsequently fault-detected, the timestamp when the last read / write process that was successfully executed without timing out can be quickly and accurately obtained (such as obtaining the current timestamp of the first read / write process in a subsequent stage), thereby determining whether the storage device has a fault based on the obtained timestamp (such as the current timestamp). Exemplarily, the fault detection device may store the current timestamp in the memory when the first read / write process is successfully executed.

[0111] For example, since read / write timeout detection is continuous, each time a read / write timeout is detected, only the timestamp of the last successful read / write process that did not time out needs to be obtained. Therefore, the fault detection device can set a timestamp parameter to store the timestamp of the last successful read / write process that did not time out. After each successful read / write process that did not time out, the corresponding timestamp of the successful execution is updated to the timestamp parameter. This eliminates the need to store the timestamp of each successful read / write process that did not time out, and instead stores a single, real-time updated timestamp parameter, thereby reducing the amount of data stored and alleviating the storage burden.

[0112] S306: Determine whether the read / write timeout detection result of the first read / write process indicates that the first read / write process has timed out.

[0113] When the read / write timeout detection result of the first read / write process indicates that the first read / write process has timed out, the fault detection device may perform the following steps S307:

[0114] S307: Determine that the fault detection result of the storage device is a first fault detection result.

[0115] A first read / write process timeout refers to the failure of the first read / write process to execute successfully within the target time. The target time is associated with a time difference threshold. Specifically, an execution time range can be determined using the time difference threshold and the start time of the first read / write process. The end time of the execution time range is the start time plus the time difference threshold. Therefore, the target time can refer to the aforementioned execution time range.

[0116] When the read / write timeout detection result of the first read / write process indicates that the first read / write process has not timed out, the fault detection device may perform the following steps S308:

[0117] S308: Determine that the fault detection result of the storage device is a second fault detection result.

[0118] The second fault detection result indicates that there is no fault in the storage device, and the first read / write process not timing out means that the first read / write process is successfully executed within the target time.

[0119] In steps S306-S308, if the read / write timeout detection result indicates that the first read / write process has timed out, the storage device fault detection result is determined to be the first fault detection result (indicating that the storage device has a fault); if the read / write timeout detection result indicates that the first read / write process has not timed out, the storage device fault detection result is determined to be the second fault detection result (indicating that the storage device has not a fault). This method can achieve automated fault detection, reduce manual intervention, and improve the efficiency and accuracy of fault detection.

[0120] In one possible implementation, the fault detection device can determine the fault detection result of the storage device based on the read / write timeout detection result and the reference detection result. The specific method can be as follows:

[0121] (1) Determine a reference detection result of the storage device based on the auxiliary detection information.

[0122] In one possible implementation, the auxiliary detection information may include one or more of throughput indicator information, bandwidth indicator information, unit read / write number indicator information, and response latency indicator information. The fault detection device may determine a reference detection result of the storage device based on the auxiliary detection information.

[0123] In one possible implementation, the throughput indicator information can be used to indicate the amount of data read and write actually completed by the storage device within a certain period of time, which reflects the actual data processing capability of the storage device. The bandwidth indicator information can be used to indicate the maximum amount of data that the storage device can transmit per unit time, which represents the theoretical transmission capability of the storage device. The unit read and write times indicator information can refer to the number of input / output operations (IPOS), which indicates the number of read and write requests that the storage device can process per second. The response latency indicator information can be used to indicate the time required from initiating a read and write request to completing the request.

[0124] (2) Determine the fault detection result of the storage device based on the read / write timeout detection result and the reference detection result.

[0125] In one possible implementation, the fault detection device may determine that the fault detection result for the storage device is the second fault detection result (i.e., the storage device is not faulty) when the read / write timeout detection result indicates that the first read / write process has not timed out and the reference detection result indicates that no abnormality has occurred in the storage device. If one or both of the conditions of the read / write timeout detection result indicating that the first read / write process has not timed out and the reference detection result indicating that no abnormality has occurred in the storage device are not met, the fault detection result for the storage device may be determined to be the first fault detection result (i.e., the storage device is faulty).

[0126] In the above steps (1)-(2), the fault detection device determines the fault detection result of the storage device by combining the read and write timeout detection results and the reference detection results. This can further improve the accuracy of fault detection and reduce false alarms or missed alarms by comprehensively considering multiple performance indicators.

[0127] In one possible implementation, the first read / write process may include: a data read process and a data write process. Based on this, the implementation of simulating the execution of the first read / write process on the storage device may be as follows: simulating the execution of a data read process on the storage device regarding the first data, and simulating the execution of a data write process on the storage device regarding the second data.

[0128] In one possible implementation, the first data (or second data) may be text data, image data, video data, audio data, or fused data composed of one or more of the above data, which is not limited in this embodiment of the present application. For example, the first data (or second data) may be a timestamp, such as a timestamp corresponding to the current system time.

[0129] In one possible implementation, the first data and the second data may be the same or different. When the first data and the second data are the same, the read / write process may involve first writing the second data to the storage device and then reading the second data from the storage device. This eliminates the need to prepare different data and optimizes resource utilization. Furthermore, the fault detection device can determine whether the read second data matches the previously stored second data, thereby detecting issues such as data corruption or transmission errors.

[0130] Based on this, the time difference threshold can be determined as follows:

[0131] (1) Determine a read time difference threshold based on data-related information of the first data, and determine a write time difference threshold based on the second data.

[0132] In one possible implementation, the fault detection device can determine a reasonable read time difference threshold based on data-related information of the first data (such as data volume, data type, etc.). Similarly, the fault detection device can determine a reasonable write time difference threshold based on data-related information of the second data. For example, the larger the data volume of the first data, the larger the read time difference threshold, because large files require longer time to read and process, and the same applies to the write time difference threshold.

[0133] (2) Determine the time difference threshold based on the read time difference threshold and the write time difference threshold.

[0134] In a possible implementation, the fault detection device may add the read time difference threshold and the write time difference threshold to obtain the time difference threshold.

[0135] The above method combines the read time difference threshold and the write time difference threshold to determine a final time difference threshold. This final time difference threshold is used to determine whether the storage device's read and write processing has timed out. This provides a precise and comprehensive time difference threshold for storage device timeout detection, which helps improve the accuracy of timeout detection.

[0136] In one possible implementation, performing read and write processing on the storage device may only include reading the first data from the storage device. That is, only the data read function of the storage device is tested. In this case, the fault detection device may directly determine the read time difference threshold determined based on the data-related information of the first data as the final time difference threshold. This description will not be repeated here.

[0137] In one possible implementation, performing read and write processing on the storage device may only include writing the second data to the storage device. That is, only the data write function of the storage device is tested. In this case, the fault detection device may directly determine the write time difference threshold determined based on the data-related information of the second data as the final time difference threshold. This description is omitted here.

[0138] In one possible implementation, the storage device is configured on a target service server among multiple service servers included in a server cluster, where the target service server is a master server or a slave server in the server cluster. The master server can be responsible for management and coordination, while the slave server can be responsible for actual data storage and processing.

[0139] For example, a large-scale data center has a large number of storage nodes, each of which runs a different type of storage system, such as HBase, MySQL, ElasticSearch, HDFS, etc. A large number of storage nodes and management nodes together constitute a server cluster. Then, the business server in the embodiment of the present application may refer to a storage node, and the management server may refer to a management node. Storage device failure (such as disk failure) is the most frequent type of hardware failure. Once it occurs, it will seriously affect the availability of the storage system, and then affect the upper-layer services. Therefore, it is crucial to perform fault detection in a timely and accurate manner, and automatically remove and repair the faulty node.

[0140] Based on this, the fault detection device may further perform the following steps: when the storage device fault detection result is the first fault detection result, the fault detection result is sent to the management server of the server cluster, so that the management server performs disaster recovery processing on the target business server. Disaster recovery processing refers to a series of measures or strategies that ensure the system can continue to operate and minimize the impact of the storage device failure when a storage device failure occurs.

[0141] In a possible implementation, when the target service server is a primary server, the disaster recovery process includes replacing the target service server; when the target service server is a secondary server, the disaster recovery process includes deleting the target service server.

[0142] If the master server fails, the entire cluster may not respond properly, resulting in service interruption. Therefore, by promptly replacing the failed master server, cluster management and services can be ensured to continue uninterrupted, helping to maintain the stability and high availability of the entire system and prevent the failure from spreading to other nodes. There are usually multiple slave servers, and the failure of one slave server does not necessarily lead to an interruption of the entire service. Therefore, removing the failed target business server from the server cluster can reclaim resources such as disk space and network bandwidth, preventing the failed slave server from affecting the overall performance of the server cluster. By taking different disaster recovery measures based on the role of the target business server (master server or slave server), the fault detection device can ensure that the server cluster maintains efficient and stable operation in the face of node failures.

[0143] In a possible implementation, the fault detection device may further perform the following steps:

[0144] (1) If the fault detection result of the storage device is the first fault detection result, fault prompt information is generated.

[0145] (2) Send the fault prompt information to the prompt object.

[0146] In a possible implementation, the fault prompt information is used to prompt the prompt object that an input or output error occurs in the storage device, or the fault prompt information is used to prompt the prompt object that the first read / write process times out.

[0147] For example, if error prompt information is detected during the first read / write process, and the storage device fault detection result is determined to be the first fault detection result, the fault prompt information is used to notify the recipient of the prompt that an input or output error has occurred in the storage device. For example, the fault prompt information may read: "The storage device has failed due to an input or output error."

[0148] For example, if no error prompt information is detected during the execution of the first read / write process, and the read / write timeout detection result of the first read / write process indicates that the first read / write process has timed out, and the storage device fault detection result is determined to be the first fault detection result, then the fault prompt information is used to notify the prompt recipient that the first read / write process has timed out. For example, the fault prompt information may read: "The storage device has failed due to a timeout in the first read / write process."

[0149] In one possible implementation, the prompt object may refer to the administrator corresponding to the terminal device (or server) where the storage device is located, or the prompt object may refer to the client or terminal device of the administrator corresponding to the terminal device (or server) where the storage device is located. This allows the administrator to quickly locate the cause of the storage device failure based on the fault prompt information and perform fault repair, thereby improving the efficiency of fault repair. In addition, the prompt object may also refer to the terminal device (or server) where the storage device is located, such as the business server where the storage device is configured. This allows the business server to record the cause of the failure by storing the fault prompt information, so as to facilitate subsequent availability analysis and performance evaluation of the business server based on the fault prompt information stored in the business server.

[0150] In an embodiment of the present application, when the fault detection result of the storage device is the first fault detection result, that is, when there is a fault in the storage device, the fault detection device generates fault prompt information and sends the fault prompt information to the prompt object, so that the prompt object can promptly know the fault condition of the storage device, quickly and accurately locate the cause of the failure of the storage device, and can take corresponding measures according to the fault indication information, thereby reducing the impact of the fault on the business.

[0151] The following describes the application scenarios of the fault detection method provided in the embodiments of the present application:

[0152] The fault detection method provided in the embodiments of the present application can be regarded as a universal fault detection method for storage device faults and can be widely applied to various products and application scenarios requiring storage device fault detection, including but not limited to the following scenarios:

[0153] 1. Server Operation and Maintenance Management: The reliability and stability of server hard disk resources are crucial to ensuring normal business operations. This solution helps operation and maintenance managers promptly detect disk failures and take appropriate measures to repair or replace them, thereby improving server availability and stability.

[0154] 2. Cloud Storage Services: Cloud storage service providers are responsible for storing user data in highly secure data centers and must ensure data security and business continuity. Therefore, they need to quickly detect and repair storage device failures, particularly by regularly monitoring disk health. This solution can help cloud storage service providers implement automated and efficient disk failure detection mechanisms, improving service quality and user experience.

[0155] 3. Large Data Centers: Large data centers typically involve massive amounts of data storage and processing. A disk failure at a node can cause severe service interruption and data loss. This solution can be applied to the automated operation and maintenance of large data centers, enabling node disk failure detection, alarming, and automatic fault elimination, improving the stability and continuity of external services provided by large data centers.

[0156] This solution has the following advantages:

[0157] 1. Achieve cross-system and cross-platform fault detection: This solution can be applied to different operating systems and hardware platforms, and is independent of the storage device model, storage device operating data, system kernel version, etc., and is universal.

[0158] 2. Comprehensive Fault Detection: This solution can detect all types of storage device failures, including software-level and hardware-level failures such as slow disks and unresponsiveness. By making decisions based on simulated storage device read and write behavior, it can comprehensively detect various types of storage device failures with high detection rates and explainability.

[0159] 3. High flexibility: This solution can be deployed on any storage node to obtain the disk health information of the storage device of the storage node (the disk health information may include fault detection results and related information such as error prompts), and report the disk health information of the storage node to the management end, which will implement fault detection and processing capabilities.

[0160] The following examples illustrate the fault detection method provided in the embodiments of the present application:

[0161] See Figure 4A , Figure 4A This is a schematic diagram of the architecture of a fault detection solution provided in an embodiment of the present application. Figure 4A Take the configuration of the fault detection device on the server node (such as each business server) as an example for explanation. Each server node includes one or more storage devices. The embodiment of the present application performs fault detection on the above-mentioned one or more storage devices. First, the fault detection device simulates the read and write processing of the storage device through the storage device read and write module, and then can obtain the execution information of the above-mentioned read and write processing, such as the current timestamp, error prompt information, etc. The fault detection module performs fault detection processing on the storage device according to the execution information to obtain a fault detection result. Finally, the communication module generates the health status information of the storage device of the server node, and reports the health status information of the storage device to the management end (the management end is such as a management server). The health status information of the storage device may include one or more of the following: error prompt information, read and write timeout detection results, and fault detection results.

[0162] See Figure 4B , Figure 4B This is a schematic diagram of a fault detection process provided by an embodiment of the present application. The main process can be as follows:

[0163] 1. After the fault detection device is started, the storage device read and write module continuously and regularly performs read and write processing on the storage device to simulate normal storage device read and write behavior (such as Figure 4B , step S401 in the process.

[0164] 2. The fault detection device detects error prompt information during the reading and writing process (such as Figure 4B In step S402), when an error message (such as a system I / O error) is detected during the read / write process, the fault detection module will immediately detect a storage device failure (such as Figure 4B Otherwise, the timestamp of the last successful read / write process in the memory is updated by the storage device read / write module, such as obtaining the timestamp of the record (such as Figure 4B It should be noted that when a storage device fails due to hardware reasons such as slow disk reads, slow disk writes, or no response, the storage device read / write module may also become unresponsive due to I / O operation blockage. Therefore, detecting whether a read / write process has timed out is a necessary capability of the fault detection module.

[0165] 3. The fault detection device continues to perform read and write timeout detection to detect whether the read and write timeout (such as Figure 4BWhen a read / write timeout is detected, it is determined that a failure of the storage device has been detected (e.g., Figure 4B In step S403), when no read / write timeout is detected, it is determined that the storage device has failed (e.g., Figure 4B In other words, if the storage device read / write module does not detect an error message during the read / write process, and the fault detection module does not detect a read / write timeout, the read / write process can be considered successful, and the storage device can be considered healthy (no fault). Otherwise, a fault in the storage device is detected.

[0166] Assume that the fault detection module is configured with a read / write timeout threshold (e.g., a time difference threshold) of X, the current timestamp is T2, and the timestamp of the most recent successful read / write transaction in the disk read / write module's memory is T1. If T2-T1 is greater than or equal to X, the fault detection module may determine that the storage device has failed. Alternatively, the determination condition may be limited to: if T2-T1 is greater than X, the fault detection module may determine that the storage device has failed.

[0167] 4. The disk detection module generates storage device health information, and the communication module reports the storage device health information to the management terminal (such as Figure 4B Step S407 in the process) enables the management end to perform subsequent fault alarm, fault troubleshooting and other processing.

[0168] It should be noted that when simulating the read and write processing for the storage device, the embodiment of the present application can read and write data to a specified file (the data can be a random string), but other disk read and write tools or disk I / O testing tools can also be used to perform read and write processing, such as using fio, dd and other tools to achieve the same purpose.

[0169] It should be noted that the timeout determination threshold (such as the time difference threshold) in the embodiment of the present application can be flexibly adjusted according to actual conditions. For example, the storage systems used by different businesses have different time requirements. Therefore, it can be adjusted to 5 seconds, 60 seconds, 2 minutes, etc. according to the service level promised by the storage node to the outside world. The embodiment of the present application does not limit this.

[0170] The fault detection method proposed in the embodiment of the present application can be applied to the operation and maintenance management of massive storage nodes in large-scale data centers. From the perspective of operational data, the above solution has the following advantages:

[0171] 1. Improved fault detection accuracy and comprehensiveness. This solution supports detection of various types of disk failures, including software-level and hardware-level failures, such as slow disks and unresponsive disks, with high accuracy.

[0172] 2. Improved storage system reliability and stability. This solution can automatically and promptly detect disk failures and report the detection results to the management side in a timely manner. The management side then takes appropriate measures to repair or eliminate the fault, thereby improving system reliability and stability.

[0173] 3. Support cross-system and cross-platform disk fault detection. Large-scale data centers often have storage nodes equipped with different storage devices (e.g., different disk models), and their system kernel versions and server models vary. However, the fault detection method proposed in this solution is independent of disk model, disk operating data, or system kernel version, making it applicable to various fault detection scenarios and ensuring high versatility.

[0174] See also Figure 5 , which is a schematic diagram of the structure of a fault detection device provided in an embodiment of the present application. The fault detection device may specifically include:

[0175] A simulation execution module 501 is used to simulate the execution of a first read / write process on a storage device;

[0176] a detection module 502 configured to, if error prompt information is detected during the execution of the first read / write process, determine that a fault detection result of the storage device is a first fault detection result; the error prompt information is used to indicate that an input or output error exists in the storage device, and the first fault detection result indicates that a fault exists in the storage device;

[0177] The detection module 502 is further configured to, if the error prompt information is not detected during the execution of the first read / write process, determine a read / write timeout detection result of the first read / write process based on a current timestamp and a recorded timestamp; the current timestamp is a timestamp during the execution of the first read / write process; the recorded timestamp is a timestamp when the second read / write process is successfully executed; the second read / write process is the read / write process that was successfully executed and did not time out before the first read / write process;

[0178] The detection module 502 is further configured to determine a fault detection result of the storage device according to a read / write timeout detection result of the first read / write process; the read / write timeout detection result is used to indicate whether the first read / write process has timed out.

[0179] In one possible implementation, when the detection module 502 is used to determine the read / write timeout detection result of the first read / write process according to the current timestamp and the timestamp of the record, it is specifically used to:

[0180] Determine the time difference between the current timestamp and the timestamp of the above record;

[0181] If the time difference is less than or equal to the time difference threshold, and the current timestamp is the timestamp when the first read-write process is successfully executed, then determining that the read-write timeout detection result of the first read-write process is a first read-write timeout detection result; the first read-write timeout detection result indicates that the first read-write process has not timed out;

[0182] If the time difference is greater than or equal to the time difference threshold and the first read / write process is not successfully executed, the read / write timeout detection result of the first read / write process is determined to be the second read / write timeout detection result; the second read / write timeout detection result indicates that the first read / write process has timed out.

[0183] In one possible implementation, when the detection module 502 is used to determine the fault detection result of the storage device according to the read / write timeout detection result of the first read / write process, it is specifically used to:

[0184] If the read / write timeout detection result of the first read / write process indicates that the first read / write process has timed out, determining that the fault detection result of the storage device is the first fault detection result; the first read / write process timeout means that the first read / write process has not been successfully executed within the target time;

[0185] If the read-write timeout detection result of the above-mentioned first read-write process indicates that the above-mentioned first read-write process has not timed out, then the fault detection result of the above-mentioned storage device is determined to be the second fault detection result; the above-mentioned second fault detection result indicates that there is no fault in the above-mentioned storage device, and the above-mentioned first read-write process has not timed out means that the above-mentioned first read-write process is successfully executed within the above-mentioned target time.

[0186] In a possible implementation, the simulation execution module 501 is further configured to:

[0187] Creating a simulated read / write program; the simulated read / write program is used to generate simulated read / write tasks according to a set time interval, where the set time interval is less than the time difference threshold;

[0188] The simulation execution module 501 is specifically configured to:

[0189] According to the target simulated read / write task generated by the simulated read / write program, a first read / write process is executed on the storage device.

[0190] In one possible implementation, when the simulation execution module 501 is used to simulate the first read / write process performed on the storage device, it is specifically configured to:

[0191] simulating a data read process for the first data to be executed on the storage device, and simulating a data write process for the second data to be executed on the storage device;

[0192] The detection module 502 is further configured to:

[0193] Determine a read time difference threshold based on the data-related information of the first data, and determine a write time difference threshold based on the data-related information of the second data;

[0194] The time difference threshold is determined according to the read time difference threshold and the write time difference threshold.

[0195] In one possible implementation, the storage device is configured on a target business server among a plurality of business servers included in the server cluster, where the target business server is a master server or a slave server in the server cluster; wherein the fault detection device further includes a sending module 503, which is further configured to:

[0196] When the fault detection result of the storage device is the first fault detection result, sending the fault detection result to the management server of the server cluster, so that the management server performs disaster recovery processing on the target business server;

[0197] When the target business server is the master server, the disaster recovery process includes replacing the target business server; when the target business server is the slave server, the disaster recovery process includes deleting the target business server.

[0198] In a possible implementation, the sending module 503 is further configured to:

[0199] If the fault detection result of the storage device is the first fault detection result, generating fault prompt information;

[0200] The fault prompt information is sent to the prompt object; the fault prompt information is used to prompt the prompt object that there is an input or output error in the storage device, or the fault prompt information is used to prompt the prompt object that the first read / write process has timed out.

[0201] It should be noted that the functions of the various functional modules of the fault detection device in the embodiment of the present application can be specifically implemented according to the method in the above method embodiment. The specific implementation process can refer to the relevant description of the above method embodiment, which will not be repeated here.

[0202] See also Figure 6 , which is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 6The computer device in the embodiment shown may include: a processor 601, a storage device 602, and a communication interface 603. The processor 601, the storage device 602, and the communication interface 603 may exchange data.

[0203] The above-mentioned storage device 602 may include a volatile memory (volatile memory), such as a random-access memory (RAM); the storage device 602 may also include a non-volatile memory (non-volatile memory), such as a flash memory (flash memory), a solid-state drive (SSD), etc.; the above-mentioned storage device 602 may also include a combination of the above-mentioned types of memory.

[0204] The processor 601 may be a central processing unit (CPU). In one embodiment, the processor 601 may also be a graphics processing unit (GPU). The processor 601 may also be a combination of a CPU and a GPU. In one possible implementation, the storage device 602 is used to store a computer program, and the processor 601 may call the computer program to execute the steps of each method embodiment of the present application.

[0205] In a specific implementation, the processor 601, storage device 602 and communication interface 603 described in the embodiments of the present application can execute the aforementioned embodiments of the present application. Figure 2 or Figure 3 The implementation method described in the relevant embodiments of the provided fault detection method can also be implemented in the embodiments of this application. Figure 5 The implementation methods described in the relevant embodiments of the provided fault detection device will not be repeated here.

[0206] In the several embodiments provided in this application, it should be understood that the disclosed methods, devices and systems can be implemented in other ways. The device embodiments described above are merely schematic, and the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0207] In addition, it should be pointed out here that: the embodiment of the present application also provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program executed by the aforementioned fault detection device, and the computer program includes program instructions. When the processor executes the above-mentioned program instructions, it can execute the method in the aforementioned embodiment, so it will not be described in detail here. In addition, the description of the beneficial effects of adopting the same method will not be repeated. For technical details not disclosed in the computer-readable storage medium embodiment involved in this application, please refer to the description of the method embodiment of this application. As an example, the program instructions can be deployed on a computer device, or executed on multiple computer devices located at one location, or, executed on multiple computer devices distributed at multiple locations and interconnected by a communication network. Multiple computer devices distributed at multiple locations and interconnected by a communication network can constitute a blockchain system.

[0208] According to one aspect of the present application, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, so that the computer device can perform the method of the aforementioned embodiment. Therefore, a detailed description thereof will not be given here.

[0209] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0210] It can be understood that in the specific implementation of the present application, related data such as error prompt information, recorded timestamp, current timestamp, time difference threshold, first data, second data, etc., when the above embodiments of the present application are applied to specific products or technologies, the collection, use and processing of relevant data need to comply with relevant laws and standards of the relevant regions.

[0211] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0212] It should be noted that the terms "first" and "second" in the embodiments of this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a technical feature designated as "first" or "second" may explicitly or implicitly include at least one such feature.

[0213] The above disclosure is only part of the embodiments of the present application, and it is certainly not intended to limit the scope of the rights of the present application. A person skilled in the art can understand that all or part of the processes of the above embodiments and equivalent changes made in accordance with the claims of the present application are still within the scope of the invention.

Claims

1. A fault detection method, characterized in that: The method comprises: simulating a first read / write process for the storage device; If error prompt information is detected during the execution of the first read / write process, determining that the fault detection result of the storage device is a first fault detection result; the error prompt information is used to prompt that an input or output error exists in the storage device, and the first fault detection result indicates that the storage device has a fault; If the error prompt information is not detected during the execution of the first read-write process, determining the read-write timeout detection result of the first read-write process based on the current timestamp and the recorded timestamp; the current timestamp is the timestamp during the execution of the first read-write process, the recorded timestamp is the timestamp when the second read-write process is successfully executed, and the second read-write process is the read-write process that was successfully executed and did not time out before the first read-write process; The fault detection result of the storage device is determined according to the read / write timeout detection result of the first read / write process; the read / write timeout detection result is used to indicate whether the first read / write process has timed out.

2. The method according to claim 1, wherein The determining, according to the current timestamp and the recorded timestamp, a read / write timeout detection result of the first read / write process includes: determining a time difference between a current timestamp and a timestamp of the record; If the time difference is less than or equal to the time difference threshold, and the current timestamp is the timestamp when the first read-write process is successfully executed, determining that the read-write timeout detection result of the first read-write process is a first read-write timeout detection result; the first read-write timeout detection result indicates that the first read-write process has not timed out; If the time difference is greater than or equal to the time difference threshold and the first read-write process is not executed successfully, the read-write timeout detection result of the first read-write process is determined to be a second read-write timeout detection result; the second read-write timeout detection result indicates that the first read-write process has timed out.

3. The method according to claim 1 or 2, wherein: The determining the fault detection result of the storage device according to the read / write timeout detection result of the first read / write process includes: If the read / write timeout detection result of the first read / write process indicates that the first read / write process has timed out, determining that the fault detection result of the storage device is the first fault detection result; the first read / write process timeout means that the first read / write process has not been successfully executed within the target time; If the read-write timeout detection result of the first read-write process indicates that the first read-write process has not timed out, the fault detection result of the storage device is determined to be the second fault detection result; the second fault detection result indicates that there is no fault in the storage device, and the fact that the first read-write process has not timed out means that the first read-write process is successfully executed within the target time.

4. The method according to claim 2, wherein The method further comprises: Creating a simulated read / write program; the simulated read / write program is used to generate simulated read / write tasks according to a set time interval, and the set time interval is less than the time difference threshold; The simulation performs a first read / write process on the storage device, including: According to the target simulated read / write task generated by the simulated read / write program, a first read / write process is executed on the storage device.

5. The method according to claim 2, wherein The simulation performs a first read / write process on the storage device, including: simulating a data read process for a storage device regarding first data, and simulating a data write process for the storage device regarding second data; The method further comprises: determining a read time difference threshold based on the data-related information of the first data, and determining a write time difference threshold based on the data-related information of the second data; The time difference threshold is determined according to the read time difference threshold and the write time difference threshold.

6. The method according to any one of claims 1-2 or 4-5, characterized in that The storage device is configured on a target business server among a plurality of business servers included in a server cluster, and the target business server is a master server or a slave server in the server cluster; wherein the method further includes: When the fault detection result of the storage device is the first fault detection result, sending the fault detection result to the management server of the server cluster, so that the management server performs disaster recovery processing on the target business server; When the target business server is the primary server, the disaster recovery process includes replacing the target business server; when the target business server is the secondary server, the disaster recovery process includes deleting the target business server.

7. The method according to any one of claims 1-2 or 4-5, characterized in that The method further comprises: If the fault detection result of the storage device is the first fault detection result, generating fault prompt information; The fault prompt information is sent to a prompt object; the fault prompt information is used to prompt the prompt object that an input or output error occurs in the storage device, or the fault prompt information is used to prompt the prompt object that the first read / write process times out.

8. A fault detection device, characterized in that: The device comprises: A simulation execution module, configured to simulate execution of a first read / write process on a storage device; a detection module configured to, if error prompt information is detected during the execution of the first read / write process, determine that a fault detection result of the storage device is a first fault detection result; the error prompt information is used to prompt that an input or output error exists in the storage device, and the first fault detection result indicates that a fault exists in the storage device; The detection module is further configured to, if the error prompt information is not detected during the execution of the first read-write process, determine a read-write timeout detection result of the first read-write process based on a current timestamp and a recorded timestamp; the current timestamp is a timestamp during the execution of the first read-write process, the recorded timestamp is a timestamp when the second read-write process is successfully executed, and the second read-write process is the read-write process that was successfully executed and did not time out before the first read-write process; The detection module is further configured to determine a fault detection result of the storage device according to a read / write timeout detection result of the first read / write process; the read / write timeout detection result is configured to indicate whether the first read / write process has timed out.

9. A computer device, characterized in that: include: A processor, a storage device and a communication interface, wherein the processor, the communication interface and the storage device are connected to each other, wherein the storage device stores a computer program, and the processor is used to call the computer program to implement the fault detection method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the fault detection method according to any one of claims 1 to 7 is implemented.

11. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the computer program is used to implement the fault detection method according to any one of claims 1 to 7.