Fault detection method and server

By obtaining the storage type of the virtual machine, simulating access to the target volume and judging the read and write rate, the problem of low virtual machine failure detection efficiency in the virtualization platform is solved, fast and accurate fault detection is achieved, and the stability of virtual machine business is ensured.

WO2025156737A1PCT designated stage expired Publication Date: 2025-07-31XFUSION DIGITAL TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/126739
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-22
Filing Date
2024-10-23
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

In virtualization platforms, it is difficult for the existing technology to efficiently detect virtual machine failures, resulting in system performance degradation or data loss, and the existing detection mechanism is complex and time is slow.

Method used

By obtaining the storage type of the virtual machine accessing the storage data volume, simulating the process of the virtual machine accessing the target volume, judging the access status information, and determining the fault detection result using the read rate and write rate, avoiding polling on each virtual machine, and improving detection efficiency.

Benefits of technology

It realizes fast and accurate virtual machine failure detection, reduces detection time, improves the efficiency and accuracy of virtual machine failure detection, and ensures the stable operation of virtual machine services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024126739_31072025_PF_FP_ABST
    Figure CN2024126739_31072025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to the technical field of servers. Disclosed are a fault detection method and a server, used for improving the efficiency of virtual machine fault detection. The method is applied to a server; virtual machines run on the server. The method comprises: acquiring a storage type for a virtual machine to access a storage data volume; determining a target volume on the basis of the storage type, wherein the target volume is another storage data volume that supports access by the virtual machine using the storage type; performing an access operation on the target volume to obtain access state information; and on the basis of the access state information, determining a fault detection result, wherein the fault detection result is used for indicating whether a fault occurs in the access of the virtual machine to the storage data volume.
Need to check novelty before this filing date? Find Prior Art

Description

Fault detection method and server

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on January 22, 2024, with application number 202410090204.2 and application name “A Fault Detection Method and Server”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The embodiments of the present application relate to the field of server technology, and in particular to a fault detection method and a server. Background Art

[0003] Virtualization platforms improve resource utilization and management efficiency by abstracting physical storage resources into virtual, logical storage resources and providing virtual machines with data volumes. However, using storage systems in virtualization platforms can introduce various failure risks, preventing virtual machines from accessing data, leading to system performance degradation or data loss, and ultimately impacting virtual machine services. Therefore, timely detection of virtual machine failures is crucial to ensuring stable operation of virtualization platforms.

[0004] In related technologies, whether a current virtual machine has a fault can only be determined by obtaining the network status. The implementation mechanism is relatively complex and the detection time is slow, resulting in low efficiency in virtual machine fault detection.

[0005] Summary of the Invention

[0006] The present invention provides a fault detection method and server that can improve the efficiency of virtual machine fault detection. The technical solution is as follows:

[0007] In a first aspect, a fault detection method is provided. The method is applied to a server running a virtual machine, comprising: obtaining the storage type of a storage data volume accessed by the virtual machine; determining a target volume based on the storage type; the target volume being another storage data volume that supports access to the storage type used by the virtual machine; performing an access operation on the target volume to obtain access status information; and determining a fault detection result based on the access status information, the fault detection result indicating whether a fault has occurred in the virtual machine's access to the storage data volume.

[0008] It is understandable that by obtaining the storage type of the virtual machine accessing the storage data volume, determining the target volume, and simulating the process of the virtual machine accessing the storage data volume to perform access operations on the target volume, it is determined whether there is an abnormality in the process of accessing the target volume, and based on the access status information, it is further determined whether a failure occurs in the process of the virtual machine using the storage type to access the storage data volume. At the same time, determining whether a virtual machine has failed by simulating the virtual machine using the storage type to access the target volume is not targeted at a single virtual machine in the server, but is able to monitor whether a virtual machine using the same storage type has failed in the process of accessing the storage data volume, avoiding polling each virtual machine in the server. The access status information of the target volume can be used to determine whether a failure occurs when the virtual machine accesses a storage data volume of the same storage type as supported by the target volume, greatly improving the efficiency of virtual machine fault detection.

[0009] In a possible implementation, an access rate of an access operation to a target volume is obtained, where the access rate includes a read rate and / or a write rate; and a fault detection result is determined according to the read rate and / or the write rate.

[0010] It is understood that by obtaining the access rate of the target volume, it is possible to determine whether a failure has occurred in the process of the virtual machine accessing the storage data volume. Based on the read rate and / or write rate, it is possible to determine whether there is link congestion, link interruption, or slow storage space access during the process of the virtual machine accessing the storage data volume using the storage type. Therefore, it can be determined that a failure has occurred in the process of the virtual machine accessing the storage data volume. This allows for more accurate monitoring of the status of the virtual machine accessing the storage data volume, improving the accuracy of fault detection.

[0011] In one possible implementation, if the read rate is less than a first threshold, the fault detection result is determined to be an abnormality in the virtual machine's read operation on the storage data volume; if the write rate is less than a second threshold, the fault detection result is determined to be an abnormality in the virtual machine's write operation on the storage data volume; if the read rate is less than the first threshold and the write rate is less than the second threshold, the fault detection result is determined to be an abnormality in the virtual machine's access to the storage data volume.

[0012] It is understood that by obtaining the access rate of the target volume, the type of virtual machine failure can be further determined based on the read rate and / or write rate. This can more accurately determine the state of the virtual machine accessing the storage data volume and improve the accuracy of the fault detection results.

[0013] In one possible implementation, determining a target volume according to a storage type includes generating a query command corresponding to the storage type according to the storage type; querying a server based on the query command whether the target volume is included; and obtaining a query result, where the query result is used to indicate whether the target volume is included in the server.

[0014] It can be understood that by generating query commands according to the storage type to determine whether the target volume exists in the server, not only can the query efficiency of the target volume be improved, but also the failure of fault detection caused by the target volume and the storage data volume accessed by the virtual machine supporting different storage types can be avoided.

[0015] In a possible implementation, if the query result indicates that the server does not include the target volume, the target volume is created according to the storage type.

[0016] As can be understood, by checking whether the target volume exists on the server, the accuracy of fault detection is ensured. If the target volume does not exist on the server, it is created based on the storage type, avoiding fault detection failures caused by the target volume and the storage data volume accessed by the virtual machine not supporting the same storage type. This also greatly improves the accuracy of fault detection.

[0017] In a possible implementation, if creation of the target volume fails, the access status information is determined to be creation failure; and based on the access status information, the fault detection result is determined to be target volume creation failure.

[0018] It can be understood that by determining whether the target volume creation is successful, it is possible to detect whether there are any anomalies in the simulated virtual machine's access to the target volume using the storage type, and to determine whether there is a failure in the virtual machine's access to the storage data volume. If the target volume creation fails, it can be determined that the access operation to the target volume has failed, and the access status information can be determined as access failure. Further, it can be determined that the fault detection result indicates that the virtual machine has experienced a failure in accessing the storage data volume. This improves the accuracy and efficiency of virtual machine fault detection.

[0019] In a possible implementation, if the storage type of the storage data volume accessed by the virtual machine fails to be obtained, the access status information is determined to be access failure; based on the access status information, the fault detection result is determined to be a failure of the virtual machine to use the storage type.

[0020] It is understood that by determining whether the virtual machine successfully obtains the storage type of the storage volume it is accessing, it can be determined whether the virtual machine can successfully use the storage type to access the storage volume, and further, whether a failure occurred during the virtual machine's access to the storage data volume. If the virtual machine fails to obtain the storage type of the storage volume it is accessing, it can be determined that a failure occurred during the virtual machine's access to the storage data volume, thus increasing the dimension of fault detection and improving the efficiency of fault detection.

[0021] In one possible implementation, performing an access operation on a target volume includes: generating an access command for the target volume according to a storage type; performing the access operation on the target volume based on the access command and a target period, where the target period indicates a time period for performing the access operation on the target volume, and the access operation includes a read operation and / or a write operation; and obtaining an access rate if the access operation on the target volume is successful, where the access rate includes a read rate and / or a write rate.

[0022] It can be understood that by generating write access commands according to the storage type, the server can simulate the process of the virtual machine using the storage type to access the storage data volume to perform access operations on the target volume according to the target period, which not only improves the access efficiency to the target volume, but also can periodically judge the access status of the virtual machine to the storage data volume through the status and speed of the read operation and / or write operation on the target volume, and further periodically monitor whether there is any fault in the process of the virtual machine accessing the storage data volume, which can improve the accuracy of virtual machine fault detection.

[0023] In one possible implementation, if a read operation on the target volume fails, the fault detection result is determined to be a failure to read the storage data volume; if a write operation on the target volume fails, the fault detection result is determined to be a failure to write to the storage data volume; if a read operation on the target volume fails and a write operation on the target volume fails, the fault detection result is determined to be a failure to access the storage data volume.

[0024] It is understood that the access status information can be determined by determining whether the read operation and / or write operation on the target volume is successful. The access status information can be used to determine whether the virtual machine can normally access the storage data volume. If the read operation on the target volume fails, or the write operation on the target volume fails, the fault detection result can be determined to be a failure in the read and / or write operation on the storage data volume, and further, it can be determined that a fault has occurred in the virtual machine's access to the storage data volume.

[0025] In a possible implementation, if it is determined that the fault detection result indicates that a fault occurs when the virtual machine accesses the storage data volume, the virtual machine is migrated to another server.

[0026] It is understandable that if it is determined that a failure occurs in the process of a virtual machine accessing a storage data volume, the virtual machine in the server can be promptly migrated to another server to avoid affecting the business on the virtual machine, thereby ensuring the smooth operation of the virtual machine business and improving the stability of the virtual machine business.

[0027] In the second aspect, a fault detection device is provided. In an embodiment of the present application, the fault detection device can be divided into functional modules according to the method provided in the first aspect. For example, each functional module can be divided according to each function, or two or more functions can be inherited in one processing module. Exemplarily, an embodiment of the present application can divide the fault detection device into a processing module and a detection module according to the function. The description of the possible technical solutions and beneficial effects executed by each of the functional modules divided above can refer to the technical solutions provided by the first aspect or its corresponding possible implementation methods, and will not be repeated here.

[0028] In a third aspect, an embodiment of the present application provides a server, which includes a processor and a memory for storing processor-executable instructions; the processor is configured to execute instructions so that the server executes the above-mentioned fault detection method.

[0029] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which at least one computer program is stored. The computer program is loaded and executed by a processor to implement the fault detection method as described above.

[0030] In a fifth aspect, embodiments of the present application provide a computer program product or computer program, comprising computer instructions stored in a computer-readable storage medium. A processor of a server reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the server to perform the fault detection method provided in various optional implementations of the above aspects.

[0031] For the specific description of the second to fifth aspects and their various implementations in the embodiments of the present application, reference can be made to the detailed description in the first aspect and its various implementations; and for the beneficial effects of the second to fifth aspects and their various implementations, reference can be made to the analysis of the beneficial effects in the various implementations of the first aspect, which will not be repeated here.

[0032] These and other aspects of the embodiments of the present application will be more clearly understood in the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] FIG1 is a schematic diagram of a distributed storage cluster provided in an embodiment of the present application;

[0034] FIG2 is a schematic diagram of the structure of a fault detection system provided in an embodiment of the present application;

[0035] FIG3 is a schematic diagram of a flow chart of a fault detection method provided in an embodiment of the present application;

[0036] FIG4 is a logic diagram of a fault detection method provided in an embodiment of the present application;

[0037] FIG5 is a schematic diagram of a virtual machine fault detection method including multiple storage types provided in an embodiment of the present application;

[0038] FIG6 is a schematic structural diagram of a fault detection device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0039] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.

[0040] In this document, "plurality" refers to two or more. "And / or" describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates an "or" relationship between the associated objects.

[0041] First, the application scenarios of the embodiments of the present application are exemplarily introduced.

[0042] Storage devices vary widely in capabilities and protocol interfaces, resulting in limited scalability of storage capacity and performance. Virtualization technology, however, can format different storage devices, transforming diverse storage resources into uniformly managed data storage resources that can be used to store virtual machine disks, virtual machine configuration information, and more, making storage management more homogeneous for users. Storage virtualization allows the use of data volumes to meet file modification requirements. Data volumes are specially designed directories that bypass the union file system (UFS) and provide access to one or more virtual machines.

[0043] To meet the storage needs of large amounts of data, storage systems within virtualization platforms can be distributed clusters consisting of multiple servers. Each server is an independent computer system capable of running applications and services. Servers communicate and collaborate with each other over a network. Distributed clusters offer advantages such as high availability, high performance, and scalability, making them widely used in large-scale computing and storage scenarios. Therefore, efficiently detecting and repairing cluster failures ensures the stable operation of the virtualization platform and improves business stability.

[0044] In the related technologies, there are problems such as complex mechanisms, slow detection progress, and low efficiency in detecting virtual machine data volume failures.

[0045] In view of this, an embodiment of the present application provides a fault detection method that determines a target volume based on the storage type used by the virtual machine on the server. On the server, a virtual machine accessing the target volume is simulated based on the storage type used by the virtual machine. Based on the target volume access status information and whether there are any abnormalities during the target volume access process, it is determined whether the virtual machine on the server has a fault.

[0046] Next, the system architecture of the embodiment of the present application is exemplarily introduced.

[0047] For example, Figure 1 shows a schematic diagram of a distributed storage cluster provided by an embodiment of the present application. The above-mentioned fault detection method can be applied to the distributed storage cluster shown in Figure 1 .

[0048] As shown in FIG. 1 , the distributed storage cluster may include a server 101 and a distributed storage space 102 .

[0049] The server 101 is an independent computer system, which can be a standard general-purpose server for running applications and services. The servers 101 can communicate and collaborate with each other through a network, and the communication methods can include point-to-point communication, multicast communication, broadcast communication, etc.

[0050] Distributed storage space 102 can be a storage space formed by a distributed storage system (Hadoop distributed file system, HDFS), a distributed storage system Ceph, or a distributed file system GlusterFS, and is used to store large amounts of data. Among them, Ceph is an open, self-repairing, and self-managing open source distributed storage system with the advantages of high scalability, high performance, and high reliability. It can support a scale of thousands of storage nodes, including block storage interfaces (RBD) and file storage interfaces (NFS), and can be applied to different application scenarios.

[0051] Optionally, the distributed storage cluster may further include a distributed computing system for managing the scheduling and execution of tasks; and may further include a load balancing system for balancing the load between the servers 101 to ensure the stability and reliability of the distributed storage cluster.

[0052] For example, Figure 2 shows a schematic diagram of a fault detection system structure provided by an embodiment of the present application. The fault detection method provided by an embodiment of the present application can be applied to a fault detection system. The fault detection system can be a portion of a distributed storage cluster or the entire distributed storage cluster.

[0053] As shown in FIG. 2 , the fault detection system may include a first server 201 , a second server 202 , and a distributed storage space 102 .

[0054] The first server 201 may include a processor 203, virtual machine A, virtual machine B, and a target volume. The second server may include a processor 203, virtual machine 1, virtual machine 2, and a target volume. If virtual machines A and B access storage data volumes of the same storage type, the first server 201 may obtain the storage type of the storage data volumes accessed by virtual machines A and B and determine the target volume based on the storage type. The target volume may be another storage data volume that supports access by the storage type used by virtual machines A and B.

[0055] The first server 201 can simulate the process of virtual machines A and B accessing the storage data volume according to the storage type of the storage data volume accessed by virtual machines A and B, and perform access operations on the target volume; if an abnormality occurs in the process of virtual machines A and B accessing the target volume, it can be determined based on the access status information that a failure occurred in the process of virtual machines A and B using the storage type to access the storage data volume.

[0056] Processor 203 may include a detection module 204. Detection module 204 may obtain the storage type of the storage data volume accessed by the virtual machine on the node, and based on the storage type, query whether a target volume exists on the server. If the target volume does not exist, detection module 204 may create the target volume based on the storage type. If the target volume exists, detection module 204 may also simulate the process of the virtual machine accessing the storage data volume, perform access operations on the target volume, and determine a fault detection result based on the obtained access status information of the target volume. The fault detection result indicates whether a fault has occurred during the process of the current virtual machine accessing the storage data volume using the storage type.

[0057] In one possible implementation, if it is determined that a failure occurs in the process of virtual machine A and virtual machine B in the first server 201 accessing the storage data volume, the first server 201 can interact with the second server 202 through the communication network. The first server 201 can send an idle virtual machine judgment task to the second server 202, and the detection module 204 in the second server 202 can perform an access operation on the target volume by simulating the process of virtual machine 1 and virtual machine 2 accessing the storage data volume. If it is determined that there is no failure in the process of virtual machine 1 and virtual machine 2 in the second server 202 accessing the storage data volume using the storage type. At this time, when the first server 201 receives feedback that the idle virtual machine status in the second server 202 is normal, the first server 201 can trigger the migration of virtual machine A and virtual machine B, and migrate the business in virtual machine A and virtual machine B from the first server 201 to virtual machine 1 and virtual machine 2 in the second server 202.

[0058] It is understood that by simulating the process of a virtual machine on first server 201 accessing a storage data volume and performing access operations on the target volume, it is possible to determine whether a failure has occurred in the process of virtual machine A and virtual machine B accessing the storage data volume based on whether there are any anomalies in the simulated virtual machine access to the target volume. Because the target volume is a different storage data volume that supports access to the storage type used by virtual machines A and B, the simulated process of accessing the target volume can determine whether a failure has occurred in the process of accessing the storage data volume for all virtual machines using that storage type, rather than for a single virtual machine. Therefore, fault detection can be greatly accelerated, improving the efficiency of virtual machine fault detection.

[0059] It should be noted that the system architecture and application scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0060] For ease of understanding, the fault detection method provided in the present application is exemplarily introduced below with reference to the accompanying drawings. The fault detection method is applicable to the system shown in FIG2 .

[0061] FIG3 shows a schematic flow chart of a fault detection method provided by an embodiment of the present application. The fault detection method includes the following steps:

[0062] S101: Obtain the storage type of a storage data volume accessed by a virtual machine.

[0063] In an embodiment of the present application, a server in a distributed storage cluster may include multiple virtual machines.

[0064] Exemplarily, the storage type of the storage data volume accessed by the virtual machine may be the storage technology used by the virtual machine. Classification of storage types according to storage technology may include block storage RBD, block storage device mapping technology (internet small computer system interface, ISCSI) and file storage NFS.

[0065] Block storage breaks data into blocks and stores each block separately. Each block has an identifier, enabling fast retrieval. In file storage, data is stored as individual pieces of information in folders. When a server wants to access and use data, it needs to obtain the corresponding search path.

[0066] That is, the server can obtain the storage type of the storage data volume accessed by the virtual machine according to the storage type of the storage technology used by the virtual machine in the server.

[0067] In one possible implementation, if the storage type of the virtual machine accessing the storage data volume fails, it can be determined that the access operation to the target volume has failed, and the fault detection result is determined to be access failure. Based on the fault detection result, it is further determined that a failure occurred during the process of the virtual machine accessing the storage data volume.

[0068] For example, if the storage type of the virtual machine accessing the storage data volume cannot be obtained, it means that there is an abnormality in the storage technology being used by the virtual machine, which makes it impossible to simulate the storage type used by the virtual machine to access the target volume. It can also be determined that a failure occurred in the process of the virtual machine accessing the storage data volume.

[0069] For example, if a virtual machine on a server is using block storage technology to access data, if there is an anomaly in the RBD driver or iSCSI protocol, such as a link error or installation package error, the server will be unable to obtain the storage type when obtaining the storage volume type used by the virtual machine. In this case, it will be impossible to simulate the storage type used by the virtual machine to access the target volume.

[0070] S102: Determine a target volume according to the storage type used by the virtual machine.

[0071] In the embodiment of the present application, the target volume is different from the storage data volume currently being used by the virtual machine, wherein the target volume is another storage data volume that supports access by the storage type used by the virtual machine.

[0072] In one possible implementation, the server generates a query command corresponding to the storage type according to the storage type. The server queries the server based on the query command whether the target volume is included in the server, and obtains a query result. The query result may indicate whether the target volume is included in the server.

[0073] In other words, the server can generate a query command corresponding to the storage type based on the storage type of the data volume accessed by the virtual machine. The server can respond to the query command and query whether the target volume is included in the server. If the query result obtained by the server indicates that the target volume is included in the server, the target volume can be determined.

[0074] For example, the virtual machine in the server uses block storage technology, and the storage type of the virtual machine accessing the storage data volume is RBD. The server generates a query command corresponding to the storage type based on the storage type RBD of the virtual machine accessing the storage data volume. The target volume is named by the server identification code plus a fixed suffix. The query command can be qemu-img info rbd:sys / 461D0959-beatheart, which is used to query whether the target volume is included in the server. 461D0959 is the server identification code, and 461D0959-beatheart is used to indicate the target volume of the computing device.

[0075] The server can respond to the query command to check whether the target volume is included in the server. If the query result received by the server is as follows:

[0076] image:json:{"driver":"raw","file":{"pool":"sys","image":"461D0959-beatheart","driver":

[0077] "rbd","namespace":""}} / / Server 461D0959 has a volume that supports the RBD storage type

[0078] File format:raw / / unused volume

[0079] virtual size:1GiB(1073741824bytes) / / The volume size is 1GiB(1073741824bytes)

[0080] The query result may indicate that the image file is returned in the form of a character string, indicating that there is an unused target volume that supports the RBD storage type in the server 461D0959, and the size of the target volume is 1 GiB.

[0081] In a possible implementation, if the obtained query result indicates that the server does not include the target volume, the server may create the target volume according to the storage type.

[0082] That is, the server responds to the query command and queries whether the server includes the target volume; if the server indicates that the target volume does not exist in the server according to the query result obtained, the server can create the target volume according to the storage type.

[0083] For example, based on the query results, the server can determine that the target volume does not exist. Based on the storage type, the server can generate a command to create the target volume. The command might be qemu-img create-f rbd rbd:sys / 461D0959-beatheart -osize=1G, which instructs the server to create a target volume 461D0959-beatheart that supports the RBD storage type and has a volume size of 1GB. The server responds to the creation command. If the server receives a creation result (Formatting 'rbd:sys / test-beatheart', fmt=rbd size=1073741824), the target volume is successfully created.

[0084] In a possible implementation, if creation of the target volume fails, the access status information is determined to be access failure; based on the access status information, it is determined that a fault detection result indicates that a failure occurs in accessing the storage data volume by the virtual machine.

[0085] Exemplarily, the server responds to a creation command. If the creation result obtained is a failure, it can be determined that the access operation to the target volume has failed. The server cannot simulate the virtual machine to use the storage type to access the target volume. The fault detection result is determined to be an operation failure. Based on the fault detection result, it is further determined that a failure has occurred in the process of the virtual machine accessing the storage data volume.

[0086] For example, if a virtual machine on a server uses the RBD storage type to access a data volume, and an RBD driver or installation package exception occurs, the RBD storage cannot be used, potentially causing the server to fail to create the target volume. In this case, the server cannot simulate the virtual machine's use of the storage type to access the target volume.

[0087] S103: The server performs an access operation on the target volume to obtain access status information.

[0088] In an embodiment of the present application, the server can simulate the virtual machine accessing the target volume according to the storage type of the storage data volume accessed by the virtual machine, and perform access operations on the target volume according to the same storage type.

[0089] In one possible implementation, the server may generate a target volume access command based on the storage type. Based on the access command, the server may simulate a virtual machine accessing the target volume using the storage type. If the simulated virtual machine successfully accesses the target volume using the storage type, the server obtains an access rate for the simulated virtual machine accessing the target volume using the storage type.

[0090] In other words, the server can generate an access command based on the storage type of the storage volume being accessed by the virtual machine. The server can respond to the access command, simulating the process of the virtual machine accessing the storage volume using the storage type, and access the target volume. If the target volume is successfully accessed, the server can obtain the access speed.

[0091] For example, if a virtual machine on a server uses the storage type RBD to access a storage data volume, and the target volume on the server is named 461D0959-beatheart, the read command for the target volume generated based on the storage type of the virtual machine accessing the storage data volume is as follows:

[0092] rbd bench sys / 461D0959-beatheart --io-type read --io-size 1M / / The input and output type of the target volume 461D0959-beatheart is read, and the input and output size is 1M

[0093] bench type read io_size 1048576 io_threads 16 bytes 1073741824 pattern sequential / / Read in 16-byte order

[0094] The above command may indicate that the input / output operation type for the target volume in the server 461D0959-beatheart is read through the bench command, the input / output size is 1M, and the reading is performed sequentially in 16 bytes.

[0095] The server can respond to the above access command, simulate the process of the virtual machine accessing the storage data volume, and perform a read operation on the target volume. If the read operation is successful, the server can receive the read operation result, which is as follows:

[0096] elapsed:0ops:1024ops / sec:1098.71bytes / sec:1.07GiB / s / / Read rate is 1.07GiB / s

[0097] The access result may indicate that at the recording time, the number of operations per second is 1024, and the number of bytes per second is 1.07 GiB, that is, the read rate of the read operation on the target volume is 1.07 GiB.

[0098] For example, if a virtual machine on a server uses the storage type RBD to access a storage data volume, and the target volume on the server is named 461D0959-beatheart, the write command to the target volume is generated as follows based on the storage type of the storage data volume accessed by the virtual machine:

[0099] rbd bench sys / 461D0959-beatheart --io-type write --io-size 1M / / Set the input / output type of the target volume 461D0959-beatheart to write and the input / output size to 1M

[0100] bench type write io_size 1048576 io_threads 16 bytes 1073741824 pattern sequential / / Write in 16-byte order

[0101] The above command may indicate that the input / output operation type to be performed on the target volume in the server 461D0959-beatheart through the bench command is write, the input / output size is 1M, and the writing is performed in sequence in 16 bytes.

[0102] The server can respond to the above access command, simulate the process of the virtual machine accessing the storage data volume, and perform a write operation on the target volume. If the write is successful, the server can receive the write result, which is as follows:

[0103] The access results may include write rates completed at different times.

[0104] S104: Determine a fault detection result based on the access status information.

[0105] In an embodiment of the present application, the process of a server accessing a storage data volume can be simulated. The server performs an access operation on the target volume to obtain access status information. The access status information may include the access rate and access status. Therefore, the fault detection result can be determined based on the access status information.

[0106] The fault detection result indicates whether a fault has occurred in the virtual machine accessing the storage data volume. In one possible implementation, the access rate of the target volume access operation is obtained, where the access rate includes a read rate and / or a write rate; and the fault detection result is determined based on the read rate and / or the write rate.

[0107] That is to say, the server can obtain the access rate of the simulated virtual machine using the storage type to access the target volume; the access rate includes the read rate and / or the write rate. If the read rate is less than the first threshold, it means that there may be link congestion, link interruption or slow storage space reading in the process of the virtual machine using the storage type to access the storage data volume, and it can be determined that the fault detection result is that the virtual machine performs an abnormal read operation on the storage data volume; and / or, if the write rate is less than the second threshold, it means that there may be link congestion, link interruption, storage drive abnormality or slow storage space writing in the process of the virtual machine using the storage type to access the storage data volume, and it can be determined that the fault detection result is that the virtual machine performs an abnormal write operation on the storage data volume; and / or, if the read rate is less than the first threshold and the write rate is less than the second threshold, it means that there may be link congestion, link interruption, storage drive abnormality or slow storage space access in the process of the virtual machine using the storage type to access the storage data volume, and the fault detection result is determined to be abnormal access of the virtual machine to the storage data volume, and it is further determined that a fault occurs in the process of the virtual machine using the storage type to access the storage data volume.

[0108] The first threshold and the second threshold may be different values ​​or the same value, and are not specifically limited herein. Exemplarily, the server responds to the access command, simulates a virtual machine accessing a storage data volume, and performs an access operation on the target volume. Based on the access result received by the server, an access rate of the simulated virtual machine accessing the target volume using the storage type is obtained. Based on the access rate, it is determined whether a failure has occurred in the virtual machine accessing the storage data volume.

[0109] For example, if the read rate is less than the first threshold, it can be determined that an abnormality has occurred in the process of the simulated virtual machine using the storage type to perform a read operation on the target volume; and the fault detection result is determined to be an abnormality in the virtual machine's read operation on the storage data volume. If the write rate is lower than the second threshold, it can be determined that an abnormality has occurred in the process of the simulated virtual machine using the storage type to perform a write operation on the target volume; and the fault detection result is determined to be an abnormality in the virtual machine's write operation on the storage data volume. If the read rate exceeds the first threshold and the write rate is lower than the second threshold, it can be determined that an abnormality has occurred in the process of the simulated virtual machine using the storage type to perform an access operation on the target volume, and the fault detection result is determined to be an abnormality in the virtual machine's access operation on the storage data volume. Based on the detection result, it is further determined that a fault has occurred in the process of the virtual machine accessing the storage data volume. Among them, the first threshold can be 1.0MiB / s, and the second threshold can be 1.5MiB / s. Based on the read operation and / or write operation results received by the server, the read rate of the simulated virtual machine using the storage type to access the target volume is 0.73MiB / s, and the write rate is 0.6MiB / s. At this time, the virtual machine may have a network port failure, storage link congestion or link disconnection; therefore, it is determined that an abnormality has occurred in the process of the simulated virtual machine using the storage type to access the target volume, and the fault detection result is determined to be an abnormality in the virtual machine's access operation to the storage data volume. Based on the detection result, it is further determined that a fault has occurred in the process of the virtual machine accessing the storage data volume.

[0110] In a possible implementation, the server may respond to the access command according to the target period, and obtain an access rate of the simulated virtual machine accessing the target volume using the storage type.

[0111] For example, the server may generate an access command based on the storage type of the storage data accessed by the virtual machine, respond to the access command based on the target period, access the target volume, and periodically obtain the access rate.

[0112] For example, if the server determines that the target volume exists, it will generate an access command according to the storage type used by the virtual machine. The server can respond to the access command every 30 seconds, simulate the virtual machine using the storage type to access the target volume, obtain the access rate, and determine the detection result based on the access rate in the access status information. It can dynamically monitor whether there are any abnormalities in the process of using the storage data volume of the virtual machine.

[0113] In one possible implementation, if the read rate and / or write rate of the simulated virtual machine using the storage type to access the target volume obtained according to a specified period is lower than the first threshold and / or the second threshold for more than a third threshold, the fault detection result is determined to be an abnormality in the virtual machine's read and / or write operations on the storage data volume, and it is determined that a failure has occurred in the process of the virtual machine accessing the storage data volume.

[0114] For example, using a server reading a target volume as an example, the server periodically responds to read commands and obtains the read rate of the target volume by the simulated virtual machine using the storage type. If the obtained rate falls below the first threshold more than a third threshold number of times, the fault detection result is determined to be an abnormality in the virtual machine's read operation on the storage data volume. Based on the detection result, it can be further determined that an abnormality occurred in the virtual machine's read operation on the storage data volume.

[0115] For example, the server can respond to read commands every 30 seconds to simulate a virtual machine reading the target volume using the storage type. If the read rates obtained are 1.04 MiB / s, 0.91 MiB / s, 0.24 MiB / s, 0.35 MiB / s, 0.66 MiB / s, and 0.84 MiB / s, and the target volume read rate falls below 1.0 MiB / s five times in a row, it can be determined that an abnormality has occurred during the simulated virtual machine reading operation using the storage type, and that a failure has occurred during the virtual machine reading operation on the storage data volume.

[0116] For example, using a server writing to a target volume as an example, the server periodically responds to write commands and obtains the write rate of the simulated virtual machine using the storage type to write to the target volume. If the obtained rate falls below the second threshold more than a third threshold number of times, the fault detection result is determined to be an abnormality in the virtual machine's write operation to the storage data volume. Based on the detection result, it can be further determined that a fault has occurred in the virtual machine's write operation to the storage data volume.

[0117] For example, the server can respond to read commands every 30 seconds to simulate a virtual machine writing to the target volume using the storage type. If the consecutively obtained write rates are 1.52 MiB / s, 0.98 MiB / s, 0.54 MiB / s, 0.63 MiB / s, 0.42 MiB / s, and 1.04 MiB / s, and the write rate to the target volume falls below 1.5 MiB / s five times in a row, it can be determined that an abnormality has occurred during the simulated virtual machine writing to the target volume using the storage type, and that a failure has occurred during the virtual machine writing to the storage data volume.

[0118] For example, if within a cycle time, the number of times the obtained read rate is lower than the first threshold exceeds the third threshold, and the number of times the obtained write rate is lower than the second threshold exceeds the third threshold, it can be determined that the fault detection result is an abnormal access operation of the virtual machine to the storage data volume, and based on the detection result, it can be further determined that a fault has occurred in the process of the virtual machine writing to the storage data volume.

[0119] In one possible implementation, a fault detection result may be determined based on the access status in the access status information. If a read operation on the target volume fails, the fault detection result is determined to be a read operation failure on the target volume. If a write operation on the target volume fails, the fault detection result is determined to be a write operation failure on the target volume. If both a read operation on the target volume fails and a write operation on the target volume fails, the fault detection result is determined to be a target volume access operation failure. Based on the detection result, it is further determined that a fault has occurred in the process of the virtual machine accessing the storage data volume.

[0120] Exemplarily, the server may generate an access command based on the storage type of the storage data accessed by the virtual machine. If the access command is to perform a read operation on the target volume, the server may respond to the access command and perform a read operation on the target volume. If the server fails to respond to the access command, the simulated virtual machine fails to perform a read operation on the target volume using the storage type, and the fault detection result is determined to be a failure to perform a read operation on the storage data volume. If the access command is to perform a write operation on the target volume, the server may respond to the access command and perform a read operation on the target volume. If the server fails to respond to the access command, the simulated virtual machine fails to perform a write operation on the target volume using the storage type, and the fault detection result is determined to be a failure to perform a read operation on the storage data volume. If the server fails to respond to the access command, the simulated virtual machine fails to perform a read operation on the target volume using the storage type, and the simulated virtual machine fails to perform a write operation on the target volume using the storage type, the fault detection result is determined to be a failure to perform an access operation on the storage data volume. Based on the fault detection result, it is further determined that a fault occurs in the process of the virtual machine accessing the storage data volume.

[0121] For example, if the virtual machine's storage driver or protocol is abnormal, the server may fail to respond to access commands, making it impossible to simulate the virtual machine's use of the storage type to access the target volume. In this case, the fault detection results can be used to determine that a failure has occurred in the virtual machine's access to the storage data volume.

[0122] S105: If it is determined that a failure occurs in accessing the storage data volume by the virtual machine, the virtual machine is migrated to another server.

[0123] If, during a virtual machine's access to a target volume using a storage type, the simulated virtual machine access to the target volume fails, or the simulated virtual machine access to the target volume using the storage type fails, or the rate at which the virtual machine accesses the target volume using the storage type is abnormal, a failure in the virtual machine's access to storage data can be determined. In this case, to ensure that the virtual machine's services are not affected, a virtual machine recovery mechanism can be triggered.

[0124] In one possible implementation, if a virtual machine cannot access a storage data volume normally or a failure occurs during access, it can severely impact services on the virtual machine, even leading to service interruption. Therefore, when a failure is detected during a virtual machine's access to a storage data volume, a virtual machine recovery mechanism needs to be triggered to ensure that the virtual machine's services are not affected and can continue to operate stably.

[0125] In a possible implementation, if it is determined that a failure occurs in the process of the virtual machine accessing the storage data volume, the virtual machine is migrated to another server.

[0126] That is, the server may include a first server and a second server. If it is determined that a failure occurs during access of a storage data volume by a virtual machine on the first server, the first server may trigger migration of the virtual machine from the first server to the second server, where the second server is different from the first server.

[0127] For example, if it can be determined that a failure occurred while a virtual machine on the first server was accessing a storage data volume using a storage type, the first server can send a migration command to the second server to directly migrate the business of the virtual machine on the first server to the second server. Alternatively, if it can be determined that a failure occurred while a virtual machine on the first server was accessing a storage data volume using a storage type, it can be determined that there was no failure while a virtual machine on the second server was accessing a storage data volume using the storage type. In this case, the first server can send a migration command to the second server to directly migrate the business of the virtual machine on the first server to the second server.

[0128] For example, the first server includes virtual machine A. Through the target volume of the detection module in the first server, it can be determined that virtual machine A fails in the process of using the storage type to access the storage data volume. The first server can interact with the second server through the communication network. The first server can send a migration command to the second server. The second server responds to the migration command and migrates the business in virtual machine A in the first server to the second server. Alternatively, if it can be determined that virtual machine A fails in the process of using the storage type to access the storage data volume, the first server can send a query command to the second server to determine that virtual machine 1 in the second server can normally use the same storage type as virtual machine A to access the storage data volume. The first server can send a migration command to the second server. The migration command can instruct the business in virtual machine A in the first server to migrate to virtual machine 1 in the second server. The second server responds to the migration command to complete the migration of virtual machine A and ensure the stable operation of the business in virtual machine A.

[0129] It's important to note that different VM migration methods are available depending on the virtualization platform or deployment environment of the distributed storage cluster, and this is not a limitation. For example, if you use Docker, container images can be quickly migrated between different servers. If you use cloud computing, you can also use the VM migration features and services provided by your cloud server provider for cloud migration.

[0130] It is understandable that if a failure occurs in the process of the virtual machine in the first server accessing the storage data volume, the performance of the virtual machine will be greatly reduced. In order to avoid affecting the business on the virtual machine, the virtual machine in the first server is migrated to the second server in time to ensure the smooth operation of the virtual machine business and improve the stability of the virtual machine business.

[0131] Taking the detection module that can be included in the server as an example, the detection module can simulate the process of the virtual machine using the storage type to access the target volume, and determine whether a fault occurs in the process of the virtual machine using the storage type to access the storage data volume. Exemplarily, Figure 4 shows a logical schematic diagram of a fault detection method provided by an embodiment of the present application. As shown in Figure 4, in a distributed storage cluster, the storage type of the virtual machine accessing the storage data volume in the server is obtained (S401). According to the storage type, a query command corresponding to the storage type is generated, and according to the query command, it is determined whether there is a target volume in the server (S402); if there is a target volume in the server, the storage type used by the virtual machine is simulated to access the target volume (S405). If it is determined that the target volume does not exist, the server creates the target volume according to the storage type (S403). It is determined whether the creation of the target volume is successful (S404). If the creation of the target volume fails, the access status information of the target volume is determined to be an access failure, it is determined that a fault occurs in the process of the virtual machine accessing the storage data volume, the virtual machine recovery mechanism is triggered, and the virtual machine migration is performed (S408). If the target volume is created successfully, the server generates an access command for the target volume based on the storage type, and simulates the storage type used by the virtual machine to access the target volume based on the access command (S405). The status and rate of access to the target volume are obtained (S406), wherein the status of access to the target volume includes whether the storage type used by the simulated virtual machine successfully accesses the target volume; if the storage type used by the simulated virtual machine successfully accesses the target volume, the rate of access to the target volume is obtained; if the storage type used by the simulated virtual machine fails to access the target volume or the read rate of accessing the target volume is less than a first threshold or the write rate is less than a second threshold, it indicates that an abnormality has occurred in the process of the virtual machine using the storage type to access the target volume, and it is determined that the process of the virtual machine accessing the storage data volume has failed. It is determined whether the process of the virtual machine accessing the storage data volume has failed (S407); if it is determined that the process of the virtual machine accessing the storage data volume has failed, the virtual machine recovery mechanism is triggered and the virtual machine is migrated (S408); the virtual machine is migrated to another server in the distributed storage cluster. If the process of the determined virtual machine accessing the storage data volume has not failed, the storage type of the virtual machine accessing the storage data volume in the server is continued to be obtained (S401) for dynamic monitoring.

[0132] Because the detection module in the server simulates a virtual machine in the server using a storage type to access a target volume, and further simulates whether an abnormality occurs when the virtual machine uses the storage type to access the target volume, it can be determined whether a failure occurs during the process of the virtual machine using the storage type to access the storage data volume. Not only can it quickly determine whether a failure occurs during the process of the virtual machine using the storage type to access the storage data volume, but it can also simultaneously determine whether a failure occurs when all virtual machines in the server using the same storage type access the storage data volume, greatly improving the efficiency of virtual machine fault detection.

[0133] It should be noted that by simulating the process of a virtual machine accessing a storage data volume to access the target volume, and judging whether a failure occurs in the process of the virtual machine accessing the storage data volume based on the access status and access rate, it is not aimed at a single virtual machine in the server, but is able to monitor whether a virtual machine using the same storage type fails in the process of accessing the storage data volume.

[0134] In order to improve the efficiency of virtual machine fault detection and simplify the virtual machine fault detection implementation mechanism, the detection module in the server can create different target volumes according to the storage type used by each virtual machine in the server, and dynamically monitor whether each virtual machine in the computing device fails in the process of using the storage type to access the storage data volume. For example, Figure 5 shows a schematic diagram of a virtual machine fault detection including multiple storage types provided by an embodiment of the present application. As shown in Figure 5, the first server 201 is a server in a distributed storage cluster, and the first server 201 can be a server including a processor 203, virtual machine A, virtual machine B, virtual machine C, virtual machine D and virtual machine E. Among them. The processor 203 includes a detection module 204. The detection module 204 can obtain the storage type of the storage data volume accessed by virtual machines A, B, C, D and E; wherein, the storage type of the storage data volume accessed by virtual machines A and B is the same, which is the first storage type; the storage type of the storage data volume accessed by virtual machines C, D and E is the same, which is the second storage type.

[0135] A first target volume and a second target volume are determined based on the first storage type and the second storage type. The first target volume supports the first storage type, and the second target volume supports the second storage type. The detection module 204 can generate access commands based on the first storage type and the second storage type. The detection module 204 simulates the process of virtual machines A and B accessing the storage data volume to access the first target volume, and simulates the process of virtual machines C, D, and E accessing the storage data volume to access the second target volume. Based on the access status and speed of the first target volume and the second volume, it is determined whether a failure occurs during the process of virtual machines A, B, C, D, and E accessing the storage data volume.

[0136] In summary, an embodiment of the present application provides a fault detection method, which is applied to a server, and a virtual machine is running on the server. The server can obtain the storage type of the virtual machine accessing the storage data volume, and determine the target volume based on the storage type. Among them, the storage type supported by the target volume is the same as the storage type of the virtual machine accessing the storage data volume. The server can simulate the process of the virtual machine accessing the storage data volume to access the target volume based on the storage type. If an abnormality occurs in the process of simulating the virtual machine accessing the target volume, it can be determined that the process of the virtual machine accessing the storage data volume has failed. Since the target volume can monitor whether a virtual machine using the same storage type has failed in the process of accessing the storage data volume, the time for virtual machine fault detection is greatly shortened and the efficiency of virtual machine fault detection is improved.

[0137] The above mainly introduces the scheme of the embodiment of the present application from the perspective of method. It can be understood that in order to realize the above functions, the fault detection device includes at least one of the hardware structure and software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0138] The embodiment of the present application can divide the functional units of the fault detection device according to the above method example. For example, each functional unit can be divided according to each function, or two or more functions can be integrated into one processing unit. The above integrated unit can be implemented in the form of hardware or in the form of software functional units. It should be noted that the division of units in the embodiment of the present application is schematic and is only a logical functional division. In actual implementation, there may be other division methods.

[0139] For example, FIG6 shows a schematic diagram of the structure of a fault detection device provided in an embodiment of the present application. As shown in FIG6 , the fault detection device can be applied to a server, and the fault detection device 400 includes:

[0140] Processing module 401 is used to obtain the storage type of the storage data volume accessed by the virtual machine; determine the target volume according to the storage type; the target volume is other storage data volume that supports access by the virtual machine using the storage type.

[0141] The judgment module 402 is used to perform an access operation on the target volume to obtain access status information; based on the access status information, determine a fault detection result, which is used to indicate whether a fault occurs when the virtual machine accesses the storage data volume.

[0142] As an example, in conjunction with FIG2 , part or all of the functions of the processing module 401 and the judgment module 402 of the fault detection device 400 may be implemented by the processor in FIG2 .

[0143] In a possible implementation, the fault detection device 400 further includes a migration module, which is configured to migrate the virtual machine to another server if it is determined that a fault occurs in the process of the virtual machine accessing the storage data volume.

[0144] In a possible implementation, the processing module 401 is further configured to generate a query command corresponding to the storage type according to the storage type; query whether the server includes the target volume according to the query command; and obtain a query result indicating whether the server includes the target volume.

[0145] In a possible implementation, the processing module 401 is further configured to create a target volume according to the storage type if the obtained query result indicates that the server does not include the target volume.

[0146] In one possible implementation, the processing module 401 is further used to generate an access command for the target volume according to the storage type; perform an access operation on the target volume according to the access command and the target period, where the target period is used to indicate the time period for performing the access operation on the target volume, and the access operation includes a read operation and / or a write operation; if the access operation on the target volume is successful, obtain an access rate, where the access rate includes a read rate and / or a write rate.

[0147] In a possible implementation, the judgment module 402 is further configured to obtain an access rate of access operations to the target volume, where the access rate includes a read rate and / or a write rate; and determine a fault detection result based on the read rate and / or the write rate.

[0148] In one possible implementation, the judgment module 402 is further configured to determine a fault detection result based on the read rate and / or write rate, including: if the read rate is less than a first threshold, determining that the fault detection result is an abnormal read operation of the virtual machine on the storage data volume; if the write rate is less than a second threshold, determining that the fault detection result is an abnormal write operation of the virtual machine on the storage data volume; and if the read rate is less than the first threshold and the write rate is less than the second threshold, determining that the fault detection result is an abnormal access of the virtual machine to the storage data volume.

[0149] In a possible implementation, the judgment module 402 is further configured to determine, if the target volume creation fails, that the access status information is creation failure. Based on the access status information, determine that the fault detection result is target volume creation failure.

[0150] In one possible implementation, the determination module 402 is further configured to determine, if a read operation on the target volume fails, that the fault detection result is a read operation failure on the storage data volume. If a write operation on the target volume fails, the determination module 402 is further configured to determine, if a write operation on the target volume fails, that the fault detection result is a write operation failure on the storage data volume. If both a read operation on the target volume fails and a write operation on the target volume fails, the determination module 402 is further configured to determine, if a read operation on the target volume fails and a write operation on the target volume fails, that the fault detection result is a access operation failure on the storage data volume.

[0151] For the detailed description of the above optional methods, please refer to the above method embodiments, which will not be repeated here. In addition, the explanation and description of the beneficial effects of any of the above fault detection devices 400 can refer to the above corresponding method embodiments, which will not be repeated here.

[0152] The present application also provides a computer-readable storage medium storing at least one computer instruction, which is loaded and executed by a processor to implement the fault detection method described in each of the above embodiments. For explanations of the relevant contents and descriptions of the beneficial effects of any of the above-mentioned computer-readable storage media, please refer to the corresponding embodiments above and will not be repeated here.

[0153] The present application also provides a computer program product comprising instructions, which, when executed on a computer, causes the computer to perform any of the methods described in the above embodiments. The computer program product comprises one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated therein. The available media may be magnetic media (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or semiconductor media (e.g., SSD).

[0154] It should be noted that the above-mentioned device for storing computer instructions or computer programs provided in the embodiments of the present application, such as but not limited to, the above-mentioned memory, computer-readable storage medium and communication chip, etc., all have non-volatile (non-transitory). Those skilled in the art should be aware that in the above-mentioned one or more examples, the functions described in the embodiments of the present application can be implemented with hardware, software, firmware or any combination thereof. When implemented using software, these functions can be stored in a computer-readable storage medium or transmitted as one or more instructions or codes on a computer-readable storage medium. Computer-readable storage media include computer storage media and communication media, wherein the communication medium includes any medium that is convenient for transmitting a computer program from one place to another. The storage medium can be any available medium that a general or special-purpose computer can access.

[0155] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A fault detection method, characterized in that Applied to a server; There is a virtual machine running on the server, and the method includes: Obtain the storage type of the virtual machine accessing the storage data volume; Determine the target volume according to the storage type; the target volume is another storage data volume that supports the virtual machine to access using the storage type; Perform an access operation on the target volume to obtain access status information; Based on the access status information, determine a fault detection result, where the fault detection result is used to indicate whether a fault occurs when the virtual machine accesses the storage data volume.

2. The method according to claim 1, wherein The method further includes: Obtain the access rate of the access operation on the target volume, where the access rate includes a read rate and / or a write rate; Determine the fault detection result according to the read rate and / or the write rate.

3. The method according to claim 2, characterized in that, The determining the fault detection result according to the read rate and / or the write rate includes: If the read rate is less than a first threshold, determine that the fault detection result is that the read operation of the virtual machine on the storage data volume is abnormal; And / or, if the write rate is less than a second threshold, determine that the fault detection result is that the write operation of the virtual machine on the storage data volume is abnormal; And / or, if the read rate is less than the first threshold and the write rate is less than the second threshold, determine that the fault detection result is that the virtual machine accesses the storage data volume abnormally.

4. The method according to claim 1 or 2, characterized in that, The determining the target volume according to the storage type includes: Generate a query command corresponding to the storage type according to the storage type; Query whether the target volume is included in the server according to the query command; Obtain a query result; the query result is used to indicate whether the target volume is included in the server.

5. The method according to claim 4, characterized in that, The method further includes: If the query result indicates that the target volume is not included in the server, create the target volume according to the storage type.

6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: If the creation of the target volume fails, determine that the access status information is creation failure; Based on the access status information, determine that the fault detection result is that the creation of the target volume fails.

7. The method according to any one of claims 1 to 6, characterized in that, The performing an access operation on the target volume includes: Generate an access command for the target volume according to the storage type; Perform an access operation on the target volume according to the access command and a target period, where the target period is used to indicate the time period for performing the access operation on the target volume, and the access operation includes a read operation and / or a write operation; If the access operation on the target volume is successful, obtain the access rate, where the access rate includes a read rate and / or a write rate.

8. The method according to claim 7, characterized in that, The method further includes: If the read operation on the target volume fails, determine that the fault detection result is that the read operation on the storage data volume fails; If the write operation on the target volume fails, determine that the fault detection result is that the write operation on the storage data volume fails; If the read operation on the target volume fails and the write operation on the target volume fails, determine that the fault detection result is that the access operation on the storage data volume fails.

9. The method according to any one of claims 1 to 8, characterized in that, The method further includes: If it is determined that the fault detection result indicates that there is a fault in the virtual machine accessing the storage data volume, migrate the virtual machine to another server.

10. A server, characterized in that, The server includes a processor and a memory for storing executable instructions of the processor; The processor is configured to execute the instructions such that the server performs the fault detection method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Dynamically modifying durability properties for individual data volumes

    CN106068507A

  • Storage space data processing method and device of distributed file system

    CN111736772A

  • Fault detection method and server

    CN118113552A

  • Root cause detection and monitoring for storage systems

    US10282245B1

  • Method and system for migrating a virtual machine

    US20110321041A1