A method and device for health monitoring of a storage domain
By determining the health status of the storage domain by determining the multiple states of the storage domain, the problem of slow speed and high workload of storage domain health problems in the prior art is solved, and faster and more efficient storage domain health monitoring is achieved.
Patent Information
- Application Number
- CN202111264360.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-28
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-10-28
AI Technical Summary
In the prior art, when users initiate problem feedback or staff discover storage domain health problems, the speed of poor storage domain health conditions is slow and staff work hard.
The health status of the storage domain is determined by determining the storage domain state, including network state, node state, storage pool state, virtual machine state, heartbeat state and data disk state, so as to detect unhealthy conditions of the storage domain more quickly before users or staff find problems.
It realizes faster detection of unhealthy conditions in the storage domain, reduces workload for workers, and improves the efficiency of problem discovery and handling.
Smart Images

Figure CN114168402B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data storage, and particularly to a method and device for health monitoring of a storage domain. Background Art
[0002] The CFS (Cloud File Storage) storage domain is a cloud storage solution in which multiple node nodes share a heartbeat disk and data disks. When the health of the storage domain is poor, problems such as the heartbeat disk being unable to be normally formatted, the data disks being unable to be normally accessed, the storage pool being unable to be created, and the storage pools or virtual machines in the storage domain being unable to work will occur. Therefore, the health of the storage domain has a great impact on the user's operations and experience. In the prior art, when the user initiates a problem feedback or the staff discovers the above problems, the staff can discover the reason for the poor health of the storage domain only after locating the problem of the storage domain, resulting in a slow speed of discovering the poor health of the storage domain and a large workload for the staff. Summary of the Invention
[0003] The object of the present invention is to provide a method and device for health monitoring of a storage domain, which can discover the health problems of the storage domain before the user initiates a problem feedback or the staff discovers, can more quickly discover the unhealthy situation of the storage domain, and also reduces the workload.
[0004] To solve the above technical problems, the present invention provides a method for health monitoring of a storage domain, including:
[0005] Determine the storage domain state of the storage domain, where the storage domain state includes a combination of one or more of the network state used by the storage domain, the node state of each node in the storage domain, the storage pool state of the storage pool, the virtual machine state of the virtual machine, the heartbeat disk state of the heartbeat disk, and the data disk state of the data disk;
[0006] Determine the health degree state of the storage domain according to the storage domain state, where the health degree state includes a healthy state and an unhealthy state.
[0007] Preferably, when the storage domain state includes the node state, determining the storage domain state of the storage domain includes:
[0008] Judge whether the network of each node in the storage domain is normal, whether each node is powered off, and whether the agent service of each node is normal;
[0009] If all are yes, it is determined that the node state of each node in the storage domain is normal;
[0010] Otherwise, it is determined that the node status of the node in the storage domain is abnormal;
[0011] Determine the health status of the storage domain according to the storage domain status, including:
[0012] Judge whether there is an abnormal node status of the node in the storage domain;
[0013] If there is an abnormal node status of the node, it is determined that the health status is an unhealthy state;
[0014] If there is no abnormal node status of the node, it is determined that the health status is a healthy state.
[0015] Preferably, when the storage domain status includes the network status, determining the storage domain status of the storage domain includes:
[0016] Obtain the communication configuration file of the storage domain;
[0017] Judge whether the communication configuration information in the communication configuration file is communication configuration information indicating that the network of the storage domain is normal;
[0018] If it is communication configuration information indicating that the network of the storage domain is normal, it is determined that the network status is a normal state;
[0019] If it is not communication configuration information indicating that the network of the storage domain is normal, it is determined that the network status is an abnormal state;
[0020] Determine the health status of the storage domain according to the storage domain status, including:
[0021] Judge whether the network status is a normal state;
[0022] If it is a normal state, it is determined that the health status is a healthy state;
[0023] If it is an abnormal state, it is determined that the health status is an unhealthy state.
[0024] Preferably, when the storage domain status includes the storage pool status, determining the storage domain status of the storage domain includes:
[0025] Judge whether the storage pool is unmounted;
[0026] If it is not unmounted, it is determined that the storage pool status is a normal state;
[0027] If it is unmounted, judge whether a user unmount instruction is obtained;
[0028] If the user uninstallation instruction is not obtained, mount the storage pool and determine whether the mounting of the storage pool is successful;
[0029] If the mounting is successful, determine that the status of the storage pool is the normal status;
[0030] If the mounting fails, determine that the status of the storage pool is the abnormal status;
[0031] Determine the health status of the storage domain according to the storage domain status, including:
[0032] Judge whether the status of the storage pool is the normal status;
[0033] If it is the normal status, determine that the health status is the healthy status;
[0034] If it is the abnormal status, determine that the health status is the unhealthy status.
[0035] Preferably, when the storage domain status includes the virtual machine status, determining the storage domain status of the storage domain includes:
[0036] Judge whether all the virtual machines in the storage domain are in the shutdown state;
[0037] If all the virtual machines are in the shutdown state, power on a preset virtual machine and determine whether the power-on of the preset virtual machine is successful;
[0038] If the power-on is successful, determine that the virtual machine status is the normal status;
[0039] If the power-on fails, determine that the virtual machine status is the abnormal status;
[0040] Determine the health status of the storage domain according to the storage domain status, including:
[0041] Judge whether the virtual machine status is the normal status;
[0042] If it is the normal status, determine that the health status is the healthy status;
[0043] If it is the abnormal status, determine that the health status is the unhealthy status.
[0044] Preferably, when the storage domain status includes the heartbeat disk status, determining the storage domain status of the storage domain includes:
[0045] When it is determined that the connection between the heartbeat disk and the data disk is normal and the status of the data disk is the normal status, control the heartbeat disk to generate a preset first instruction and send it to the data disk;
[0046] Determine whether the heartbeat disk has successfully generated the preset first instruction and sent it to the data disk;
[0047] If the preset first instruction is successfully generated and sent to the data disk, determine that the status of the heartbeat disk is the normal status; otherwise, determine that the status of the heartbeat disk is the abnormal status;
[0048] Determine the health status of the storage domain according to the storage domain status, including:
[0049] Determine whether the status of the heartbeat disk is the normal status;
[0050] If it is the normal status, determine that the health status is the healthy status;
[0051] If it is the abnormal status, determine that the health status is the unhealthy status.
[0052] Preferably, when the storage domain status includes the heartbeat disk status, determining the storage domain status of the storage domain includes:
[0053] Send a preset access instruction to the heartbeat disk;
[0054] Determine whether feedback information generated by the heartbeat disk when receiving the preset access instruction is obtained within a preset time period;
[0055] If the feedback information is received, determine that the status of the heartbeat disk is the normal status;
[0056] If the feedback information is not received, determine that the status of the heartbeat disk is the abnormal status;
[0057] Determine whether the status of the heartbeat disk is the normal status;
[0058] If it is the normal status, determine that the health status is the healthy status;
[0059] If it is the abnormal status, determine that the health status is the unhealthy status.
[0060] Preferably, when the storage domain status includes the data disk status, determining the storage domain status of the storage domain includes:
[0061] When it is determined that the connection between the heartbeat disk and the data disk is normal and the status of the heartbeat disk is the normal status, control the heartbeat disk to generate a preset second instruction and send it to the data disk;
[0062] Determine whether the data disk has successfully received the preset second instruction sent by the heartbeat disk;
[0063] If the preset second instruction is received, determine that the status of the data disk is the normal status;
[0064] If the preset second instruction is not received, it is determined that the status of the data disk is an abnormal status;
[0065] Determine the health status of the storage domain according to the storage domain status, including:
[0066] Judge whether the status of the data disk is a normal status;
[0067] If it is a normal status, it is determined that the health status is the healthy status;
[0068] If it is an abnormal status, it is determined that the health status is the unhealthy status.
[0069] Preferably, after determining the health status of the storage domain according to the storage domain status, it further includes:
[0070] If the health status of the storage domain is an unhealthy status, determine the status type of the unhealthy status in the storage domain status;
[0071] Obtain a fault log including the status type;
[0072] Based on the fault log and the relationship between the preset log and the solution, determine the solution to the health status of the storage domain.
[0073] The present invention also provides a health monitoring device for a storage domain, including:
[0074] A memory for storing a computer program;
[0075] A processor for implementing the steps of the health monitoring method of the storage domain as described above when executing the computer program.
[0076] The present invention provides a health monitoring method and device for a storage domain, which obtain the storage domain status of the storage domain. The storage domain status includes a combination of one or more of the network status used by the storage domain, the node status of each node in the storage domain, the storage pool status of the storage pool, the virtual machine status of the virtual machine, the heartbeat disk status of the heartbeat disk, and the data disk status of the data disk. Then, based on the storage domain status, the health status of the storage domain is determined. It can discover the health problems of the storage domain before the user initiates a problem feedback or the staff discovers them, can more quickly discover the unhealthy situation of the storage domain, and also reduces the workload. Description of the Drawings
[0077] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required in the prior art and the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0078] Figure 1 It is a flowchart of a method for health monitoring of a storage domain provided by the present invention;
[0079] Figure 2 It is a flowchart of another method for health monitoring of a storage domain provided by the present invention;
[0080] Figure 3 It is a schematic structural diagram of a device for health monitoring of a storage domain provided by the present invention. Detailed implementation manners
[0081] The core of the present invention is to provide a method and device for health monitoring of a storage domain, which can discover the health problems of the storage domain before the user initiates problem feedback or the staff discovers them, can discover the unhealthy situation of the storage domain faster, and also reduces the workload.
[0082] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0083] Please refer to Figure 1 , Figure 1 It is a flowchart of a method for health monitoring of a storage domain provided by the present invention, including:
[0084] S11: Determine the storage domain status of the storage domain, where the storage domain status includes one or a combination of multiple of the network status used by the storage domain, the node status of each node in the storage domain, the storage pool status of the storage pool, the virtual machine status of the virtual machine, the heartbeat disk status of the heartbeat disk, and the data disk status of the data disk;
[0085] S12: Determine the health status of the storage domain according to the storage domain status, where the health status includes a healthy state and an unhealthy state.
[0086] In order to be able to perceive the health status of the storage domain before the user initiates feedback or the staff discovers a problem with the health status of the storage domain, in this embodiment, the storage domain status of the storage domain is determined, and then the health status of the storage domain is determined according to the storage domain status.
[0087] Specifically, there are many types of storage domain statuses, including the network status used by the storage domain, the node status of each node in the storage domain, the storage pool status of the storage pool, the virtual machine status of the virtual machine, the heartbeat disk status of the heartbeat disk, and the data disk status of the data disk. Since these statuses can all affect the health status of the storage domain, determining the storage domain status of the storage domain is also to determine a combination of one or more of these statuses. Which specific statuses need to be determined depends on the actual situation in work. Determining the health status of the storage domain according to the storage domain status is also to determine the health status of the storage domain according to the status types included in the storage domain status. When the storage domain status includes multiple statuses, when all of these statuses are normal, the health status of the storage domain is determined to be the healthy status. For example, when the storage domain status includes the network status and the node status, the health status of the storage domain is determined according to the node status and the network status. When both the node status and the network status are in the normal state, the health status of the storage domain is the healthy status.
[0088] In addition, for how to determine the storage domain status, the node can report the current status, or the current status can be monitored through a software platform. The present invention does not limit this. For the time to determine the storage domain status, the various statuses in the storage domain status can be periodically determined according to a preset period. When there are multiple different statuses in the storage domain status, the same determination period can be set for all statuses, or different determination periods can be set for different statuses. The present application does not limit this.
[0089] In summary, by obtaining the storage domain status of the storage domain, where the storage domain status includes a combination of one or more of the network status used by the storage domain, the node status of each node in the storage domain, the storage pool status of the storage pool, the virtual machine status of the virtual machine, the heartbeat disk status of the heartbeat disk, and the data disk status of the data disk, and then determining the health status of the storage domain based on the storage domain status, it is possible to discover the health problems of the storage domain before the user initiates problem feedback or the staff discovers them, and it is possible to discover the unhealthy situation of the storage domain faster, and the workload is also reduced.
[0090] Based on the above embodiment:
[0091] Please refer to Figure 2 , Figure 2 which is the flowchart of another method for health monitoring of a storage domain provided by the present invention;
[0092] As a preferred embodiment, when the storage domain state includes node states, determining the storage domain state of the storage domain includes:
[0093] Determine whether the network of each node in the storage domain is normal, whether each node is not powered off, and whether the agent service of each node is normal;
[0094] If all are yes, it is determined that the node states of each node in the storage domain are normal;
[0095] Otherwise, it is determined that there are nodes in the storage domain with abnormal node states;
[0096] Determine the health status of the storage domain according to the storage domain state, including:
[0097] Determine whether there are nodes in the storage domain with abnormal node states;
[0098] If there are nodes with abnormal node states, it is determined that the health status is an unhealthy state;
[0099] If there are no nodes with abnormal node states, it is determined that the health status is a healthy state.
[0100] When the storage domain state includes node states, in order to be able to simply and directly determine the health status according to the storage domain state, in this embodiment, because there are three problems with node nodes: network anomalies, node power outages, and agent services that cannot be started, these problems will all cause abnormal node states. Therefore, first determine whether the network of each node in the storage domain is normal, whether each node is not powered off, and whether the agent service is normal. When the network of each node is normal, each node is not powered off, and the agent service is normal, it is determined that the node state is normal. When there is one or more node nodes with network anomalies and / or power outages and / or agent service anomalies, it is determined that the node state is abnormal at this time. At this time, determining the health status of the storage domain according to the storage domain state is to determine the health status of the storage domain according to the node state. When the node state is normal, the health status of the storage domain is a healthy state. When the node state is abnormal, the status of the storage domain is an unhealthy state.
[0101] In summary, when the storage domain state includes node states, determine whether the network of each node is normal, whether each node is not powered off, and whether the agent service is normal. Based on this, the health status of the storage domain can be simply and directly determined according to the storage domain state.
[0102] As a preferred embodiment, when the storage domain state includes network states, determining the storage domain state of the storage domain includes:
[0103] Obtain the communication configuration file of the storage domain;
[0104] Determine whether the communication configuration information in the communication configuration file is the communication configuration information indicating that the network of the storage domain is normal;
[0105] If it is the communication configuration information indicating that the network of the storage domain is normal, determine that the network status is the normal status;
[0106] If it is not the communication configuration information indicating that the network of the storage domain is normal, determine that the network status is the abnormal status;
[0107] Determine the health status of the storage domain according to the storage domain status, including:
[0108] Determine whether the network status is the normal status;
[0109] If it is the normal status, determine that the health status is the healthy status;
[0110] If it is the abnormal status, determine that the health status is the unhealthy status.
[0111] When the storage domain status includes the network status, in order to simply determine the health status according to the storage domain status, in this embodiment, first obtain the communication configuration file in the storage domain. The communication configuration information in the communication configuration file is not fixed. The communication configuration file of the storage domain can reflect the current network status of the storage domain. At this time, determine whether the communication configuration information in the communication configuration file is the communication configuration information indicating that the network of the storage domain is in the normal state. If so, it means that the network status of the storage domain is normal at this time. If not, it means that the network status of the storage domain is abnormal at this time. The specific reason for the abnormality can, but is not limited to, manually repairing the network status. However, no matter what the reason is, it indicates that there is a problem with the network status of the storage domain. At this time, determining the health status of the storage domain according to the storage domain status is to determine the health status of the storage domain according to the network status. When the network status is normal, determine that the health status of the storage domain is the healthy status. When the network status is abnormal, determine that the health status of the storage domain is the unhealthy status.
[0112] In summary, when the storage domain status includes the network status, determine whether the communication configuration information in the communication configuration file of the storage domain is the communication configuration information indicating that the network status of the storage domain is normal. Based on this, the health status can be simply and directly determined according to the storage domain status.
[0113] As a preferred embodiment, when the storage domain status includes the storage pool status, determine the storage domain status of the storage domain, including:
[0114] Determine whether the storage pool is unmounted;
[0115] If it is not unmounted, the storage pool status is determined to be the normal status;
[0116] If it is unmounted, it is judged whether a user unmount instruction is obtained;
[0117] If a user unmount instruction is not obtained, the storage pool is mounted, and it is judged whether the storage pool is successfully mounted;
[0118] If the mount is successful, the storage pool status is determined to be the normal status;
[0119] If the mount fails, the storage pool status is determined to be the abnormal status;
[0120] Determine the health status of the storage domain according to the storage domain status, including:
[0121] Judge whether the storage pool status is the normal status;
[0122] If it is the normal status, the health status is determined to be the healthy status;
[0123] If it is the abnormal status, the health status is determined to be the unhealthy status.
[0124] When the storage domain status includes the storage pool status, in order to simply and directly determine the health status according to the storage domain status, in this embodiment, the storage pool in the normal situation will be in the mounted state so that virtual machines or other software systems on the storage pool can work. First, judge whether the storage pool is unmounted. When an unmounting situation occurs, it means that the storage pool is not in the normal situation at this time. After the storage pool is unmounted, judge whether a user unmount instruction is obtained. Considering that there may be a situation where the storage pool is unmounted manually. For example, when the storage pool needs to be updated or replaced, the staff will first unmount the storage pool and then perform other operations. If a user unmount instruction is not obtained, it means that there may be a problem with the storage pool status at this time. So, the storage pool is mounted, and it is judged whether the storage pool is successfully mounted. If the mount is successful, it means that the storage pool status is normal, and this unmounting may be an accidental unmounting. If the mount fails, it means that the storage pool status is abnormal, and there is a problem causing the storage pool to not be successfully mounted. At this time, determining the health status of the storage domain according to the storage domain status is to determine the health status of the storage domain according to the storage pool status. When the storage pool status is normal, the health status of the storage domain is determined to be the healthy status. When the storage pool status is abnormal, the health status of the storage domain is determined to be the unhealthy status.
[0125] In summary, when the storage domain status includes the storage pool status, it is necessary to determine whether the storage pool is unmounted. When it is unmounted, it is necessary to determine whether a user unmount instruction is obtained. When the user unmount instruction is not obtained, the storage pool is mounted and it is determined whether the storage pool is successfully mounted. Based on this, the health status can be simply and directly determined according to the storage domain status. In addition, since the storage pool is mounted when the user unmount instruction is not obtained, some situations where the storage pool is accidentally unmounted can be repaired.
[0126] As a preferred embodiment, when the storage domain status includes the virtual machine status, determining the storage domain status of the storage domain includes:
[0127] Determine whether all virtual machines in the storage domain are in the shutdown state;
[0128] If all virtual machines are in the shutdown state, power on the preset virtual machine and determine whether the preset virtual machine is successfully powered on;
[0129] If the power-on is successful, determine that the virtual machine status is the normal state;
[0130] If the power-on fails, determine that the virtual machine status is the abnormal state;
[0131] Determining the health status of the storage domain according to the storage domain status includes:
[0132] Determine whether the virtual machine status is the normal state;
[0133] If it is the normal state, determine that the health status is the healthy state;
[0134] If it is the abnormal state, determine that the health status is the unhealthy state.
[0135] When the storage domain state includes the virtual machine state, in order to simply and directly determine the health status based on the storage domain state, in this embodiment, a task for monitoring the virtual machine state can be established, but is not limited to, on the management platform of the storage domain, to determine whether all virtual machines in the storage domain are in the shutdown state. The virtual machines in the storage domain include the virtual machines created by users and the virtual machines created by staff for testing or other work. In order to be able to work at any time, the virtual machines of users or staff will remain powered on. Usually, not all of these virtual machines will be in the shutdown state. When all virtual machines are in the shutdown state, power on the preset virtual machine and determine whether the preset virtual machine is successfully powered on. Since the staff cannot interfere with or use the virtual machines created by users, the preset virtual machine is the virtual machine created by the staff. When the preset virtual machine is successfully powered on, it indicates that the virtual machine state is normal. When the preset virtual machine fails to power on, it indicates that the virtual machine state is abnormal. At this time, determining the health status of the storage domain based on the storage domain state is to determine the health status of the storage domain based on the virtual machine state. When the virtual machine state is normal, the health status of the storage domain is determined to be the healthy state. When the virtual machine state is abnormal, the health status of the storage domain is determined to be the unhealthy state.
[0136] In summary, determining whether all virtual machines in the storage domain are in the shutdown state, powering on the preset virtual machine and determining whether it is successfully powered on when all are in the shutdown state can simply and directly determine the health status based on the storage domain state.
[0137] As a preferred embodiment, when the storage domain state includes the heartbeat disk state, determining the storage domain state of the storage domain includes:
[0138] When it is determined that the connection between the heartbeat disk and the data disk is normal and the data disk state is the normal state, control the heartbeat disk to generate a preset first instruction and send it to the data disk;
[0139] Determine whether the heartbeat disk successfully generates a preset first instruction and sends it to the data disk;
[0140] If the preset first instruction is successfully generated and sent to the data disk, then determine that the heartbeat disk state is the normal state; otherwise, determine that the heartbeat disk state is the abnormal state;
[0141] Determining the health status of the storage domain based on the storage domain state includes:
[0142] Determine whether the heartbeat disk state is the normal state;
[0143] If it is the normal state, then determine that the health status is the healthy state;
[0144] If it is the abnormal state, then determine that the health status is the unhealthy state.
[0145] When the storage domain status includes the heartbeat disk status, in order to simply and directly determine the health status based on the storage domain status, in this embodiment, the control heartbeat disk generates a preset first instruction and sends it to the data disk. Since the connection and data transmission between the heartbeat disk and the data disk are composed of the heartbeat disk, the data disk, and the connection relationship between the two, when the heartbeat disk sends an instruction to the data disk, in addition to the problem that the heartbeat disk itself may not be able to send the preset first instruction to the data disk, there may also be a situation where the connection between the two is disconnected or the data disk cannot receive the preset first instruction. Since these two problems have nothing to do with the heartbeat disk itself, it is necessary to determine that both the connection relationship and the data disk are in a normal state at this time. If the preset first instruction can be generated and sent to the data disk, it is determined that the heartbeat disk status is a normal state. If the preset first instruction cannot be generated and sent to the data disk, it may be that the preset first instruction cannot be generated, or it cannot be sent to the data disk after the preset first instruction is generated, and it is determined that the heartbeat disk status is an abnormal state. At this time, determining the health status of the storage domain based on the storage domain status is to determine the health status of the storage domain based on the heartbeat disk status. When the heartbeat disk status is normal, it is determined that the health status of the storage domain is a healthy state. When the heartbeat disk status is abnormal, it is determined that the health status of the storage domain is an unhealthy state.
[0146] In addition, the preset first instruction can be a write instruction or other instructions that can be received by the data disk. This application does not make a limitation here.
[0147] In summary, when it is determined that the connection between the heartbeat disk and the data disk is normal and the data disk status is normal, the control heartbeat disk generates a preset first instruction and sends it to the data disk, and then determines whether the heartbeat disk successfully generates the preset first instruction and sends it to the data disk, so that the health status can be simply and directly determined based on the storage domain status.
[0148] As a preferred embodiment, when the storage domain status includes the heartbeat disk status, determining the storage domain status of the storage domain includes:
[0149] Sending a preset access instruction to the heartbeat disk;
[0150] Determining whether feedback information generated by the heartbeat disk when receiving the preset access instruction is obtained within a preset time period;
[0151] If the feedback information is received, it is determined that the heartbeat disk status is a normal state;
[0152] If the feedback information is not received, it is determined that the heartbeat disk status is an abnormal state;
[0153] Determining whether the heartbeat disk status is a normal state;
[0154] If it is a normal state, it is determined that the health status is a healthy state;
[0155] If it is an abnormal state, the health status is determined to be an unhealthy state.
[0156] When the storage domain status includes the heartbeat disk status, considering the situation where the heartbeat disk itself may be damaged and unable to be accessed, resulting in the health status of the storage domain becoming unhealthy. In this embodiment, first, a preset access instruction is sent to the heartbeat disk, and it is determined whether feedback information generated by the heartbeat disk when receiving the preset access instruction can be obtained within a preset time period. The preset time period can be, but is not limited to, a time period set manually. Different heartbeat disks may provide feedback information at different speeds after receiving the preset access instruction, and the length of the preset time period can be adjusted according to the specific situation. If the information can be received, it indicates that the heartbeat disk can be accessed normally at this time. The normal access of the heartbeat disk indicates that there is no problem with the heartbeat disk at this time. If the feedback information cannot be received, it indicates that the heartbeat disk cannot be accessed normally at this time, indicating that there may be hardware damage or internal data damage to the heartbeat disk, etc. At this time, determining the health status of the storage domain according to the storage domain status is to determine the health status of the storage domain according to the heartbeat disk status. When the heartbeat disk status is normal, the health status of the storage domain is determined to be a healthy state. When the heartbeat disk status is abnormal, the health status of the storage domain is determined to be an unhealthy state.
[0157] In summary, sending a preset access instruction to the heartbeat disk and determining whether feedback information generated by the heartbeat disk when receiving the preset access instruction can be obtained within the preset time period can simply and directly determine the health status according to the storage domain status.
[0158] As a preferred embodiment, when the storage domain status includes the data disk status, determining the storage domain status of the storage domain includes:
[0159] When it is determined that the connection between the heartbeat disk and the data disk is normal and the heartbeat disk status is in a normal state, control the heartbeat disk to generate a preset second instruction and send it to the data disk;
[0160] Determine whether the data disk has successfully received the preset second instruction sent by the heartbeat disk;
[0161] If the preset second instruction is received, the data disk status is determined to be a normal state;
[0162] If the preset second instruction is not received, the data disk status is determined to be an abnormal state;
[0163] Determining the health status of the storage domain according to the storage domain status includes:
[0164] Determine whether the data disk status is a normal state;
[0165] If it is a normal state, the health status is determined to be a healthy state;
[0166] If it is an abnormal state, it is determined that the health status is an unhealthy state.
[0167] When the storage domain status includes the data disk status, in order to simply and directly determine the health status according to the storage domain status, in this embodiment, the heartbeat disk is controlled to generate a preset second instruction and send it to the data disk. Since the connection and data transmission between the heartbeat disk and the data disk are composed of the heartbeat disk, the data disk, and the connection relationship between the two, when the heartbeat disk sends an instruction to the data disk, in addition to the problem that the data disk itself may not be able to receive the preset second instruction normally, there may also be a situation where the connection between the two is disconnected or the heartbeat disk cannot send the preset second instruction. Since these two problems have nothing to do with the data disk itself, it is necessary to determine that both the connection relationship and the heartbeat disk are in a normal state at this time. If the data disk can receive the preset second instruction, it is determined that the data disk status is a normal state. If the data disk does not receive the preset second instruction, it is determined that the data disk status is an abnormal state. At this time, determining the health status of the storage domain according to the storage domain status is to determine the health status of the storage domain according to the data disk status. When the data disk status is normal, it is determined that the health status of the storage domain is a healthy state. When the data disk status is abnormal, it is determined that the health status of the storage domain is an unhealthy state.
[0168] In addition, the preset second instruction can be the same as the preset first instruction, or it can be an instruction that other data disks can receive. This application does not make a limitation here.
[0169] In summary, when it is determined that the connection between the data disks is normal and the data disk status is normal, controlling the data disk to generate a preset first instruction and send it to the data disk, and then determining whether the data disk successfully generates the preset first instruction and sends it to the data disk can simply and directly determine the health status according to the storage domain status.
[0170] As a preferred embodiment, after determining the health status of the storage domain according to the storage domain status, it further includes:
[0171] S13: If the health status of the storage domain is an unhealthy state, determine the status type of the unhealthy state in the storage domain status;
[0172] S14: Obtain a fault log including the status type;
[0173] S15: Determine a solution for the health status of the storage domain based on the fault log and the relationship between the preset log and the solution.
[0174] In order to simply and directly understand which states in the storage domain have problems and the corresponding solutions to the problems, in this embodiment, after determining the health status of the storage domain, if the health status of the storage domain is an unhealthy state, first determine the state types with unhealthy states in the storage domain status, then obtain the fault logs including these unhealthy types, and then determine the solution to the health status of the storage domain based on the fault logs and the solutions corresponding to the preset fault logs.
[0175] Specifically, the health status of the storage domain has two types: healthy state and unhealthy state. When the health status is an unhealthy state, since each state in the storage domain status has its own normal state and abnormal state, at this time, the state types in the abnormal state in the storage domain status can be determined, and then the fault logs including the state types in the abnormal state are obtained. The preset log-solution relationship includes the solutions corresponding to different state types in the abnormal state. Based on the fault logs and the preset log-solution relationship, the solution to the health status at this time can be determined. For example, when the storage domain status includes the storage pool status and the virtual machine status, if the storage pool status is normal but the virtual machine status is abnormal, at this time, it is determined that the state type in the abnormal state in the storage domain status is the virtual machine status, then the fault logs including the virtual machine status are obtained, and then the solution corresponding to the virtual machine status is determined, and the solution to the health status at this time can be determined.
[0176] In addition, in order to be able to understand the specific situation of the unhealthy state, in addition to including the state types in the abnormal state, the fault logs can also include the number of times the state type is in the abnormal state. When the storage domain status includes the node status of the node, it can also include the name of the node with the abnormal state, and an alarm can also be issued when the unhealthy state occurs, so as to better handle these abnormal states according to the severity of the abnormal state.
[0177] In summary, by determining the state types with unhealthy states in the storage domain status, then obtaining the fault logs including the state types, and finally determining the solution to the health status of the storage domain based on the relationship between the fault logs and the preset logs and solutions, it is possible to simply and directly understand which states in the storage domain have problems and the corresponding solutions to the problems.
[0178] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a health monitoring device for a storage domain provided by the present invention, including:
[0179] A memory 11 for storing computer programs;
[0180] A processor 12, configured to implement the steps of the health monitoring method for the storage domain as described above when executing a computer program.
[0181] For a detailed introduction to a health monitoring device for a storage domain provided by the present invention, please refer to the embodiments of the health monitoring method for the storage domain above, and details are not described herein again in this application.
[0182] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference may be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference may be made to the description in the method part for related parts.
[0183] It should also be noted that in this specification, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or further elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
Claims
1. A method for health monitoring of a storage domain, characterized in that Including: Determine the storage domain status of the storage domain, where the storage domain status includes a combination of multiple of the network status used by the storage domain, the node status of each node in the storage domain, the storage pool status of the storage pool, the virtual machine status of the virtual machine, the heartbeat disk status of the heartbeat disk, and the data disk status of the data disk; Determine the health status of the storage domain according to the storage domain status, where the health status includes a healthy state and an unhealthy state; Among them, when the storage domain status includes the virtual machine status, if all the virtual machines are in the shutdown state, power on a preset virtual machine. If the power-on is successful, determine that the virtual machine status is the normal state; if the power-on fails, determine that the virtual machine status is the abnormal state; if the virtual machine status is the normal state, determine that the health status is the healthy state; if the virtual machine status is the abnormal state, determine that the health status is the unhealthy state; When it is in the unhealthy state and the status type of the unhealthy state in the storage domain status is the virtual machine status, obtain a fault log including the virtual machine status and the number of abnormal states, and determine a solution corresponding to the virtual machine status according to the fault log.
2. The health monitoring method for a storage domain according to claim 1, characterized in that When the storage domain status includes the node status, determining the storage domain status of the storage domain includes: Judge whether the network of each node in the storage domain is normal, whether each node is not powered off, and whether the agent service of each node is normal; If all are yes, determine that the node status of each node in the storage domain is normal; Otherwise, determine that there is a node in the storage domain whose node status is abnormal; Determine the health status of the storage domain according to the storage domain status, including: Judge whether there is a node in the storage domain whose node status is abnormal; If there is a node in the storage domain whose node status is abnormal, determine that the health status is the unhealthy state; If there is no node in the storage domain whose node status is abnormal, determine that the health status is the healthy state.
3. The health monitoring method of the storage domain according to claim 1, characterized in that, When the storage domain status includes the network status, determining the storage domain status of the storage domain includes: Obtain the communication configuration file of the storage domain; Judge whether the communication configuration information in the communication configuration file is communication configuration information indicating that the network of the storage domain is normal; If it is communication configuration information indicating that the network of the storage domain is normal, determine that the network status is the normal state; If it is not communication configuration information indicating that the network of the storage domain is normal, determine that the network status is the abnormal state; Determine the health status of the storage domain according to the storage domain status, including: Judge whether the network status is the normal state; If it is the normal state, determine that the health status is the healthy state; If it is the abnormal state, determine that the health status is the unhealthy state.
4. The health monitoring method of the storage domain according to claim 1, characterized in that, When the storage domain status includes the storage pool status, determining the storage domain status of the storage domain includes: Judge whether the storage pool is unmounted; If it is not unmounted, determine that the storage pool status is the normal status; If it is unmounted, determine whether a user unmount instruction is obtained; If the user unmount instruction is not obtained, mount the storage pool and determine whether the storage pool is successfully mounted; If the mount is successful, determine that the storage pool status is the normal status; If the mount fails, determine that the storage pool status is the abnormal status; Determine the health status of the storage domain according to the storage domain status, including: Determine whether the storage pool status is the normal status; If it is the normal status, determine that the health status is the healthy status; If it is the abnormal status, determine that the health status is the unhealthy status.
5. The health monitoring method of the storage domain according to claim 1, wherein When the storage domain status includes the virtual machine status, determine the storage domain status of the storage domain, including: Determine whether all virtual machines in the storage domain are in the shutdown state; If all the virtual machines are in the shutdown state, power on a preset virtual machine and determine whether the preset virtual machine is successfully powered on; If the power on is successful, determine that the virtual machine status is the normal status; If the power on fails, determine that the virtual machine status is the abnormal status; Determine the health status of the storage domain according to the storage domain status, including: Determine whether the virtual machine status is the normal status; If it is the normal status, determine that the health status is the healthy status; If it is the abnormal status, determine that the health status is the unhealthy status.
6. The health monitoring method for a storage domain according to claim 1, wherein, When the storage domain status includes the heartbeat disk status, determine the storage domain status of the storage domain, including: When it is determined that the connection between the heartbeat disk and the data disk is normal and the data disk status is the normal status, control the heartbeat disk to generate a preset first instruction and send it to the data disk; Determine whether the heartbeat disk successfully generates the preset first instruction and sends it to the data disk; If the preset first instruction is successfully generated and sent to the data disk, determine that the heartbeat disk status is the normal status; otherwise, determine that the heartbeat disk status is the abnormal status; Determine the health status of the storage domain according to the storage domain status, including: Determine whether the heartbeat disk status is the normal status; If it is the normal status, determine that the health status is the healthy status; If it is the abnormal status, determine that the health status is the unhealthy status.
7. The health monitoring method for a storage domain according to claim 1, wherein When the storage domain status includes the heartbeat disk status, determine the storage domain status of the storage domain, including: Send a preset access instruction to the heartbeat disk; Determine whether feedback information generated by the heartbeat disk when receiving the preset access instruction is obtained within a preset time period; If the feedback information is received, determine that the heartbeat disk status is the normal status; If the feedback information is not received, determine that the heartbeat disk status is the abnormal status; Determine whether the heartbeat disk status is the normal status; If it is the normal status, determine that the health status is the healthy status; If it is the abnormal status, determine that the health status is the unhealthy status.
8. The health monitoring method of the storage domain according to claim 1, characterized in that When the storage domain status includes the data disk status, determine the storage domain status of the storage domain, including: When it is determined that the connection between the heartbeat disk and the data disk is normal and the status of the heartbeat disk is in a normal state, control the heartbeat disk to generate a preset second instruction and send it to the data disk; Determine whether the data disk successfully receives the preset second instruction sent by the heartbeat disk; If the preset second instruction is received, determine that the status of the data disk is in a normal state; If the preset second instruction is not received, determine that the status of the data disk is in an abnormal state; Determine the health status of the storage domain according to the storage domain status, including: Determine whether the status of the data disk is in a normal state; If it is in a normal state, determine that the health status is the healthy state; If it is in an abnormal state, determine that the health status is the unhealthy state.
9. The method for health monitoring of a storage domain according to any one of claims 1 to 8, characterized in that, After determining the health status of the storage domain according to the storage domain status, it further includes: If the health status of the storage domain is in an unhealthy state, determine the status type of the unhealthy state in the storage domain status; Obtain a fault log including the status type; Based on the fault log and the relationship between the preset log and the solution, determine the solution to the health status of the storage domain.
10. A health monitoring device for a storage domain, characterized in that, Including: A memory for storing a computer program; A processor for implementing the steps of the health monitoring method of the storage domain according to any one of claims 1 to 9 when executing the computer program.
Citation Information
Patent Citations
Storage pool monitoring method and device
CN107453951A
Virtualization system after-disaster recovery system, method and device and readable storage medium
CN110941508A