Fault Tolerance Method, Product, Device and Storage Medium for Distributed Storage File System
By storing the read and write statistics of nodes on the heartbeat disk of the OCFS2 cluster and judging the number of consecutive exceptions, the problem that the cluster is unavailable due to occasional IO errors is solved, and higher stability and reliability are achieved.
Patent Information
- Application Number
- CN202510240715.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-03
AI Technical Summary
The existing distributed storage file system OCFS2 cluster has too strict requirements on IO reading and writing, and cannot tolerate occasional IO errors, resulting in nodes being evicted from the cluster due to occasional IO errors, resulting in frequent reconstruction nodes or unavailable for the entire cluster.
By storing the read and write statistics of each node on the heartbeat disk, including normal or abnormal situations of read and write input and output and the number of consecutive abnormalities, it is determined whether the number of consecutive abnormalities in the slot data of the node is greater than the threshold. If so, the control node is offline to avoid occasional IO errors affecting the entire cluster.
The IO fault tolerance mechanism is added to avoid cluster unavailability caused by occasional IO errors, and improve the stability and reliability of the distributed storage file system.
Smart Images

Figure CN119718762B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of computer data storage, and particularly to a method, product, computer device, and storage medium for fault tolerance of a distributed storage file system. Background Art
[0002] Cluster systems are usually designed to share network and disk resources and use node heartbeats to communicate with each other to maintain services within the cluster. Cluster systems typically include the function of detecting non-responsive nodes and removing them from the cluster to maintain their ability to continue providing services and avoid failures or data corruption. OCFS2 is a distributed storage file system that utilizes shared storage and can be used for many general computing tasks that require shared cluster storage. OCFS2 uses disk-based heartbeats to determine the availability of nodes within the cluster. If a node does not respond, other nodes in the cluster can continue to provide available services by removing the node from the cluster. If a certain node cannot communicate through the shared storage, the node will be isolated, which means it will be evicted from the cluster so that the cluster can return to a running state. The reason for the abnormal function of the OCFS2 cluster is that the existing heartbeat mechanism has overly strict requirements for IO reads and writes and cannot tolerate occasional IO errors, resulting in nodes in the OCFS2 cluster being evicted from the cluster due to occasional IO errors, causing frequent reconstruction of nodes or rendering the entire OCFS2 cluster unavailable. Summary of the Invention
[0003] Based on this, there is provided a method, apparatus, computer device, and storage medium for fault tolerance of a distributed storage file system that can prevent nodes in a distributed storage file system cluster from being evicted from the cluster due to occasional IO errors, adding an IO fault tolerance mechanism to avoid the problem of the entire distributed storage file system cluster becoming unavailable due to occasional IO errors.
[0004] On the one hand, there is provided a method for fault tolerance of a distributed storage file system, the method comprising:
[0005] Obtain all nodes in the shared cluster connected to the heartbeat disk, and store the read and write statistical data of each node during each disk heartbeat on the heartbeat disk as slot data, where the read and write statistical data includes whether the read and write input / output is normal or abnormal and the number of consecutive occurrences of abnormal read and write input / output;
[0006] During each disk heartbeat of the heartbeat disk, in response to the first node receiving read and write input / output data, update the slot data of the first node, read the slot data of the second node in the shared cluster, filter and save the valid generation information of the second node according to the root field in the slot data of the second node, and count the number of consecutive occurrences of abnormal read and write input / output;
[0007] Determine whether the number of consecutive occurrences of read / write input / output exceptions in the slot data of the second node is greater than the first threshold. If so, control the second node to go offline. If not, verify the slot data of the next node.
[0008] In one embodiment, storing the read / write statistical data of each node during each disk heartbeat through the slot data on the heartbeat disk includes:
[0009] Set the initial value of the number of consecutive occurrences of read / write input / output exceptions to zero;
[0010] Determine whether the read / write input / output of the current node is normal or abnormal during the current disk heartbeat;
[0011] In response to the read / write input / output of the current node being normal during the current disk heartbeat, determine whether the read / write input / output of the current node is normal or abnormal during the next heartbeat read / write;
[0012] In response to the read / write input / output of the current node being abnormal during the current disk heartbeat, increment the number of consecutive occurrences of read / write input / output exceptions of the current node by one.
[0013] In one embodiment, determining whether the read / write input / output of the current node is normal or abnormal during the current disk heartbeat includes:
[0014] Set the generation information of each node to be stored on the heartbeat disk. The read / write statistical data further includes the generation information of the current node performing the read / write;
[0015] Read the heartbeat read / write statistical data of the second node sharing a connection with the first node, and determine whether the generation information of the current node performing the read / write in the read / write statistical data is the same as the generation information of the second node stored on the heartbeat disk. If they are different, determine that the current read / write input / output is abnormal. If they are the same, determine that the current read / write input / output is normal.
[0016] In one embodiment, determining whether the read / write input / output of the current node is normal or abnormal during the current disk heartbeat further includes:
[0017] Determine whether there is an error in the data being read / written in the read / write statistical data;
[0018] If there is no error and the generation information of the current node performing the read / write in the read / write statistical data is the same as the generation information of the second node stored on the heartbeat disk, determine that the current read / write input / output is normal;
[0019] If there is an error and the generation information of the current node performing the read / write in the read / write statistical data is the same as the generation information of the second node stored on the heartbeat disk, ignore the current read / write input / output error.
[0020] In one embodiment, reading the slot data of the second node in the shared cluster, filtering and saving the valid generation information of the second node according to the root field in the slot data of the second node, and counting the number of consecutive read / write input / output exceptions includes:
[0021] Setting the slot data further includes a first root field, a second root field, a third root field, and a fourth root field. The initial value of the slot data is zero. The first root field is used to record the valid generation information of the corresponding node. The second root field is used to record the number of consecutive times the generation information of the corresponding node is the same when read continuously. The third root field is used to record the number of consecutive times the generation information of each node is different. The fourth root field is used to record whether the generation information of each node has been authenticated as legal;
[0022] In response to reading the slot data of the second node, obtaining the value of the fourth root field of the second node. If the value of the fourth root field of the second node is zero, it is determined that the first node has not saved the valid generation information of the second node. If the value of the fourth root field of the second node is not zero, it is determined that the first node has saved the valid generation information of the second node;
[0023] In response to the first node not saving the valid generation information of the second node, obtaining the value of the second root field of the second node, and writing the corresponding node generation information when the read / write input / output is normal into the first root field as the valid generation information of the second node by determining whether the second root field of the second node is greater than a third threshold;
[0024] In response to the first node having saved the generation information of the second node, obtaining the value of the third root field of the second node, and setting the number of consecutive read / write input / output exceptions in the slot data of the second node by determining whether the third root field of the second node is greater than a second threshold.
[0025] In one embodiment, the fourth root field is used to record whether the generation information of each node has been authenticated as legal, including:
[0026] Setting the fourth root field to zero or one;
[0027] In response to obtaining that the fourth root field is one, it is determined that the generation information of the corresponding node has been authenticated as legal;
[0028] In response to obtaining that the fourth root field is zero, it is determined that the generation information of the corresponding node has not been authenticated as legal.
[0029] In one embodiment, the step of writing the corresponding node generation information when the read / write input / output is normal into the first root field as the valid generation information of the second node by determining whether the second root field of the second node is greater than a third threshold includes:
[0030] Determine whether the value of the second root field of the second node is greater than the third threshold. If so, set the fourth root field to one. If not, determine whether the second root field of the second node is zero;
[0031] If the second root field of the second node is zero, read the generation information of the second node from the heartbeat disk and store the generation information of the second node read from the heartbeat disk into the first root field, and at the same time increment the second root field by one;
[0032] If the second root field of the second node is not zero, indicating that the generation information of the second node has been stored in the first root field, then determine whether the value of the first root field of the second node in the slot data is the same as the generation information of the second node read from the heartbeat disk. If the same, increment the second root field by one. If different, save the generation information of the second node read from the heartbeat disk to the first root field, and at the same time set the second root field to one;
[0033] If the value of the second root field of the second node after incrementing is greater than the third threshold, set the fourth root field to one.
[0034] In one embodiment, when the value of the second root field of the second node is less than or equal to the third threshold, before determining whether the second root field of the second node is zero, further include:
[0035] Determine whether the generation information of the second node read from the heartbeat disk is zero;
[0036] If it is zero, end. If not, determine whether the second root field of the second node is zero.
[0037] In one embodiment, the method of setting the number of consecutive read / write input / output exceptions in the slot data of the second node by determining whether the third root field of the second node is greater than the second threshold includes:
[0038] If the value of the fourth root field of the second node is one, determine whether the value of the first root field of the second node in the slot data is the same as the generation information of the second node read from the heartbeat disk, and simultaneously determine whether the value of the third root field of the second node is zero;
[0039] If the value of the first root field of the second node in the slot data is the same as the generation information of the second node read from the heartbeat disk, and the value of the third root field of the second node is not zero and less than the second threshold, set the third root field to zero;
[0040] If the value of the first root field of the second node in the slot data is different from the generation information of the second node read from the heartbeat disk, increment the count of the third root field by one, and determine whether the incremented third root field is greater than the second threshold. If so, increment the number of consecutive read / write input / output exceptions in the slot data of the second node by one; if not, end.
[0041] In one embodiment, when controlling the second node to go offline, it further includes:
[0042] Clear the slot data of the second node, and trigger a command to kick or isolate the second node from the shared cluster.
[0043] In one embodiment, the distributed storage file system fault tolerance method further includes:
[0044] Before the heartbeat disk performs the next disk heartbeat, determine whether the slot data of all nodes has been verified. If so, end; if not, verify the slot data of the next node.
[0045] In one embodiment, the distributed storage file system fault tolerance method further includes:
[0046] In response to the first node completing the verification of the slot data of all nodes, control the second node to receive read / write input / output data, update the slot data of the second node, read and verify the slot data of the third node in the shared cluster until all nodes in the shared cluster connected to the heartbeat disk have completed receiving read / write input / output data.
[0047] On the other hand, a computer program product is provided, including a computer program that, when executed by a processor, implements the following steps:
[0048] Obtain all nodes in the shared cluster connected to the heartbeat disk, and store the read / write statistical data of each node during each disk heartbeat on the heartbeat disk as slot data. The read / write statistical data includes whether the read / write input / output is normal or abnormal, and the number of consecutive read / write input / output exceptions;
[0049] During each disk heartbeat of the heartbeat disk, in response to the first node receiving read / write input / output data, update the slot data of the first node, read the slot data of the second node in the shared cluster, filter and save the valid generation information of the second node according to the root field in the slot data of the second node, and count the number of consecutive read / write input / output exceptions;
[0050] Determine whether the number of consecutive occurrences of read / write input / output exceptions in the slot data of the second node is greater than the first threshold. If so, control the second node to go offline. If not, verify the slot data of the next node.
[0051] In another aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:
[0052] Obtain all nodes in the shared cluster connected to the heartbeat disk, and store the read / write statistical data of each node during each disk heartbeat on the heartbeat disk as slot data. The read / write statistical data includes whether the read / write input / output is normal or abnormal, and the number of consecutive occurrences of read / write input / output exceptions;
[0053] During each disk heartbeat of the heartbeat disk, in response to the first node receiving read / write input / output data, update the slot data of the first node, read the slot data of the second node in the shared cluster, filter and save the valid generation information of the second node according to the root field in the slot data of the second node, and count the number of consecutive occurrences of read / write input / output exceptions;
[0054] Determine whether the number of consecutive occurrences of read / write input / output exceptions in the slot data of the second node is greater than the first threshold. If so, control the second node to go offline. If not, verify the slot data of the next node.
[0055] In yet another aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0056] Obtain all nodes in the shared cluster connected to the heartbeat disk, and store the read / write statistical data of each node during each disk heartbeat on the heartbeat disk as slot data. The read / write statistical data includes whether the read / write input / output is normal or abnormal, and the number of consecutive occurrences of read / write input / output exceptions;
[0057] During each disk heartbeat of the heartbeat disk, in response to the first node receiving read / write input / output data, update the slot data of the first node, read the slot data of the second node in the shared cluster, filter and save the valid generation information of the second node according to the root field in the slot data of the second node, and count the number of consecutive occurrences of read / write input / output exceptions;
[0058] Determine whether the number of consecutive occurrences of read / write input / output exceptions in the slot data of the second node is greater than the first threshold. If so, control the second node to go offline. If not, verify the slot data of the next node.
[0059] The above distributed storage file system fault tolerance method, product, computer device, and storage medium adopt the method of interacting with all nodes in the shared cluster connected to the heartbeat disk to check whether the read / write input / output data is abnormal, and set that when the number of consecutive read / write input / output exceptions is greater than the first threshold, the corresponding node is controlled to go offline, thereby increasing the input / output fault tolerance mechanism and avoiding the problem that the entire distributed storage file system cluster becomes unavailable due to occasional input / output errors. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0061] Figure 1 It is a schematic diagram of the heartbeat mechanism of the OCFS2 file system in the prior art;
[0062] Figure 2 It is a schematic flowchart of the distributed storage file system fault tolerance method in an embodiment of the present application;
[0063] Figure 3 It is a logic diagram of the distributed storage file system fault tolerance method in an embodiment of the present application;
[0064] Figure 4 It is a schematic structural diagram of adding M groups of gen fields in the memory to temporarily store the status information of the gen fields in the node slot data in an embodiment of the present application;
[0065] Figure 5 It is a schematic flowchart of the step of writing the node generation information corresponding to normal read / write input / output into the first root field as the valid generation information of the second node by judging whether the second root field of the second node is greater than the third threshold in an embodiment of the present application;
[0066] Figure 6 It is a schematic flowchart of the step of setting the number of consecutive read / write input / output exceptions in the slot data of the second node by judging whether the third root field of the second node is greater than the second threshold in an embodiment of the present application;
[0067] Figure 7 It is a structural block diagram of the distributed storage file system fault tolerance device in an embodiment of the present application;
[0068] Figure 8 It is an internal structure diagram of a computer device in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0069] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0070] As described in the background art, OCFS2 is a distributed storage file system. Utilizing shared storage, it can be used for many general computing tasks that require shared cluster storage. In OCFS2, each node has a sector slot of its own on the shared heartbeat disk. Under normal circumstances, when a node joins the shared heartbeat disk, it writes a unique identity identifier "generation information" into its own slot, and then periodically writes information such as a timestamp into its own slot. At the same time, it periodically reads the slot data of other nodes to obtain the current status of other nodes.
[0071] As Figure 1 shown, taking a two-node cluster as an example, the heartbeat processing logic and the impact after an occasional heartbeat anomaly. Node A and Node B share a heartbeat disk for disk heartbeat. First, Node A updates the data in its own slot A, and then reads the data in slot B belonging to Node B. If all 0 data is read at this time due to a shared disk failure or high disk pressure, it will be considered that Node B is unavailable and isolated from the cluster. Similarly, Node B updates the data in its own slot B, and then reads the data in slot A belonging to Node A. After the data verification is correct, it is considered that Node A is available, and then a message is sent to Node A. At this time, in the view of Node A, Node B is no longer in the cluster, so the message sent by Node B to Node A will not be received, resulting in the message sent by Node B to Node A being blocked, thereby affecting the OCFS2 cluster function. The reason for the OCFS2 cluster function anomaly is that the existing heartbeat mechanism has too strict requirements for IO read and write and cannot tolerate occasional IO errors.
[0072] Therefore, under the current OCFS2 disk heartbeat mechanism, how to increase the IO fault tolerance mechanism to avoid the problem that the entire OCFS2 cluster becomes unavailable due to occasional IO errors is a technical problem that needs to be solved.
[0073] To solve the above problems, in the embodiments of the present invention, a creative method for fault tolerance of a distributed storage file system is proposed. The classification of heartbeat situations and the expected effects are shown in Table 1.
[0074] Table 1 Classification Table of Heartbeat Situations
[0075]
[0076] In the heartbeat count series, one heartbeat includes one write I / O and one read I / O.
[0077] The caseN column represents different heartbeat read I / O exception situations, where T indicates that the read I / O is correct, F indicates that the read I / O is abnormal, and - indicates to be ignored.
[0078] The result row indicates the impact in the current case. Among them, ok indicates that the cluster function is not affected, old indicates that the cluster function is affected under the traditional solution, tolerate indicates that the cluster function is not affected under the current solution, and fail indicates that the cluster function is affected under the current solution.
[0079] The detailed description is as follows:
[0080] Case1: If each disk heartbeat read I / O is normal, then the cluster function is normal;
[0081] Case2: Under the traditional solution, if a heartbeat read I / O is occasionally abnormal, then the cluster function is affected;
[0082] Case3 - 8: Under the current solution, if the consecutive heartbeat read I / O abnormalities do not exceed 2 times, then it can be tolerated and the cluster function is normal;
[0083] Case9: Under the current solution, if the consecutive heartbeat read I / O abnormalities exceed 2 times, then the cluster function is affected.
[0084] In one embodiment, as Figure 2 、 Figure 3 shown, a fault tolerance method for a distributed storage file system is provided, including the following steps:
[0085] Step S1, obtain all nodes in the shared cluster connected to the heartbeat disk, and store the read and write statistical data of each node in each disk heartbeat as slot data on the heartbeat disk. The read and write statistical data includes whether the read and write input / output is normal or abnormal, and the number of consecutive read and write input / output abnormalities;
[0086] Step S2, when each disk heartbeat occurs on the heartbeat disk, in response to the first node receiving the read and write input / output data, update the slot data of the first node, read the slot data of the second node in the shared cluster, filter and save the valid generation information of the second node according to the root field in the slot data of the second node, and count the number of consecutive read and write input / output abnormalities;
[0087] Step S3, determine whether the number of consecutive read and write input / output abnormalities in the slot data of the second node is greater than a first threshold. If so, control the second node to go offline. If not, verify the slot data of the next node.
[0088] Among them, by adopting all nodes in the shared cluster connected to the heartbeat disk to interact and check whether the read / write input / output data is abnormal, and setting that when the number of consecutive read / write input / output abnormalities exceeds a first threshold, the corresponding node is controlled to go offline, an input / output fault tolerance mechanism is added to avoid the problem that the entire distributed storage file system cluster becomes unavailable due to occasional input / output errors.
[0089] In this embodiment, the read / write statistical data of each node at each disk heartbeat stored on the heartbeat disk through slot data includes:
[0090] Set the initial value of the number of consecutive read / write input / output abnormalities to zero;
[0091] Judge whether the read / write input / output of the current node at the current disk heartbeat is normal or abnormal;
[0092] In response to the read / write input / output of the current node being normal at the current disk heartbeat, judge whether the read / write input / output of the current node is normal or abnormal during the next heartbeat read / write;
[0093] In response to the read / write input / output of the current node being abnormal at the current disk heartbeat, increment the number of consecutive read / write input / output abnormalities of the current node by one.
[0094] Among them, by continuously accumulating the number of times of abnormal read / write input / output, a first threshold is set to achieve IO exception fault tolerance. The first threshold represents the number of consecutive IO errors that can be tolerated, and is configured when the distributed storage file system (OCFS2) is started. It is required that this value is greater than 0. The larger this value is, the stronger the IO fault tolerance ability. It is recommended to configure the first threshold as 3. Implement Case3-8: In the current scheme, if the number of consecutive heartbeat read IO exceptions does not exceed 2 times, it can be tolerated and the cluster function is normal; Case9: In the current scheme, if the number of consecutive heartbeat read IO exceptions exceeds 2 times, the cluster function is affected.
[0095] In this embodiment, the judgment of whether the read / write input / output of the current node at the current disk heartbeat is normal or abnormal includes:
[0096] Set the generation information of each node stored on the heartbeat disk, and the read / write statistical data also includes the generation information of the current node performing the read / write;
[0097] Read the heartbeat read / write statistical data of the second node shared and connected with the first node, and judge whether the generation information of the current node performing the read / write in the read / write statistical data is the same as the generation information of the second node stored on the heartbeat disk. If they are different, it is determined that the current read / write input / output is abnormal. If they are the same, it is determined that the current read / write input / output is normal.
[0098] Among them, it is determined whether the current read and write input / output is normal or abnormal by using whether the generation information of the node during the read and write process is the same as the generation information of the node stored on the heartbeat disk, avoiding the interference of abnormal input / output data.
[0099] In this embodiment, determining whether the read and write input / output of the current node is normal or abnormal during the current disk heartbeat further includes:
[0100] Determine whether there is an error in the data being read and written in the read and write statistical data;
[0101] If there is no error and the generation information of the current node being read and written in the read and write statistical data is the same as the generation information of the second node stored on the heartbeat disk, it is determined that the current read and write input / output is normal;
[0102] If there is an error and the generation information of the current node being read and written in the read and write statistical data is the same as the generation information of the second node stored on the heartbeat disk, the current read and write input / output error is ignored.
[0103] Combined with Table 1, where "-" indicates ignoring. At this time, ignoring the current read and write input / output error is the situation when there is an error in the data being read and written and the generation information of the current node being read and written in the read and write statistical data is the same as the generation information of the second node stored on the heartbeat disk, avoiding the interference of abnormal input / output data.
[0104] In this embodiment, reading the slot data of the second node in the shared cluster, screening and saving the valid generation information of the second node according to the root field in the slot data of the second node, and counting the number of consecutive read and write input / output exceptions includes:
[0105] Setting the slot data further includes a first root field, a second root field, a third root field, and a fourth root field. The initial value of the slot data is zero. The first root field is used to record the valid generation information of the corresponding node. The second root field is used to record the number of consecutive times the generation information of the corresponding node is the same when read continuously. The third root field is used to record the number of consecutive times the generation information of each node is different. The fourth root field is used to record whether the generation information of each node has been authenticated as legal;
[0106] In response to reading the slot data of the second node, obtain the value of the fourth root field of the second node. If the value of the fourth root field of the second node is zero, it is determined that the first node has not saved the valid generation information of the second node. If the value of the fourth root field of the second node is not zero, it is determined that the first node has saved the valid generation information of the second node;
[0107] In response to the first node not saving the valid generation information of the second node, obtain the second root field value of the second node, and determine whether the second root field of the second node is greater than a third threshold to write the node generation information corresponding to normal read / write input / output into the first root field as the valid generation information of the second node;
[0108] In response to the first node having saved the generation information of the second node, obtain the third root field value of the second node, and set the number of consecutive occurrences of abnormal read / write input / output in the slot data of the second node by determining whether the third root field of the second node is greater than a second threshold.
[0109] Specifically, as Figure 3 shown, in fact, the distributed storage file system (OCFS2) periodically checks the node slot data to obtain the status of each node. Here, an example of checking the node slot data once is used for illustration.
[0110] First, the node updates its own slot data, then reads the slot data of other nodes, and then sequentially verifies the slot data of other nodes. If the gen field in the slot is continuously incorrect for the first threshold, the node is considered unavailable and is kicked out of the cluster; otherwise, it is determined whether the slot data of other nodes has been verified. If not, continue to verify the gen field in the slot data of the next node, otherwise this disk heartbeat ends.
[0111] Among them, the first root field, the second root field, the third root field, and the fourth root field are gen fields. The first root field is gen_record, the second root field is gen_equal, the third root field is gen_change, and the fourth root field is gen_valid.
[0112] As Figure 4 shown, add M groups of gen fields in the memory to temporarily store the status information of the gen fields in the node slot data. Here, M represents the maximum number of nodes supported in the ocfs2 cluster. When the number of nodes is 16, M is 15, and the slot data of each node is slot_0, slot_1,..., slot_16.
[0113] In this embodiment, the fourth root field is used to record whether the generation information of each node has been authenticated as legal, including:
[0114] Set the fourth root field to zero or one;
[0115] In response to obtaining that the fourth root field is one, determine that the generation information of the corresponding node has been authenticated as legal;
[0116] When it is determined that the fourth root field is zero, it is determined that the generation information of the corresponding node is not authenticated as legal.
[0117] Combined with Table 1, in Table 1, when the fourth root field is zero, if the generation information of the corresponding node is not authenticated as legal, it shows false; when the fourth root field is one, if the generation information of the corresponding node is authenticated as legal, it shows true.
[0118] As Figure 5 shown, in this embodiment, the method of writing the generation information of the corresponding node when the read-write input / output is normal into the first root field as the valid generation information of the second node by determining whether the second root field of the second node is greater than the third threshold includes:
[0119] Determine whether the value of the second root field of the second node is greater than the third threshold. If it is, set the fourth root field to one; if not, determine whether the second root field of the second node is zero.
[0120] If the second root field of the second node is zero, read the generation information of the second node from the heartbeat disk, store the generation information of the second node read from the heartbeat disk into the first root field, and at the same time increment the second root field by one.
[0121] If the second root field of the second node is not zero, indicating that the generation information of the second node has been stored in the first root field, then determine whether the value of the first root field of the second node in the slot data is the same as the generation information of the second node read from the heartbeat disk. If they are the same, increment the second root field by one; if they are different, save the generation information of the second node read from the heartbeat disk into the first root field, and at the same time set the second root field to one.
[0122] If the value of the second root field of the second node after incrementing is greater than the third threshold, set the fourth root field to one, indicating that the generation information of the second node is authenticated as legal.
[0123] As Figure 5 shown, in this embodiment, when the value of the second root field of the second node is less than or equal to the third threshold, before determining whether the second root field of the second node is zero, it further includes:
[0124] Determine whether the generation information of the second node read from the heartbeat disk is zero.
[0125] If it is zero, end; if it is not zero, determine whether the second root field of the second node is zero.
[0126] Among them, judging whether the generation information of the second node read from the heartbeat disk is zero can avoid misjudgment when the generation information of the second node is zero, and can distinguish the information with zero generation information from the initially set zero value.
[0127] It can be understood that in order to avoid confusion with the initially set zero value caused by whether the generation information of the node is zero, the present application can further set the generation information of each node to a non-zero value.
[0128] Specifically, the distributed storage file system fault tolerance method further includes:
[0129] Sort all the nodes in the shared cluster connected to the heartbeat disk and obtain the serial number of each node;
[0130] Set the generation information of each node to a non-zero value;
[0131] Set multiple slots on the heartbeat disk in sequence according to the serial number of each node, and slot data corresponding to one node is stored in each slot, where the number of slots is equal to the number of nodes;
[0132] In response to the slot data in the slot having a storage data exception, create a new slot on the heartbeat disk to replace the slot with the storage data exception;
[0133] In response to the target node going offline, clear or delete the slot data of the slot corresponding to the target node on the heartbeat disk;
[0134] In response to adding a new node to the shared cluster, create a new slot on the heartbeat disk to store the slot data of the new node.
[0135] Among them, setting the number of slots equal to the number of nodes by using the heartbeat disk. When the total number of all nodes in the shared cluster is M, the number of slots is M, which realizes synchronous management of the slot data and the corresponding nodes, avoids generating a large amount of invalid data on the heartbeat disk, and at the same time setting the generation information of each node to a non-zero value can avoid confusion with the initially set zero value caused by whether the generation information of the node is zero.
[0136] Such as Figure 5As shown, check the gen_valid field. If it is true, it means that the generation field of this slot has been successfully cached. Exit directly. Otherwise, check if gen_equal is greater than the third threshold (indicating how many consecutive times the same generation batch value is read before it is considered legal. It is configured when ocfs2 starts and is required to be greater than 1. The purpose of configuring the third threshold is to debounce. Only when the generation value is stable is it considered a legal value). If gen_equal is greater than STABLE_CNT, the generation value is considered legal and gen_valid is set to true. Otherwise, check if the generation value gen_ondisk read from the heartbeat disk is 0. If it is 0, exit. Otherwise, further check the cached gen_equal value. If gen_equal is 0, it means that the generation value has not started to be cached yet. Cache the gen_ondisk read from the heartbeat disk into gen_record, and at the same time increment the gen_equal count by 1, indicating the number of valid times of the generation value. If the generation value (gen_equal) is not 0, it means that the generation value has been cached but is not yet considered a legal value. It is necessary to further compare gen_ondisk and gen_record. If the two are equal, increment the gen_equal count by 1. Otherwise, overwrite the content in gen_record with the latest generation value gen_ondisk read from the heartbeat disk, and at the same time increment the gen_equal count by 1.
[0137] As Figure 6 shown, in this embodiment, setting the number of consecutive read / write input / output exceptions in the slot data of the second node by determining whether the third root field of the second node is greater than the second threshold includes:
[0138] If the value of the fourth root field of the second node is 1, determine whether the value of the first root field of the second node in the slot data is the same as the generation information of the second node read from the heartbeat disk, and simultaneously determine whether the value of the third root field of the second node is 0;
[0139] If the value of the first root field of the second node in the slot data is the same as the generation information of the second node read from the heartbeat disk, and the value of the third root field of the second node is not 0 and less than the second threshold, set the third root field to 0;
[0140] If the value of the first root field of the second node in the slot data is different from the generation information of the second node read from the heartbeat disk, increment the count of the third root field by one, and determine whether the incremented third root field is greater than the second threshold. If so, increment the number of consecutive read / write input / output exceptions in the slot data of the second node by one. If not, end the process.
[0141] As Figure 6 shown, this solution can tolerate N consecutive (first threshold) heartbeat IO errors. For non-consecutive heartbeat IO errors (less than N times), once a normal heartbeat IO is read, the heartbeat IO error count needs to be cleared to prevent the accumulation of heartbeat IO errors. The gen check process involves the following several judgments: Determine whether the gen_valid field is true. If true, it means the data temporarily stored in gen_record is valid; determine whether the generation value gen_ondisk read from the heartbeat disk is consistent with the temporarily stored gen_record. If consistent, it means the current heartbeat IO is normal; determine whether the temporarily stored gen_change is 0. If 0, it means the previous several heartbeat IOs are all normal, that is, the gen check has passed.
[0142] If gen_valid is true, gen_ondisk and gen_record are consistent, and gen_change is not equal to 0, it means the temporarily stored generation value is valid, the current heartbeat IO is normal, and the previous heartbeat IOs are also normal. Then clear the gen_change count and end the current gen check process.
[0143] Furthermore, if gen_ondisk and gen_record are consistent, it means the current heartbeat IO is normal, and end the current gen check process. Otherwise, consider the current heartbeat IO abnormal, and increment the gen_change count by 1. Then determine whether gen_change is greater than the first threshold (the number of heartbeat IO errors that can be tolerated, corresponding to 'N'). If greater, it means the number of consecutive heartbeat IO errors has reached the maximum tolerable number. Then clear gen_record, gen_change, and gen_equal to 0, set gen_valid to false, and at the same time trigger a node offline event, and the current gen check ends.
[0144] As Figure 2 shown, in this embodiment, when controlling the second node to go offline, it further includes:
[0145] Step S4, clear the slot data of the second node and trigger a command to kick or isolate the second node from the shared cluster.
[0146] Among them, the slot data of the second node is cleared. The purpose of resetting gen_record and gen_valid here is to be compatible with the scenario of the node going online again. Because when the node goes online again, a new generation value will be generated again.
[0147] As Figure 2 shown, in this embodiment, the distributed storage file system fault tolerance method further includes:
[0148] Step S5, before the next disk heartbeat of the heartbeat disk, determine whether the slot data of all nodes has been verified. If so, end; if not, verify the slot data of the next node.
[0149] Among them, by judging whether the slot data of all other nodes has been verified at each node, interactive check can be realized, avoiding the malfunction of the overall cluster nodes caused by occasional read and write errors of multiple nodes themselves.
[0150] In this embodiment, the distributed storage file system fault tolerance method further includes:
[0151] In response to the first node completing the verification of the slot data of all nodes, control the second node to receive read and write input / output data, update the slot data of the second node, read and verify the slot data of the third node in the shared cluster until all nodes in the shared cluster connected to the heartbeat disk complete receiving read and write input / output data.
[0152] Taking the process of this solution from the perspective of node A in a two-node cluster composed of node A and node B as an example.
[0153] Before OCFS2 starts, configure the first threshold to represent the number of consecutive IO errors that can be tolerated, the second threshold to represent the number of consecutive different node generation information that can be tolerated in gen_change (the consecutive number of times inconsistent with the determined node generation information), and the third threshold to represent how many times the same generation information value is continuously read in gen_equal before it is considered legal. Preferably, configure the total number of cluster nodes to be 16, the first threshold to be 3, the second threshold to be 2, and the third threshold to be 2.
[0154] Under this solution, 16 groups of gen-related fields are added in memory to temporarily store the status information of the gen field in the node slot. In fact, only 2 groups are used, and the rest are used as redundancy for subsequent expansion. The gen-related fields include the following four: gen_record, gen_equal, gen_change, and gen_valid. gen_record is used to record the valid (i.e., legal) generation information in the node slot. gen_equal is used to record the number of times the same generation is continuously read in the node slot. gen_change is used to record the number of times an illegal (inconsistent with the determined generation information) generation is continuously read in the node slot. gen_valid is used to record the legality of the data in the aforementioned gen_record. If it is legal, it is recorded as true; otherwise, it is recorded as false. When ocfs2 starts, the above values are all 0.
[0155] The first heartbeat.
[0156] First, node A updates its own slot data and writes 0x11111111. Then it reads the slot data of node B. Here, it is assumed that the generation value in the slot of node B is 0xbbbbbbbb. Then it verifies the slot data of node B, that is, the following steps.
[0157] First, check the gen_valid field. At this time, the gen_valid field is 0 (i.e., false), indicating that there is no legal generation field of node B's slot temporarily stored yet. Then check gen_equal. At this time, gen_equal is 0, which is less than the third threshold 2. Then further judge that the generation value gen_ondisk read from the heartbeat disk is 0xbbbbbbbb, not 0. Then check that gen_equal is 0. At this time, the value 0xbbbbbbbb of gen_ondisk read from the heartbeat disk is temporarily stored in gen_record, that is, the gen_record field changes from 0 to 0xbbbbbbbb, and at the same time, the gen_equal count is incremented by 1, that is, the gen_equal field changes from 0 to 1. All node slots have been verified, and the loop ends. At this time, gen_record is 0xbbbbbbbb, gen_equal is 1, gen_change is 0, and the gen_vaild field is 0.
[0158] The second heartbeat.
[0159] First, node A updates its own slot data, writes 0x11111112, then reads the slot data of node B. Here, it is assumed that the generation value in the slot of node B is still 0xbbbbbbbb. Then, it verifies the slot data of node B. The method is the same as the previous checking process. After the loop ends, at this time, gen_record is 0xbbbbbbbb, gen_equal is 2, gen_change is 0, and the gen_vaild field is 0.
[0160] The third heartbeat.
[0161] First, node A updates its own slot data, writes 0x11111113, then reads the slot data of node B. Here, it is assumed that the generation value in the slot of node B is still 0xbbbbbbbb. Then, it verifies the slot data of node B. The method is the same as the previous checking process. After the loop ends, at this time, gen_record is 0xbbbbbbbb, gen_equal is 3, gen_change is 0, and the gen_vaild field is 0.
[0162] The fourth heartbeat.
[0163] First, node A updates its own slot data, writes 0x11111114, then reads the slot data of node B. Here, it is assumed that the generation value in the slot of node B is still 0xbbbbbbbb. Then, it verifies the slot data of node B, which is the following process.
[0164] First, check the gen_valid field. At this time, the gen_valid field is 0 (i.e., false), indicating that there is no legal generation field for the temporarily stored slot of node B yet. Then, check gen_equal. At this time, gen_equal is 3, which is greater than the third threshold of 2. Then, set gen_valid to true, indicating that the value in gen_record is already legal at this time. At this time, gen_record is 0xbbbbbbbb, gen_equal is 3, gen_change is 0, and the gen_vaild field is 1.
[0165] For the first three heartbeats, since gen_valid is false, the process of checking gen (i.e., Figure 5 process) will not be executed. But starting from the fourth heartbeat, the gen_valid field is 1 (i.e., true), and the process of recording the valid gen (i.e., Figure 4 process) will no longer be carried out, but the process of checking gen (i.e., Figure 5) It will be executed, and since then, the heartbeat can effectively perform IO fault tolerance.
[0166] Here, it is assumed that during the Xth heartbeat (any heartbeat during the normal operation of ocfs2), the heartbeat IO is abnormal, and the generation value read from the slot of node B becomes 0. During the (X + 1)th heartbeat, the IO returns to normal, and the generation value read from the slot of node B is 0xbbbbbbbb, which is the scenario of case 3 in Table 1.
[0167] The Xth heartbeat.
[0168] It is checked and found that gen_valid is true, the gen_ondisk read from the heartbeat disk is 0, which is inconsistent with the effectively cached gen_record. Then, the gen_change count is incremented by 1. At this time, the gen_change count value is 1, which is not greater than the first threshold MAX_TOLERATE, so the node offline time will not be triggered, and the IO of this heartbeat is effectively fault-tolerant without affecting the cluster function.
[0169] The (X + 1)th heartbeat.
[0170] It is checked and found that gen_valid is true, the gen_ondisk read from the heartbeat disk is 0xbbbbbbbb, which is consistent with the effectively cached gen_record, and gen_change is not 0. Then, the gen_change count value is cleared to 0, and the gen check process ends.
[0171] Through the description of this embodiment, it can be found that this solution can effectively tolerate heartbeat IO errors.
[0172] In the above distributed storage file system fault tolerance method, by using all nodes in the shared cluster connected to the heartbeat disk to interact and check whether the read and write input and output data is abnormal, and setting that when the number of consecutive read and write input and output exceptions is greater than the first threshold, the corresponding node is controlled to go offline, an input and output fault tolerance mechanism is added to avoid the problem that the entire distributed storage file system cluster becomes unavailable due to occasional input and output errors.
[0173] It should be understood that although Figures 2 - 6 the steps in the flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figures 2 - 6At least a part of the steps therein may include multiple sub - steps or multiple stages. These sub - steps or stages are not necessarily executed and completed at the same moment, but can be executed at different moments. The execution order of these sub - steps or stages is not necessarily sequential, but can be executed alternately or in turns with at least a part of other steps or sub - steps or stages of other steps.
[0174] In one embodiment, as Figure 7 shown, a distributed storage file system fault - tolerance device 10 is provided, including: a node slot data management module 1, a slot data verification module 2, and an abnormal node offline control module 3.
[0175] The node slot data management module 1 is used to obtain all nodes in the shared cluster connected to the heartbeat disk, and store the read - write statistical data of each node during each disk heartbeat on the heartbeat disk as slot data. The read - write statistical data includes whether the read - write input / output is normal or abnormal, and the number of consecutive occurrences of abnormal read - write input / output.
[0176] The slot data verification module 2 is used to, during each disk heartbeat of the heartbeat disk, in response to the first node receiving read - write input / output data, update the slot data of the first node, read the slot data of the second node in the shared cluster, screen and save the valid generation information of the second node according to the root field in the slot data of the second node, and count the number of consecutive occurrences of abnormal read - write input / output.
[0177] The abnormal node offline control module 3 is used to determine whether the number of consecutive occurrences of abnormal read - write input / output in the slot data of the second node is greater than a first threshold. If so, control the second node to go offline; if not, verify the slot data of the next node.
[0178] In this embodiment, storing the read - write statistical data of each node during each disk heartbeat on the heartbeat disk through slot data includes:
[0179] Setting the initial value of the number of consecutive occurrences of abnormal read - write input / output to zero;
[0180] Judging whether the read - write input / output of the current node is normal or abnormal during the current disk heartbeat;
[0181] In response to the read - write input / output of the current node being normal during the current disk heartbeat, judging whether the read - write input / output of the current node is normal or abnormal during the next heartbeat read - write;
[0182] In response to the read - write input / output of the current node being abnormal during the current disk heartbeat, incrementing the number of consecutive occurrences of abnormal read - write input / output of the current node by one.
[0183] In this embodiment, the determination of whether the read / write input / output of the current node is normal or abnormal during the current disk heartbeat includes:
[0184] Set the generation information of each node to be stored on the heartbeat disk, and the read / write statistical data also includes the generation information of the current node performing the read / write;
[0185] Read the heartbeat read / write statistical data of the second node sharing a connection with the first node, and determine whether the generation information of the current node performing the read / write in the read / write statistical data is the same as the generation information of the second node stored on the heartbeat disk. If they are different, it is determined that the current read / write input / output is abnormal; if they are the same, it is determined that the current read / write input / output is normal.
[0186] In this embodiment, the determination of whether the read / write input / output of the current node is normal or abnormal during the current disk heartbeat further includes:
[0187] Determine whether there is an error in the data being read / written in the read / write statistical data;
[0188] If there is no error and the generation information of the current node performing the read / write in the read / write statistical data is the same as the generation information of the second node stored on the heartbeat disk, it is determined that the current read / write input / output is normal;
[0189] If there is an error and the generation information of the current node performing the read / write in the read / write statistical data is the same as the generation information of the second node stored on the heartbeat disk, the current read / write input / output error is ignored.
[0190] In this embodiment, the reading of the slot data of the second node in the shared cluster, the screening and saving of the valid generation information of the second node according to the root field in the slot data of the second node, and the counting of the number of consecutive occurrences of read / write input / output anomalies include:
[0191] Set the slot data to further include a first root field, a second root field, a third root field, and a fourth root field. The initial value of the slot data is zero. The first root field is used to record the valid generation information of the corresponding node, the second root field is used to record the number of consecutive times the generation information of the corresponding node is read continuously and is the same, the third root field is used to record the number of consecutive times the generation information of each node is different, and the fourth root field is used to record whether the generation information of each node has been authenticated as legal;
[0192] In response to reading the slot data of the second node, obtain the value of the fourth root field of the second node. If the value of the fourth root field of the second node is zero, it is determined that the first node has not saved the valid generation information of the second node; if the value of the fourth root field of the second node is not zero, it is determined that the first node has saved the valid generation information of the second node;
[0193] In response to the first node not saving the valid generation information of the second node, obtain the second root field value of the second node, and determine whether the second root field of the second node is greater than a third threshold to write the node generation information corresponding to normal read / write input / output into the first root field as the valid generation information of the second node;
[0194] In response to the first node having saved the generation information of the second node, obtain the third root field value of the second node, and determine whether the third root field of the second node is greater than a second threshold to set the number of consecutive abnormal read / write input / output occurrences in the slot data of the second node.
[0195] In this embodiment, the fourth root field is used to record whether the generation information of each node has been authenticated as legal, including:
[0196] Set the fourth root field to zero or one;
[0197] In response to obtaining that the fourth root field is one, determine that the generation information of the corresponding node has been authenticated as legal;
[0198] In response to obtaining that the fourth root field is zero, determine that the generation information of the corresponding node has not been authenticated as legal.
[0199] In this embodiment, the step of determining whether the second root field of the second node is greater than a third threshold to write the node generation information corresponding to normal read / write input / output into the first root field as the valid generation information of the second node includes:
[0200] Judge whether the value of the second root field of the second node is greater than the third threshold. If so, set the fourth root field to one. If not, judge whether the second root field of the second node is zero;
[0201] If the second root field of the second node is zero, read the generation information of the second node from the heartbeat disk and store the generation information of the second node read from the heartbeat disk into the first root field, and at the same time increment the second root field by one;
[0202] If the second root field of the second node is not zero, indicating that the generation information of the second node has been stored in the first root field, then judge whether the value of the first root field of the second node in the slot data is the same as the generation information of the second node read from the heartbeat disk. If they are the same, increment the second root field by one. If they are different, save the generation information of the second node read from the heartbeat disk to the first root field, and at the same time set the second root field to one;
[0203] If the value after adding one to the second root field of the second node is greater than the third threshold, set the fourth root field to one, indicating that the generation information of the second node has been authenticated as legal.
[0204] In this embodiment, when the value of the second root field of the second node is less than or equal to the third threshold, before determining whether the second root field of the second node is zero, it further includes:
[0205] Determine whether the generation information of the second node read from the heartbeat disk is zero;
[0206] If it is zero, end; if it is not zero, determine whether the second root field of the second node is zero.
[0207] In this embodiment, the method of setting the number of consecutive read / write input / output exceptions in the slot data of the second node by determining whether the third root field of the second node is greater than the second threshold includes:
[0208] If the value of the fourth root field of the second node is one, determine whether the value of the first root field of the second node in the slot data is the same as the generation information of the second node read from the heartbeat disk, and simultaneously determine whether the value of the third root field of the second node is zero;
[0209] If the value of the first root field of the second node in the slot data is the same as the generation information of the second node read from the heartbeat disk, and the value of the third root field of the second node is not zero and less than the second threshold, set the third root field to zero;
[0210] If the value of the first root field of the second node in the slot data is different from the generation information of the second node read from the heartbeat disk, increment the count of the third root field by one, determine whether the incremented third root field is greater than the second threshold, if so, increment the number of consecutive read / write input / output exceptions in the slot data of the second node by one, if not, end.
[0211] In this embodiment, as Figure 7 shown, the distributed storage file system fault tolerance device 10 further includes: a slot data clearing module 4.
[0212] When controlling the second node to go offline, the slot data clearing module 4 is used to clear the slot data of the second node and trigger a command to kick or isolate the second node from the shared cluster.
[0213] In this embodiment, as Figure 7 shown, the distributed storage file system fault tolerance device 10 further includes: a full node verification and management module 5.
[0214] The above-mentioned all-node verification management module 5 is used to determine whether the slot data of all nodes has been verified before the heartbeat disk performs the next disk heartbeat. If so, the process ends; otherwise, it verifies the slot data of the next node.
[0215] In this embodiment, as Figure 7 shown, the distributed storage file system fault tolerance device 10 further includes: a cyclic verification control module 6.
[0216] After the cyclic verification control module 6 responds to the first node to complete the verification of the slot data of all nodes, it controls the second node to receive read / write input / output data, updates the slot data of the second node, reads and verifies the slot data of the third node in the shared cluster until all nodes in the shared cluster connected to the heartbeat disk complete receiving read / write input / output data.
[0217] In the above-mentioned distributed storage file system fault tolerance device, all nodes in the shared cluster connected to the heartbeat disk are used to interactively check whether the read / write input / output data is abnormal, and when the number of consecutive occurrences of read / write input / output exceptions is greater than the first threshold, the corresponding node is controlled to go offline, increasing the input / output fault tolerance mechanism and avoiding the problem that the entire distributed storage file system cluster becomes unavailable due to occasional input / output errors.
[0218] For the specific limitations of the distributed storage file system fault tolerance device, reference can be made to the limitations of the distributed storage file system fault tolerance method in the above text, which will not be elaborated here. Each module in the above-mentioned distributed storage file system fault tolerance device can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0219] In one embodiment, a computer device is provided. This computer device can be a server, and its internal structure diagram can be as Figure 8 shown. This computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of this computer device is used to provide computing and control capabilities. The memory of this computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of this computer device is used to store distributed storage file system fault tolerance data. The network interface of this computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a distributed storage file system fault tolerance method.
[0220] Those skilled in the art can understand that Figure 8 the structure shown in Figure 8 is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0221] In one embodiment, a computer program product is provided, including a computer program, which when executed by a processor implements the following steps:
[0222] Obtain all nodes in the shared cluster connected to the heartbeat disk, and store the read-write statistical data of each node during each disk heartbeat on the heartbeat disk as slot data. The read-write statistical data includes whether the read-write input / output is normal or abnormal, and the number of consecutive occurrences of abnormal read-write input / output;
[0223] During each disk heartbeat of the heartbeat disk, in response to the first node receiving read-write input / output data, update the slot data of the first node, read the slot data of the second node in the shared cluster, filter and save the valid generation information of the second node according to the root field in the slot data of the second node, and count the number of consecutive occurrences of abnormal read-write input / output;
[0224] Judge whether the number of consecutive occurrences of abnormal read-write input / output in the slot data of the second node is greater than a first threshold. If so, control the second node to go offline. If not, check the slot data of the next node.
[0225] For the specific limitations on the steps implemented when the computer program is executed by the processor, reference can be made to the limitations on the method for fault tolerance of the distributed storage file system in the above text, which will not be elaborated here.
[0226] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:
[0227] Obtain all nodes in the shared cluster connected to the heartbeat disk, and store the read-write statistical data of each node during each disk heartbeat on the heartbeat disk as slot data. The read-write statistical data includes whether the read-write input / output is normal or abnormal, and the number of consecutive occurrences of abnormal read-write input / output;
[0228] Upon each disk heartbeat of the heartbeat disk, in response to the first node receiving read / write input / output data, update the slot data of the first node, read the slot data of the second node in the shared cluster, filter and save the valid generation information of the second node according to the root field in the slot data of the second node, and count the number of consecutive read / write input / output exceptions;
[0229] Determine whether the number of consecutive read / write input / output exceptions in the slot data of the second node is greater than a first threshold. If so, control the second node to go offline; if not, verify the slot data of the next node.
[0230] For the specific limitations on the steps implemented when the processor executes the computer program, reference can be made to the limitations on the method for fault tolerance of the distributed storage file system in the foregoing text, which will not be elaborated here.
[0231] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0232] Obtain all nodes in the shared cluster connected to the heartbeat disk, and store the read / write statistical data of each node during each disk heartbeat on the heartbeat disk as slot data. The read / write statistical data includes whether the read / write input / output is normal or abnormal, and the number of consecutive read / write input / output exceptions;
[0233] Upon each disk heartbeat of the heartbeat disk, in response to the first node receiving read / write input / output data, update the slot data of the first node, read the slot data of the second node in the shared cluster, filter and save the valid generation information of the second node according to the root field in the slot data of the second node, and count the number of consecutive read / write input / output exceptions;
[0234] Determine whether the number of consecutive read / write input / output exceptions in the slot data of the second node is greater than a first threshold. If so, control the second node to go offline; if not, verify the slot data of the next node.
[0235] For the specific limitations on the steps implemented when the computer program is executed by the processor, reference can be made to the limitations on the method for fault tolerance of the distributed storage file system in the foregoing text, which will not be elaborated here.
[0236] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0237] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0238] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application.
Claims
1. A distributed storage file system fault tolerance method, characterized in that: include: Obtain all nodes in the shared cluster connected to the heartbeat disk, and store the read and write statistics of each node in each disk heartbeat as slot data on the heartbeat disk, wherein the read and write statistics include whether the read and write input and output are normal or abnormal, and the number of consecutive read and write input and output abnormalities; At each disk heartbeat of the heartbeat disk, in response to the first node receiving read / write input / output data, the slot data of the first node is updated, the slot data of the second node in the shared cluster is read, the valid generation information of the second node is filtered and saved according to the root field in the slot data of the second node, and the number of consecutive read / write input / output anomalies is counted; the generation information is a unique identifier written to the slot where the node is located when it joins the shared heartbeat disk; when the node is back online, a new generation information value is regenerated; Determine whether the number of consecutive read / write input / output anomalies in the slot data of the second node is greater than a first threshold. If so, control the second node to go offline. Otherwise, check the slot data of the next node.
2. The distributed storage file system fault tolerance method according to claim 1, characterized in that: The storage of read and write statistics of each node in each disk heartbeat through slot data on the heartbeat disk includes: Setting the initial value of the number of consecutive read / write input / output exceptions to zero; Determine whether the read and write input and output of the current node is normal or abnormal during the current disk heartbeat; In response to the current node reading and writing input and output being normal during the current disk heartbeat, determining whether the current node reading and writing input and output is normal or abnormal during the next heartbeat reading and writing; In response to a read / write input / output exception of the current node during a current disk heartbeat, the number of consecutive read / write input / output exceptions of the current node is increased by one.
3. The distributed storage file system fault tolerance method according to claim 2, characterized in that: The determination of whether the read / write input / output of the current node is normal or abnormal during the current disk heartbeat includes: The generation information of each node is stored on the heartbeat disk, and the read and write statistics also include the generation information of the current node performing the reading and writing; Read the heartbeat read and write statistics of the second node that shares a connection with the first node, and determine whether the generation information of the current node performing reading and writing in the read and write statistics is the same as the generation information of the second node stored on the heartbeat disk; if they are different, determine that the current read and write input and output are abnormal; if they are the same, determine that the current read and write input and output are normal.
4. The distributed storage file system fault tolerance method according to claim 3, characterized in that: The determining whether the read / write input / output of the current node is normal or abnormal during the current disk heartbeat also includes: Determine whether there is an error in the data read and written in the read and write statistical data; If there is no error and the generation information of the current node being read and written in the read and write statistical data is the same as the generation information of the second node stored on the heartbeat disk, it is determined that the current read and write input and output are normal; If there is an error and the generation information of the current node being read and written in the read and write statistical data is the same as the generation information of the second node stored on the heartbeat disk, the current read and write input and output error is ignored.
5. The distributed storage file system fault tolerance method according to claim 3, characterized in that: The reading of the slot data of the second node in the shared cluster, screening and saving the valid generation information of the second node according to the root field in the slot data of the second node, and counting the number of consecutive occurrences of read / write input / output anomalies includes: The slot data is set to further include a first root field, a second root field, a third root field and a fourth root field, the initial value of the slot data is zero, the first root field is used to record the valid generation information of the corresponding node, the second root field is used to record the number of consecutive identical generation information read from the corresponding node, the third root field is used to record the number of consecutive different generation information of each node, and the fourth root field is used to record whether the generation information of each node has been authenticated as legal; In response to reading the slot data of the second node, obtaining the value of the fourth root field of the second node, if the value of the fourth root field of the second node is zero, determining that the first node does not store the valid generation information of the second node, and if the value of the fourth root field of the second node is not zero, determining that the first node has stored the valid generation information of the second node; In response to the first node not storing the valid generation information of the second node, obtaining the value of the second root field of the second node, and writing the node generation information corresponding to the normal read and write input and output into the first root field as the valid generation information of the second node by judging whether the second root field of the second node is greater than a third threshold; In response to the first node having saved the generation information of the second node, the value of the third root field of the second node is obtained, and the number of consecutive read and write input and output exceptions in the slot data of the second node is set by determining whether the third root field of the second node is greater than a second threshold.
6. The distributed storage file system fault tolerance method according to claim 5, characterized in that: The fourth root field is used to record whether the generation information of each node has been authenticated as legal, including: Setting the fourth root field to zero or one; In response to obtaining that the fourth root field is one, determining that the generation information of the corresponding node has been authenticated as legal; In response to obtaining that the fourth root field is zero, it is determined that the generation information of the corresponding node is not authenticated as legal.
7. The distributed storage file system fault tolerance method according to claim 5, characterized in that: The step of judging whether the second root field of the second node is greater than a third threshold value to write the node generation information corresponding to the normal read / write input and output into the first root field as the valid generation information of the second node comprises: Determine whether the value of the second root field of the second node is greater than a third threshold, if so, set the fourth root field to one, if not, determine whether the second root field of the second node is zero; If the second root field of the second node is zero, the generation information of the second node is read from the heartbeat disk and stored in the first root field, and the second root field is incremented by one; If the second root field of the second node is not zero, indicating that the generation information of the second node has been stored in the first root field, then determine whether the value of the first root field of the second node in the slot data is the same as the generation information of the second node read from the heartbeat disk, if they are the same, add one to the second root field, if they are different, save the generation information of the second node read from the heartbeat disk to the first root field, and set the second root field to one; If the value of the second root field of the second node plus one is greater than the third threshold, the fourth root field is set to one.
8. The distributed storage file system fault tolerance method according to claim 7, characterized in that: When the value of the second root field of the second node is less than or equal to a third threshold, before determining whether the second root field of the second node is zero, the method further includes: Determine whether the generation information of the second node read from the heartbeat disk is zero; If it is zero, end; if it is not zero, determine whether the second root field of the second node is zero.
9. The distributed storage file system fault tolerance method according to claim 5, characterized in that: The step of setting the number of consecutive read / write input / output exceptions in the slot data of the second node by judging whether the third root field of the second node is greater than a second threshold value comprises: If the value of the fourth root field of the second node is one, determine whether the value of the first root field of the second node in the slot data is the same as the generation information of the second node read from the heartbeat disk, and simultaneously determine whether the value of the third root field of the second node is zero; If the value of the first root field of the second node in the slot data is the same as the generation information of the second node read from the heartbeat disk, and the value of the third root field of the second node is not zero and is less than the second threshold, then the third root field is set to zero; If the value of the first root field of the second node in the slot data is different from the generation information of the second node read from the heartbeat disk, the third root field count is increased by one, and it is determined whether the third root field after increase is greater than the second threshold. If so, the number of consecutive read / write input / output exceptions in the slot data of the second node is increased by one, otherwise the process ends.
10. The distributed storage file system fault tolerance method according to claim 1, characterized in that: When controlling the second node to go offline, the method further includes: The slot data of the second node is cleared to trigger a command to kick out or isolate the second node from the shared cluster.
11. The distributed storage file system fault tolerance method according to claim 1, characterized in that: The method further comprises: Before the heartbeat disk performs the next disk heartbeat, it is determined whether the slot data of all nodes have been verified, and if so, the process ends; otherwise, the slot data of the next node is verified.
12. The distributed storage file system fault tolerance method according to claim 11, characterized in that: The method further comprises: In response to the first node completing the slot data verification of all nodes, the second node is controlled to receive read and write input and output data, the slot data of the second node is updated, and the slot data of the third node in the shared cluster is read and verified until all nodes in the shared cluster connected to the heartbeat disk complete receiving the read and write input and output data.
13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
14. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 12 are implemented.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Small computer system interface (SCSI) fault-tolerant optimization method and device based on hadoop distributed file system (HDFS)
CN103220162A
Distributed storage system data processing method, device and equipment and storage medium
CN111327685A