A fault processing method, device and equipment of a cluster node and a storage medium

CN116915798BActive Publication Date: 2026-09-11CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211582773.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-09
Publication Date
2026-09-11
Estimated Expiration
2042-12-09

AI Technical Summary

Technical Problem

针对该问题,目前尚无有效解决方案

Benefits of technology

[0036]本发明实施例提供一种集群节点的故障处理方法、装置、设备和存储介质。其中,所述方法包括:获取第一节点对应的第一数据,基于所述第一数据确定所述第一节点对应的网络状态;所述第一数据表征记录所述第一节点对应的数据是否成功同步至至少一个第二节点;所述第一节点为所述集群节点中的任一节点;所述第二节点为所述集群节点中除所述第一节点以外的任一节点;在所述网络状态表征所述第一节点对应的网络出现故障的情况下,获取所述第一节点对应的第二数据;所述第二数据表征记录与所述第一节点相关的设备的连接状态对应的目标数据;重新获取所述第一节点对应的第一数据,基于所述重新获取的第一数据重新确定所述第一节点对应的网络状态;在所述重新确定的网络状态表征所述第一节点对应的网络恢复正常的情况下,基于所述第二数据在所述至少一个第二节点中确定与所述第一节点对应的目标节点;向所述目标节点发送所述目标数据。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116915798B_ABST
    Figure CN116915798B_ABST
Patent Text Reader

Abstract

A cluster node fault processing method, device, equipment and storage medium are disclosed. The method comprises: obtaining first data corresponding to a first node, determining a network state corresponding to the first node based on the first data; the first data represents whether data corresponding to the first node is successfully synchronized to at least one second node; in the case that the network state represents that a network corresponding to the first node has a fault, obtaining second data corresponding to the first node; the second data represents target data corresponding to a connection state of a device related to the first node; re-obtaining the first data corresponding to the first node, and re-determining the network state corresponding to the first node based on the re-obtained first data; in the case that the re-determined network state represents that the network corresponding to the first node has returned to normal, determining a target node corresponding to the first node in the at least one second node based on the second data; and sending the target data to the target node.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of the Internet of Things (IoT), and more particularly to a method, apparatus, device, and storage medium for handling faults in cluster nodes. Background Technology

[0002] In IoT scenarios, connectivity middleware is responsible for handling the connection between various smart devices and the cloud platform, topic subscription, and message routing. Each middleware instance is called a node. In production environments, to improve device access volume and connection stability, multiple nodes deployed on different servers, each with its own independent computing and storage resources, are used to form a unified cluster that provides services to the outside world. All nodes in the cluster can communicate with each other via the network.

[0003] As long as the routing data between cluster nodes remains consistent, devices connected to the cluster can communicate with each other through cluster message forwarding. When a network failure occurs, normal communication between cluster nodes is impossible, and the routing data corresponding to newly established or disconnected device connections cannot be synchronized to other nodes in a timely manner. When the network failure is resolved, the routing data between cluster nodes will be inconsistent.

[0004] Existing technical solutions typically achieve data consistency across all nodes by restarting a minority of nodes to rejoin the cluster, or by selecting the node with the earliest startup time as the sole node in the new cluster, restarting other nodes, and having them rejoin the new cluster to synchronize historical data. However, restarting any node in the cluster disconnects all devices that were previously connected to that node. The more connections that are forcibly disconnected, the greater the impact on business operations. Currently, there is no effective solution to this problem. Summary of the Invention

[0005] To address the existing technical problems, embodiments of the present invention provide a method, apparatus, device, and storage medium for handling faults in cluster nodes.

[0006] To achieve the above objectives, the technical solution of this invention is implemented as follows:

[0007] In a first aspect, the present invention provides a method for handling faults in cluster nodes, the method comprising:

[0008] Obtain the first data corresponding to the first node, and determine the network status corresponding to the first node based on the first data; the first data indicates whether the data corresponding to the first node has been successfully synchronized to at least one second node; the first node is any node in the cluster nodes; the second node is any node in the cluster nodes other than the first node;

[0009] When the network status indicates that the network corresponding to the first node has failed, second data corresponding to the first node is obtained; the second data indicates target data corresponding to the connection status of the device related to the first node.

[0010] Reacquire the first data corresponding to the first node, and redetermine the network state corresponding to the first node based on the reacquired first data;

[0011] When the redefined network state indicates that the network corresponding to the first node has returned to normal, a target node corresponding to the first node is determined from the at least one second node based on the second data; and the target data is sent to the target node.

[0012] In the above scheme, the first data includes first statistical data on the successful synchronization of data corresponding to the first node to each of the second nodes and second statistical data on the failure of data corresponding to the first node to be successfully synchronized to each of the second nodes; determining the network status corresponding to the first node based on the first data includes:

[0013] If the first statistical data is greater than or equal to the first preset threshold and the second statistical data is less than or equal to the second preset threshold, the network status corresponding to the first node is determined to be normal.

[0014] If the first statistical data is less than the first preset threshold, and / or the second statistical data is greater than the second preset threshold, the network state corresponding to the first node is determined to be a fault state.

[0015] In the above scheme, determining the target node corresponding to the first node among the at least one second node based on the second data includes:

[0016] Obtain the first address data from the second data;

[0017] The target node corresponding to the first node is determined in the at least one second node using the first address data.

[0018] The method in the above scheme further includes:

[0019] Obtain the third data from the first node;

[0020] The third data is used to determine the target data corresponding to the connection status of the device associated with the first node;

[0021] The second data is updated based on the target data.

[0022] In the above scheme, the connection state includes a first state and a second state; the first state represents the state corresponding to the establishment of a connection between the first node and the device; the second state represents the state corresponding to the disconnection between the first node and the device; obtaining the second data corresponding to the first node includes:

[0023] Obtain first target data corresponding to the state where the first node establishes a connection with the device, and second target data corresponding to the state where the first node disconnects from the device;

[0024] The first target data and the second target data are used as the second data.

[0025] In the above scheme, after sending the target data to the target node, the method further includes:

[0026] If, after a preset time interval, the re-determined network state indicates that the network corresponding to the first node remains normal, the target data that has been sent to the target node in the second data is deleted.

[0027] The method in the above scheme further includes:

[0028] If the target node successfully receives the target data, the third data corresponding to the target node is updated using the target data.

[0029] Secondly, the present invention also provides a fault handling device for cluster nodes, the device comprising:

[0030] A determining unit is configured to acquire first data corresponding to a first node and determine the network status corresponding to the first node based on the first data; the first data indicates whether the data corresponding to the first node has been successfully synchronized to at least one second node; the first node is any node in the cluster nodes; the second node is any node in the cluster nodes other than the first node;

[0031] The acquisition unit is configured to acquire second data corresponding to the first node when the network status characterizes a network failure corresponding to the first node; the second data characterizes target data corresponding to the connection status of devices associated with the first node.

[0032] The determining unit is further configured to reacquire the first data corresponding to the first node, and re-determine the network state corresponding to the first node based on the reacquired first data.

[0033] The determining unit is further configured to, when the redefined network state indicates that the network corresponding to the first node has returned to normal, determine a target node corresponding to the first node among the at least one second node based on the second data; and send the target data to the target node.

[0034] Thirdly, the present invention also provides a fault handling device for a cluster node, the device comprising: a processor and a memory for storing a computer program capable of running on the processor, wherein the processor, when running the computer program, performs the steps of any of the methods described above.

[0035] Fourthly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of any of the methods described above.

[0036] This invention provides a method, apparatus, device, and storage medium for handling faults in cluster nodes. The method includes: acquiring first data corresponding to a first node; determining the network status of the first node based on the first data; the first data indicating whether data corresponding to the first node has been successfully synchronized to at least one second node; the first node being any node in the cluster; the second node being any node in the cluster other than the first node; if the network status indicates a network fault in the first node, acquiring second data corresponding to the first node; the second data indicating target data corresponding to the connection status of devices associated with the first node; reacquiring the first data corresponding to the first node; re-determining the network status of the first node based on the re-acquired first data; if the re-determined network status indicates that the network corresponding to the first node has returned to normal, determining a target node corresponding to the first node among the at least one second node based on the second data; and sending the target data to the target node.

[0037] By adopting the technical solution of this invention, when there is a network failure between clusters, the target data that the first node needs to synchronize with the target node after the network is restored to normal can be obtained, so as to restore the consistency of data between clusters without restarting cluster nodes or breaking existing connections. Attached Figure Description

[0038] Figure 1 A flowchart illustrating a fault handling method for a cluster node provided in an embodiment of the present invention;

[0039] Figure 2 This is a schematic diagram of the processing flow of a topic accumulator provided in an embodiment of the present invention;

[0040] Figure 3 A schematic diagram illustrating the workflow of a routing synchronization process provided in an embodiment of the present invention;

[0041] Figure 4 A schematic diagram of a processing flow for connecting a playback table is provided in an embodiment of the present invention;

[0042] Figure 5 A schematic diagram illustrating the workflow of a health check probe process provided in an embodiment of the present invention;

[0043] Figure 6 A schematic diagram of the workflow of a connection playback process provided in an embodiment of the present invention;

[0044] Figure 7 A schematic diagram of a fault handling device for a cluster node provided in an embodiment of the present invention;

[0045] Figure 8 This is a schematic diagram of a hardware structure for a fault handling device for a cluster node according to an embodiment of the present invention. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of this disclosure clearer, the technical solutions of this disclosure are further described in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on this disclosure. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0047] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0048] The terms “first / second / third” used in this disclosure are merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that “first / second / third” may be interchanged in a specific order or sequence where permitted, so that the embodiments of this disclosure described herein can be implemented in an order other than that illustrated or described herein.

[0049] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0050] This invention provides a method for handling faults in cluster nodes. Figure 1 This is a flowchart illustrating a fault handling method for a cluster node provided in an embodiment of the present invention. The method includes the following steps:

[0051] S101: Obtain the first data corresponding to the first node, and determine the network status corresponding to the first node based on the first data; the first data indicates whether the data corresponding to the first node has been successfully synchronized to at least one second node; the first node is any node in the cluster nodes; the second node is any node in the cluster nodes other than the first node.

[0052] It should be noted that the cluster in this embodiment of the invention can be any cluster, and is not limited here. For example, the cluster can be a unified cluster providing services to the outside world, composed of nodes with their own independent computing and storage resources in an IoT scenario. All nodes in the cluster can communicate with each other through the network. Nodes in the cluster are responsible for maintaining the subscription relationships of devices connected to that node and the routing data of that node, and are responsible for synchronizing the routing data to other nodes in the cluster. The subscription relationship can be understood as the topics subscribed to by devices connected to a node, for example: device client1 connected to node1 subscribes to topic1; the routing data of a node can be understood as the correspondence between nodes and subscribed topics, for example: node1 subscribes to topic1.

[0053] In this embodiment, the first data can be understood as any data related to the first node, without limitation. For example, the first data can be the routing synchronization data of the first node. Specifically, the routing synchronization data can include statistical results of successful synchronization and statistical results of failed synchronization. The network status corresponding to the first node can be understood as the network connection status between the first node and the corresponding cluster.

[0054] In practical applications, obtaining the first data corresponding to the first node can be understood as obtaining the routing synchronization data of the first node; determining the network status corresponding to the first node based on the first data can be understood as determining the network connection status between the first node and the corresponding cluster based on the routing synchronization data; the first data characterization recording whether the data corresponding to the first node has been successfully synchronized to at least one second node can be understood as the routing synchronization data recording the statistical results of whether the data synchronization from the first node to at least one second node has been successful or failed.

[0055] S102: When the network status indicates that the network corresponding to the first node has failed, obtain the second data corresponding to the first node; the second data indicates the target data corresponding to the connection status of the device related to the first node.

[0056] It should be noted that the second data can be any data related to the first node, without any limitation. For example, the second data can be connection playback data. The target data can be any data related to the connection status of devices associated with the first node, without any limitation. For example, the target data can be data from devices that have newly established or disconnected connections with the first node.

[0057] In practical applications, network status characterization of the network corresponding to the first node can be understood as the first node losing connection with the corresponding cluster; obtaining the second data corresponding to the first node can be understood as obtaining the connection replay data of the first node; the second data characterization of the target data corresponding to the connection status of the devices related to the first node can be understood as the connection replay data recording the data of the devices that have newly established or disconnected from the first node.

[0058] S103: Reacquire the first data corresponding to the first node, and redetermine the network state corresponding to the first node based on the reacquired first data.

[0059] In practical applications, reacquiring the first data corresponding to the first node can be understood as reacquiring the routing synchronization data of the first node; re-determining the network state corresponding to the first node based on the reacquiring first data can be understood as re-determining the network connection state between the first node and the corresponding cluster based on the reacquiring routing synchronization data.

[0060] S104: When the redefined network state indicates that the network corresponding to the first node has returned to normal, a target node corresponding to the first node is determined among the at least one second node based on the second data; the target data is sent to the target node.

[0061] It should be noted that the target node can be any node among at least one second node, without any limitation. For example, the target node can be a second node in the cluster that has received synchronization data from the first node.

[0062] In practical applications, the network status representation that the network corresponding to the first node has returned to normal can be understood as the first node restoring its connection with the corresponding cluster; determining the target node corresponding to the first node in at least one second node based on the second data can be understood as determining the second node that has received the synchronization data of the first node based on the connection playback data; sending target data to the target node can be understood as the first node sending data from the device that has newly established or disconnected its connection with the first node to the target node.

[0063] This invention provides a fault handling method for cluster nodes. The method determines the cluster network status corresponding to a node through the node's routing synchronization data. When the cluster network corresponding to a node fails, the routing information corresponding to the node's newly established and disconnected connections is recorded in the node's connection replay data. This forms the smallest subset of routing data that needs to be synchronized with other nodes after the network recovers, i.e., the target data. After the network recovers, the target node for synchronizing data is determined through the connection replay data, and the target data is sent to the target node. This achieves the restoration of data consistency between clusters without restarting cluster nodes or disconnecting existing connections. Furthermore, each node can easily and autonomously transmit its own routing data, reducing the complexity of data recovery.

[0064] In an optional embodiment of the present invention, the first data includes first statistical data on the successful synchronization of data corresponding to the first node to each of the second nodes and second statistical data on the failure of data corresponding to the first node to be successfully synchronized to each of the second nodes; the step of determining the network status corresponding to the first node based on the first data includes:

[0065] If the first statistical data is greater than or equal to the first preset threshold and the second statistical data is less than or equal to the second preset threshold, the network status corresponding to the first node is determined to be normal.

[0066] If the first statistical data is less than the first preset threshold, and / or the second statistical data is greater than the second preset threshold, the network state corresponding to the first node is determined to be a fault state.

[0067] It should be noted that the first statistical data can be the statistical result value of the first node's successful synchronization, and the second statistical data can be the statistical result value of the first node's failed synchronization. The first preset threshold and the second preset threshold can be determined according to the actual situation. The first preset threshold and the second preset threshold can be the same or different, which is not limited here. As an example, the first preset threshold and the second preset threshold can be the numerical values ​​corresponding to the synchronization statistical results. For example, both the first preset threshold and the second preset threshold are 0.

[0068] In practical applications, taking the example where both the first and second preset thresholds are 0, if the first statistical data is greater than or equal to the first preset threshold and the second statistical data is less than or equal to the second preset threshold, it can be understood that the statistical result of successful synchronization is greater than 0 and the statistical result of failed synchronization is equal to 0. At this time, the cluster network status corresponding to the first node is normal, that is, the first node has not been disconnected from the cluster and can synchronize data with other nodes in the cluster. If the first statistical data is less than the first preset threshold and / or the second statistical data is greater than the second preset threshold, it can be understood that when the statistical result of failed synchronization is greater than 0, the cluster network status corresponding to the first node is faulty, that is, the first node has been disconnected from the cluster and cannot synchronize data with other nodes in the cluster.

[0069] In an optional embodiment of the present invention, determining the target node corresponding to the first node among the at least one second node based on the second data includes:

[0070] Obtain the first address data from the second data;

[0071] The target node corresponding to the first node is determined in the at least one second node using the first address data.

[0072] It should be noted that the first address data can be any address data related to the target node, and there are no restrictions here. For example, the first address data can be the Internet Protocol address (IP) of the target node.

[0073] In practical applications, obtaining the first address data in the second data can be understood as obtaining the IP address of the target node in the connection replay data. Determining the target node corresponding to the first node in at least one second node through the first address data can be understood as determining the second node in the cluster that has received the synchronization data from the first node through the IP address, and taking these nodes as the target nodes.

[0074] In an optional embodiment of the present invention, the method further includes:

[0075] Obtain the third data from the first node;

[0076] The third data is used to determine the target data corresponding to the connection status of the device associated with the first node;

[0077] The second data is updated based on the target data.

[0078] It should be noted that the third data can be any data related to the first node, without any limitations. For example, the third data can be the accumulated topic data. Specifically, the accumulated topic data can include the names of the topics subscribed to in the first node and the number of subscribers.

[0079] In practical applications, using third data to determine the target data corresponding to the connection status of devices related to the first node can be understood as using topic accumulation data to determine the topics and subscription quantity of devices that have newly established or disconnected from the first node; updating the second data based on the target data can be understood as updating the connection replay data according to the target data.

[0080] In an optional embodiment of the present invention, the connection state includes a first state and a second state; the first state represents the state corresponding to the establishment of a connection between the first node and the device; the second state represents the state corresponding to the disconnection between the first node and the device; the step of obtaining the second data corresponding to the first node includes:

[0081] Obtain first target data corresponding to the state where the first node establishes a connection with the device, and second target data corresponding to the state where the first node disconnects from the device;

[0082] The first target data and the second target data are used as the second data.

[0083] It should be noted that the first state, representing the state corresponding to the establishment of a connection between the first node and the device, can be understood as the first state recording the newly established connection between the device and the first node; the second state, representing the state corresponding to the disconnection between the first node and the device, can be understood as the second state recording the disconnection between the device and the first node; the first target data can be data from any device that has established a connection with the first node, without limitation here. As an example, the first target data can be the topics and subscription quantity subscribed by the device that has newly established a connection with the first node; the second target data can be data from any device that has disconnected from the first node, without limitation here. As an example, the first target data can be the topics and subscription quantity subscribed by the device that has disconnected from the first node.

[0084] In practical applications, using the first target data and the second target data as the second data can be understood as recording the topics and subscriptions of devices that have newly connected to the first node and the topics and subscriptions of devices that have disconnected into the connection playback data.

[0085] In an optional embodiment of the present invention, after sending the target data to the target node, the method further includes:

[0086] If, after a preset time interval, the re-determined network state indicates that the network corresponding to the first node remains normal, the target data that has been sent to the target node in the second data is deleted.

[0087] It should be noted that the preset time can be any periodic setting, and there are no restrictions here. For example, the preset time can be a setting with a period of 5 milliseconds.

[0088] In practical applications, the preset interval can be understood as each interval of one cycle; the redefined network state characterization indicates that the network corresponding to the first node has always maintained normal operation, which can be understood as the first node maintaining a connection with the corresponding cluster. In this state, the first node can synchronize data with the target node; deleting the target data that has been sent to the target node in the second data can be understood as deleting the topics and subscriptions of newly created or disconnected devices that have been successfully synchronized to the target node in the connection replay data of the first node.

[0089] In an optional embodiment of the present invention, the method further includes:

[0090] If the target node successfully receives the target data, the third data corresponding to the target node is updated using the target data.

[0091] In practical applications, the successful reception of target data by the target node can be understood as the target node receiving the topics and subscription numbers of newly established or disconnected devices sent by the first node; updating the third data corresponding to the target node using the target data can be understood as the target node updating its own topic accumulation data with the received target data, that is, updating the subscribed topic names and the number of subscribers in its own topic accumulation data.

[0092] To understand the embodiments of the present invention, a fault handling method for cluster nodes is exemplified here, the specific process of which is as follows:

[0093] Each node in the cluster maintains a simple in-memory database and some processing processes. The in-memory database is used to store data information related to that node; the processing processes are used to handle faults; the in-memory database can contain multiple in-memory data tables to store different types of data records.

[0094] 1. Topic accumulator.

[0095] The in-memory database contains a topic accumulator table named topic_acc, which records the statistical results of subscribers to all topics within this node. The topic accumulator table for topic_acc can be found in Table 1, which is a schematic diagram of the topic accumulator table.

[0096] topic String Subscribed topic name count numerical values Number of subscribers

[0097] Table 1

[0098] Figure 2 This is a schematic diagram illustrating the processing flow of a topic accumulator provided in an embodiment of the present invention. Whenever a new connection attempts to subscribe to a topic, if the topic_acc table does not contain a record for that topic, a new record is added to the table, and the count value is set to 1; if the topic exists, the count value of the corresponding record for that topic is incremented by 1. Similarly, when a connection corresponding to a topic is closed, if the topic_acc table contains a record for that topic and its count value is greater than 1, its count value is decremented by 1; otherwise, the record corresponding to that topic is deleted.

[0099] 2. Routing synchronization process.

[0100] Each node has an internal health check probe process that sets the cluster state to UP or DOWN based on routing synchronization statistics. Each node also has an internal routing synchronization process that, when the cluster state is UP, broadcasts data to other nodes in the cluster and records statistics on successful or failed data synchronization. This statistical data is stored in the node's routing synchronization data table, named sync_acc. Table 2 shows a schematic diagram of the routing synchronization data table.

[0101] sync_success numerical values Statistical results of successful synchronization sync_fail numerical values Statistical results of synchronization failure

[0102] Table 2

[0103] When node data synchronization is successful, the value of the corresponding field of sync_success in the sync_acc table is incremented by 1; when node data synchronization fails, the value of the corresponding field of sync_fail in the sync_acc table is incremented by 1, and the routing information corresponding to the new connection and disconnection of the node is recorded in the node's connection replay table, named conn_replay. A schematic diagram of the three-dimensional connection replay table is shown below.

[0104] clientId String Device Client ID nodeIP String Current node IP address topics String list The list of topics subscribed to by this device destIPs IP address list List of all target node IPs that failed to synchronize type numerical values 1 indicates a new connection; 2 indicates a disconnection.

[0105] Table 3

[0106] When the cluster status is DOWN, route synchronization is not performed; only the route information corresponding to the connection is recorded in the connection replay table. Figure 3 This is a schematic diagram illustrating the workflow of a routing synchronization process provided in an embodiment of the present invention.

[0107] 3. Generate final routing data.

[0108] When a connection type is "new" and the count value of the record corresponding to that topic in the `topic_acc` table is greater than 1, it is assumed that the current node already has other connections subscribed to that topic, and no action is taken. Otherwise, a new route record is added to the `conn_replay` table, and `type=1` is set in the `conn_replay` table to indicate a new connection. When there is a disconnection request or a node detects that a device connection has been disconnected, if there is no record corresponding to that topic in the `topic_acc` table, it means that no other connections besides that connection are subscribed to that topic. At this time, it is checked whether the route record for the corresponding topic exists in the `conn_replay` table. If it exists, its type is set to 2; otherwise, a new route record is added with a type value of 2. On the other hand, if there is a route record corresponding to that topic in the `topic_acc` table, it means that other connections besides that connection are subscribed to that topic. In this case, there is no need to update the record in `conn_replay`, indicating that no other node needs to delete this route information. Figure 4 This is a schematic diagram of a process for connecting a playback table, provided as an embodiment of the present invention.

[0109] 4. Health check probe.

[0110] The health check probe process scans the current values ​​in the `sync_acc` table every `periodMS` milliseconds to determine the cluster network status, and then sets the cluster status to UP or DOWN based on the network connectivity between nodes. After each cycle is completed, the values ​​of the `sync_fail` and `sync_success` fields are reset to 0 so that the routing synchronization results for the next scan cycle can be recalculated.

[0111] If the health check probe process detects a value greater than 0 for sync_fail, it considers a network failure to have occurred within the cluster. In this case, the cluster state is set to DOWN, and the connection replay process is notified of the current cluster state. Conversely, if, within a scan cycle, the health check probe process detects a value of 0 for sync_fail in the sync_acc table and a value greater than 0 for sync_succ, it considers the cluster network to have recovered. In this case, the cluster state is set to UP, and the connection replay process is notified of the current cluster state. Figure 5 This is a schematic diagram illustrating the workflow of a health check probe process provided in an embodiment of the present invention.

[0112] 5. Connection playback mechanism.

[0113] Figure 6This is a schematic diagram illustrating the workflow of a connection replay process according to an embodiment of the present invention. The connection replay process is a standalone process that executes only when it receives a message from the health check probe process; otherwise, it blocks until a new message arrives. If the received message indicates that the cluster health status is UP, the connection replay operation begins; otherwise, the message is discarded and the process continues to block. Its execution logic is as follows: it scans each record in the `conn_replay` table, obtaining the list of failed synchronization destination node IP addresses corresponding to the `destIPs` field and the type (addition or deletion) corresponding to the `type` field. It then broadcasts routing data `nodeIP`, `topics`, and `type` to all nodes in the address list. Upon successful execution, the data record is deleted until there is no more data in the `conn_replay` table. Nodes receiving this broadcast message perform the operation of adding or deleting routing table records locally based on the `type` field, maintaining the new mapping between `nodeIP` and `topics`.

[0114] Based on the same inventive concept as described above, Figure 7 This is a schematic diagram of a fault handling device for a cluster node provided in an embodiment of the present invention. The device 700 includes:

[0115] The determining unit 701 is used to acquire first data corresponding to the first node and determine the network status corresponding to the first node based on the first data; the first data indicates whether the data corresponding to the first node has been successfully synchronized to at least one second node; the first node is any node in the cluster nodes; the second node is any node in the cluster nodes other than the first node;

[0116] The acquisition unit 702 is used to acquire second data corresponding to the first node when the network status characterizes that the network corresponding to the first node has failed; the second data characterizes target data corresponding to the connection status of the device related to the first node.

[0117] The determining unit 701 is further configured to reacquire the first data corresponding to the first node, and re-determine the network state corresponding to the first node based on the reacquired first data.

[0118] The determining unit 701 is further configured to, when the redefined network state indicates that the network corresponding to the first node has returned to normal, determine a target node corresponding to the first node among the at least one second node based on the second data; and send the target data to the target node.

[0119] In some embodiments, the first data includes a first statistical data point showing that the data corresponding to the first node was successfully synchronized to each of the second nodes and a second statistical data point showing that the data corresponding to the first node was not successfully synchronized to each of the second nodes.

[0120] The determining unit 701 is further configured to determine that the network state corresponding to the first node is a normal state when the first statistical data is greater than or equal to the first preset threshold and the second statistical data is less than or equal to the second preset threshold.

[0121] The determining unit 701 is further configured to determine the network state corresponding to the first node as a fault state when the first statistical data is less than the first preset threshold and / or the second statistical data is greater than the second preset threshold.

[0122] In some embodiments, the acquisition unit 702 is further configured to acquire the first address data in the second data;

[0123] The determining unit 701 is further configured to determine, through the first address data, the target node corresponding to the first node in the at least one second node.

[0124] In some embodiments, the acquisition unit 702 is further configured to acquire third data of the first node;

[0125] The determining unit 701 uses the third data to determine the target data corresponding to the connection status of the device related to the first node;

[0126] The device 700 further includes an update unit for updating the second data based on the target data.

[0127] In some embodiments, the connection state includes a first state and a second state; the first state represents the state corresponding to the establishment of a connection between the first node and the device; the second state represents the state corresponding to the disconnection between the first node and the device.

[0128] The acquisition unit 702 is further configured to acquire first target data corresponding to the state in which the first node establishes a connection with the device and second target data corresponding to the state in which the first node disconnects from the device; and use the first target data and the second target data as the second data.

[0129] In some embodiments, the apparatus 700 further includes: a deletion unit, configured to, after sending the target data to the target node, delete the target data already sent to the target node in the second data after a preset time interval, if the re-determined network state indicates that the network corresponding to the first node has remained normal.

[0130] In some embodiments, the updating unit is further configured to update the third data corresponding to the target node using the target data if the target node successfully receives the target data.

[0131] Figure 8 This is a schematic diagram of a hardware structure for a fault handling device for a cluster node according to an embodiment of the present invention. The device 80 includes at least one processor 801 and a memory 802. Optionally, the device 80 may further include at least one communication interface 803. The various components in the device 80 are coupled together through a bus system 804. It is understood that the bus system 804 is used to realize communication between these components. In addition to a data bus, the bus system 804 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 8 The general labeled all buses as Bus System 804.

[0132] It is understood that memory 802 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), Sync Link Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memory 802 described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.

[0133] The memory 802 in this embodiment of the invention is used to store various types of data to support the operation of the fault handling device 80 of the cluster nodes. Examples of such data include any computer program for operation on the device 80, and programs implementing the methods of this embodiment of the invention may be included in the memory 802.

[0134] The methods disclosed in the above embodiments of the present invention can be applied to or implemented by processor 801. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of the present invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in memory. The processor reads information from the memory and, in conjunction with its hardware, completes the steps of the aforementioned method.

[0135] In an exemplary embodiment, the fault handling 80 of the cluster node can be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to perform the above-described method.

[0136] The present invention also provides a computer-readable storage medium storing a computer program thereon; when executed by a processor, the computer program implements the steps of any of the methods described above. The computer-readable storage medium may be a magnetic random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM), etc.; or it may be a device comprising one or any combination of the above-mentioned memories.

[0137] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0138] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0139] In addition, in the various embodiments of the present invention, each functional unit can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0140] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0141] Alternatively, if the integrated units of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.

[0142] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for handling faults in cluster nodes, characterized in that, The method includes: Obtain the first data corresponding to the first node, and determine the network status corresponding to the first node based on the first data; the first data indicates whether the data corresponding to the first node has been successfully synchronized to at least one second node; the first node is any node in the cluster nodes; the second node is any node in the cluster nodes other than the first node; When the network status indicates that the network corresponding to the first node has failed, second data corresponding to the first node is obtained; the second data indicates target data corresponding to the connection status of the device related to the first node. Reacquire the first data corresponding to the first node, and redetermine the network state corresponding to the first node based on the reacquired first data; If the redefined network state indicates that the network corresponding to the first node has returned to normal, a target node corresponding to the first node is determined from the at least one second node based on the second data; and the target data is sent to the target node. The connection state includes a first state and a second state; the first state represents the state corresponding to the establishment of a connection between the first node and the device; the second state represents the state corresponding to the disconnection between the first node and the device; obtaining the second data corresponding to the first node includes: Obtain first target data corresponding to the state where the first node establishes a connection with the device, and second target data corresponding to the state where the first node disconnects from the device; The first target data and the second target data are used as the second data.

2. The method according to claim 1, characterized in that, The first data includes a first statistical data point showing that the data corresponding to the first node was successfully synchronized to each of the second nodes, and a second statistical data point showing that the data corresponding to the first node was not successfully synchronized to each of the second nodes. Determining the network status corresponding to the first node based on the first data includes: If the first statistical data is greater than or equal to the first preset threshold and the second statistical data is less than or equal to the second preset threshold, the network status corresponding to the first node is determined to be normal. If the first statistical data is less than the first preset threshold, and / or the second statistical data is greater than the second preset threshold, the network state corresponding to the first node is determined to be a fault state.

3. The method according to claim 1, characterized in that, The step of determining the target node corresponding to the first node among the at least one second node based on the second data includes: Obtain the first address data from the second data; The target node corresponding to the first node is determined in the at least one second node using the first address data.

4. The method according to claim 1, characterized in that, The method further includes: Obtain the third data from the first node; The third data is used to determine the target data corresponding to the connection status of the device associated with the first node; The second data is updated based on the target data.

5. The method according to claim 1, characterized in that, After sending the target data to the target node, the method further includes: If, after a preset time interval, the re-determined network state indicates that the network corresponding to the first node remains normal, the target data that has been sent to the target node in the second data is deleted.

6. The method according to claim 1, characterized in that, The method further includes: If the target node successfully receives the target data, the third data corresponding to the target node is updated using the target data.

7. A fault handling device for cluster nodes, characterized in that, The device includes: A determining unit is configured to acquire first data corresponding to a first node and determine the network status corresponding to the first node based on the first data; the first data indicates whether the data corresponding to the first node has been successfully synchronized to at least one second node; the first node is any node in the cluster nodes; the second node is any node in the cluster nodes other than the first node; The acquisition unit is configured to acquire second data corresponding to the first node when the network status characterizes a network failure corresponding to the first node; the second data characterizes target data corresponding to the connection status of devices associated with the first node. The determining unit is further configured to reacquire the first data corresponding to the first node, and re-determine the network state corresponding to the first node based on the reacquired first data. The determining unit is further configured to, when the redefined network state indicates that the network corresponding to the first node has returned to normal, determine a target node corresponding to the first node among the at least one second node based on the second data; and send the target data to the target node; The connection state includes a first state and a second state; the first state represents the state corresponding to the establishment of a connection between the first node and the device; the second state represents the state corresponding to the disconnection between the first node and the device; obtaining the second data corresponding to the first node includes: Obtain first target data corresponding to the state where the first node establishes a connection with the device, and second target data corresponding to the state where the first node disconnects from the device; The first target data and the second target data are used as the second data.

8. A fault handling device for cluster nodes, characterized in that, The processing device includes a processor and a memory for storing a computer program capable of running on the processor, wherein the processor, when running the computer program, performs the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Fault processing method and device of database cluster, and terminal

    CN108599996A

  • Cluster node fault recovery method and device, electronic equipment and storage medium

    CN111124755A