Automatic fault transfer method for time sequence database of park integrated energy monitoring system
By using the arbitration and voting mechanism to achieve automatic election and rapid failover of the master node in the time series database cluster of the comprehensive energy monitoring system at the campus level, the problems of complex automatic switching and difficult data consistency in the existing technology are solved, and the high availability and fault tolerance of the system are improved.
Patent Information
- Application Number
- CN202411767943.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-04
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing campus-level comprehensive energy monitoring system timing database clusters have complex automatic switching when facing failure scenarios such as main node failure and network partitioning, and the failure recovery speed is slow, and data consistency is difficult to ensure.
Automatic election of the master node is achieved through the arbitration voting mechanism, ensuring that the system can quickly switch master-slave roles when a failure occurs, and ensuring data consistency. The specific implementation includes: when the master node fails, all slave nodes participate in arbitration voting as candidate nodes, quickly determine the new master node through a decentralized arbitration mechanism, and data synchronization is performed through a service bus mechanism based on message queue.
It realizes rapid and automatic failover when the master node fails, reduces service interruption time, improves the high availability and failure tolerance of the timing database cluster, and ensures data consistency.
Smart Images

Figure CN119938638A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of distributed database and integrated energy monitoring, and in particular relates to a method for automatic transfer of time series database failures in a park integrated energy monitoring system. Background Art
[0002] In the park-level integrated energy monitoring system, the time series database plays a vital role. It is used to store and manage a large amount of time-varying data. These data are of great significance for the system's operation monitoring, fault diagnosis, and decision analysis, requiring the system to maintain high reliability and real-time performance in various fault scenarios. Due to the complexity and high reliability requirements of the power system, this can be achieved by building a time series database cluster, that is, by working together through multiple database nodes to improve the availability and performance of the system. In the time series database cluster, the master node may fail during the operation of the database cluster. Common host failure scenarios include: power outage, unexpected shutdown, unexpected system restart, system crash, network interruption, and service process failure. After the master node fails, it is crucial to realize the automatic transfer of database services from the master node to other nodes, that is, to realize automatic fault transfer. This is a key technology to achieve high availability.
[0003] The accuracy and continuity of data directly affect the decision-making and responsiveness of the system. Although the existing campus-level integrated energy monitoring system time series database cluster achieves high data availability through a master-slave architecture, when faced with failure scenarios such as master node failure and network partitioning, the traditional master-slave architecture often requires manual promotion of a slave server to a new master server. This involves complex operations such as configuration changes and data synchronization, which are prone to human errors and long downtime. Automatic switching is complex and fault recovery is slow. Since master-slave replication synchronizes data asynchronously, there are factors such as network delays and replication delays. There may be certain differences between the data on the slave server and the master server. During fault recovery or switching, data consistency is difficult to guarantee. That is, the existing campus-level integrated energy monitoring system time series database cluster method is often not flexible and efficient when dealing with failover. There are problems such as complex automatic switching strategies, difficulty in ensuring data consistency, data loss, and delays or long interruptions in data services.
[0004] Therefore, there is an urgent need for an automatic failover solution that can automatically switch between master and slave nodes and quickly restore data consistency when the master node fails or the network environment fluctuates. Summary of the invention
[0005] The purpose of the present invention is to provide a method for automatic transfer of time series database failures in a park integrated energy monitoring system. The core idea is to realize automatic election of a master node through an arbitration voting mechanism to ensure that when a failure occurs, the system can quickly switch between master and slave roles and ensure data consistency. It solves the problem of being able to quickly and automatically switch when a master node fails, and making a slave node become a new master node, thereby realizing automatic transfer of failures, reducing the service interruption time caused by master node failures, and improving the high availability and fault tolerance of the time series database cluster.
[0006] In order to achieve the above object, the solution of the present invention is:
[0007] A method for automatic transfer of time series database failures in a park-level integrated energy monitoring system, wherein the time series database cluster nodes of the park-level integrated energy monitoring system form a distributed network topology; the master node is responsible for processing all write requests and synchronizing time series data to other replica nodes, the slave node is responsible for receiving synchronization data from the master node and performing persistent storage, and the voting process adopts a decentralized arbitration mechanism. When a master node failure is detected, any slave node / arbitration node can initiate an arbitration voting mechanism, and the slave node immediately becomes a candidate node; the method comprises:
[0008] When the master node fails, all slave nodes serve as candidate nodes and participate in the arbitration vote of the master node;
[0009] Each slave node has one vote to select a candidate node to vote for or against failover;
[0010] The votes of each candidate node are counted. If a candidate node obtains more than half of the total votes in favor of failover before the election times out, the candidate node will serve as the new master node.
[0011] The number of nodes in the time series database cluster is counted. If the number of nodes is an even number, an arbitration node is added to the time series database cluster. The arbitration node does not contain a data set, cannot be used as a candidate node, and cannot be promoted to a master node. It has only one voting right.
[0012] Among them, the method to determine whether the master node fails is:
[0013] A heartbeat connection is established between each node in the time series database cluster and other nodes, and the source node periodically sends heartbeat information to the target node at a time interval T1; wherein the target node is the master node;
[0014] The target node processes the heartbeat information and sends a response to the source node;
[0015] The source node receives the heartbeat response and updates the target node status;
[0016] A heartbeat detection method based on a sliding time window is used. If the source node fails to receive a heartbeat response from the target node for n2 consecutive times within the set time interval T2, the target node is considered to be faulty, where T2 is the length of the sliding time window, and T2≥n1T1 is satisfied.
[0017] The heartbeat information must include the data, timestamp, database node identifier and status code, and optionally, health parameters;
[0018] The timestamp marks the message sending time and is used to detect network delays or whether the node has timed out. The database node identifier is a unique ID or name that identifies the sender and receiver. The status code is "normal", "overloaded", or "unreachable", which is used to reflect the operating status of the current node. The health parameters include CPU, memory usage, disk usage, etc.
[0019] Among them, the arbitration vote of the master node includes the following specific contents:
[0020] The master node initializes the current term;
[0021] When the master node fails, all slave nodes act as candidate nodes, increase their current terms by one, and send voting request messages to other nodes;
[0022] When each of the other nodes receives the voting request, if the term in the voting request is less than the term of the current node, it votes against it; otherwise, it updates the term of the current node to the term in the voting request, and then compares the data update date; if the data update date of the candidate node is later than that of the current node, it votes in favor; otherwise, it votes against it.
[0023] Among them, when the master node fails, any slave node can initiate an arbitration vote.
[0024] The voting status of each candidate node is counted. If no candidate node obtains more than half of the total votes in favor of failover before the election times out, each candidate node returns to the slave node state and waits for the next election.
[0025] This also includes, when it is determined that the number of normally surviving nodes in the time series database cluster of the park-level comprehensive energy monitoring system is less than 1 / 2 of the total number of clusters, the cluster is unavailable, that is, it only supports read operations but not write operations.
[0026] This also includes repairing and restoring the failed master node, and automatically rejoining the cluster as a slave node after recovery to accept data replication and synchronization from the master node.
[0027] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor; when the processor executes the computer program, the steps of the method for automatic transfer of a time series database failure in a park integrated energy monitoring system as described above are implemented.
[0028] A computer-readable storage medium stores a computer program; when the computer program is executed by a processor, the steps of the method for automatic transfer of a time series database failure in a park integrated energy monitoring system as described above are implemented.
[0029] After adopting the above scheme, the present invention has the following advantages:
[0030] (1) The present invention adopts an arbitration voting mechanism, which can realize fast and efficient automatic fault transfer in a time series database cluster. By optimizing the voting mechanism and arbitration mechanism in the election process, the master node can be quickly determined, the election time can be reduced, the response speed of the system can be improved, and the availability of the time series database set can be improved. Voting decisions are made based on multiple factors to ensure the accuracy and reliability of fault transfer and reduce system instability caused by misjudgment or wrong decisions. In the event of multiple candidates competing, the arbitration mechanism can ensure that the most suitable master node is selected, avoid election deadlocks, and improve system stability. Data synchronization is performed through a service bus mechanism based on a message queue, which can ensure data consistency of all nodes in the cluster and ensure the accuracy and reliability of power monitoring data.
[0031] (2) The arbitration mechanism of the present invention takes into account the node status and network conditions, and can maintain high availability of the system in a complex network environment with strong adaptability. Through efficient voting and data synchronization mechanisms, it ensures that the system can quickly resume normal operation after the failure of the master node, reducing the impact of the failure on the system. The entire fault detection, election and data synchronization process are automatically completed by the system, reducing the dependence on manual operation and maintenance intervention and the risk of misoperation, and improving the stability and operation and maintenance efficiency of the system. In addition, it is easy to expand to a large-scale time series database cluster, and by increasing the number of node members, the system's processing power and fault tolerance can be improved.
[0032] (3) The arbitration voting mechanism designed in the present invention is simple and reliable, easy to understand and implement, and convenient to maintain and expand. Ultimately, it can effectively improve the stability and fault tolerance of the time series database cluster and reduce the development and maintenance costs of the park-level integrated energy monitoring system. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a flowchart of the voting arbitration mechanism in the present invention;
[0034] Figure 2is a schematic diagram of a multi-node arrangement structure in an embodiment of the present invention;
[0035] Figure 3 It is a schematic diagram of a multi-node arrangement structure when an arbitration node is used in an embodiment of the present invention. DETAILED DESCRIPTION
[0036] The technical solutions and beneficial effects of the present invention will be described in detail below with reference to the accompanying drawings.
[0037] The present invention provides a method for automatically transferring a time series database failure of a park integrated energy monitoring system, comprising:
[0038] To ensure the high availability of the time series database of the campus-level monitoring system, the time series database cluster node design adopts a multi-data copy redundant design mechanism. The nodes in the database cluster are connected through a high-speed network to form a distributed network topology. Data synchronization, communication, and fault detection can be performed between nodes through the network. In order to improve the reliability of the system, redundant network connections can be used to ensure that nodes can still communicate when the network fails.
[0039] The data in the time series database is stored in time series, and each data point contains a timestamp, data value, and other related information. Data can be stored on the local disk of the node or in a distributed file system to improve data reliability and availability.
[0040] The nodes in the cluster are divided into master nodes, slave nodes, candidate nodes and arbitration nodes.
[0041] The master node is responsible for processing all write requests and synchronizing time series data to other replica nodes. The slave node is responsible for receiving synchronization data from the master node and storing it persistently. When the master node fails, the slave node becomes the new master node through election. When the master node fails, the slave node can initiate an election, become a candidate node and participate in the arbitration vote of the master node. The arbitration node cannot be used alone, it does not contain a data set, and cannot be promoted to a master node. It is only used to vote for the master node. The arbitration node uses minimal resources and does not require hardware devices.
[0042] A cluster consisting of a master node and two slave nodes is as follows Figure 2 As shown in Figure 1, a cluster consisting of a master node, a slave node, and an arbitration node is as follows: Figure 3 As shown in the figure. The existence of arbitration nodes allows the total number of nodes in the cluster to exist as an even number of nodes without adding new nodes to the data set. The only communication between arbitration nodes and other nodes is: voting during the election process, heartbeat detection, and configuration data. There is no data synchronization process between arbitration nodes and master nodes.
[0043] The number of nodes in the time series data cluster is an odd number, and at least three replica nodes are included. When the number of nodes in a time series data cluster is an even number, an arbitration node is added to form an odd number of nodes. There is at most one arbitration node in the time series data cluster. When the number of normally surviving nodes in the cluster is less than 1 / 2 of the total number of clusters, the cluster is unavailable, that is, it only supports read operations and does not support write operations, to ensure that the majority of nodes in the cluster can always maintain consistency, and the system can still operate normally even if some nodes fail.
[0044] Multiple databases in a time series data cluster need to maintain a connection state. A heartbeat mechanism is used to synchronize member status information, detect the health of the database and the validity of the connection, and ensure the continuity of data synchronization or other operations. Each node in the cluster periodically sends heartbeat information to other nodes in the cluster to obtain status. The node that actively initiates a heartbeat request is called the source node, and the node that receives the heartbeat request becomes the target node. The complete steps of a heartbeat are as follows:
[0045] S201: The source node periodically sends heartbeat information to the target node according to the set time interval T1;
[0046] S202: The target node processes the heartbeat information and sends a response to the source node;
[0047] S203: The source node receives the heartbeat response and updates the target node status.
[0048] The data that the heartbeat information must include include timestamp, database node identifier and status code, and optionally, health parameters.
[0049] The timestamp marks the message sending time and is used to detect network delays or whether the node has timed out. The database node identifier is a unique ID or name that identifies the sender and receiver. The status code is "normal", "overloaded", or "unreachable", which is used to reflect the operating status of the current node. The health parameters include CPU, memory usage, disk usage, etc.
[0050] Network congestion may cause heartbeat packet delay or loss, which may cause normal nodes to be misjudged as faulty. To reduce this situation, you can set an appropriate heartbeat interval and timeout, and use the sliding time window method to improve accuracy.
[0051] The slave node sets the heartbeat timeout and uses the sliding time window method for timeout monitoring. The length of the sliding time window is T2, T2 ≥ n1T1. If the source node does not receive the heartbeat response from the target node for n2 consecutive times within the set time interval T2, the target node is considered to be faulty. Optionally, n1 = 5, n2 = 3.
[0052] To ensure the reliability of the message, you can also use the TCP protocol to ensure that the message can arrive successfully. The sending, receiving, and status changes of all heartbeat messages should be recorded in the log to facilitate subsequent system health diagnosis and debugging. If the heartbeat message delay or loss rate is too high, the log can help analyze whether it is a network problem or a system problem.
[0053] When the system is initialized, the voting priority number of each node is configured. The number cannot be repeated. The smaller the number, the greater the priority. The master node initializes the current term, and the initial term is a set positive integer.
[0054] When the master node fails, the slave nodes and arbitration nodes will start the arbitration voting mechanism. The voting rules are as follows:
[0055] 1) Slave nodes and arbitration nodes will vote based on factors such as the failure type, failure time, and data synchronization status of the master node.
[0056] 2) Each slave node and arbitration node has one vote, and the voting results are divided into two types: in favor of failover and against failover.
[0057] 3) If the number of votes in favor of failover exceeds half of the total votes, failover is performed; otherwise, failover is not performed.
[0058] The specific voting process is as follows: Figure 1 As shown, the arbitration voting steps are as follows:
[0059] S201: When the system is initialized, the voting priority number of each node is configured. The number cannot be repeated. The smaller the number, the greater the priority. The master node initializes the current term, and the initial term is a set positive integer.
[0060] S202: The nodes in the time series database cluster are connected through a high-speed network to form a distributed network topology. Each node establishes a heartbeat connection with other nodes in the cluster. The nodes can synchronize data, communicate, and detect faults through the network. Multiple databases in the time series data cluster need to maintain a connection.
[0061] S203: adopting appropriate heartbeat interval and timeout time, and designing a heartbeat mechanism in combination with a sliding time window detection method;
[0062] S204: If the master node replies that the heartbeat timeout has occurred, proceed to step S205; otherwise, proceed to step S203 to continue the heartbeat detection;
[0063] S205: The voting process adopts a decentralized arbitration mechanism. Any slave node / arbitration node can initiate voting. When a slave node / arbitration node detects that the master node fails, the slave node / arbitration node will immediately initiate an arbitration voting mechanism. The slave node will immediately transform itself into a candidate node. The candidate will increase the current term by one to mark the round of this election.
[0064] S206: The candidate node sends a voting request message to other slave nodes / arbitration nodes to ask whether they support it to become the new master node;
[0065] S207: When other nodes receive the voting request, if the data synchronization of the slave node is not completed, they proceed to step S208, otherwise they proceed to step S209;
[0066] S208: The slave node continues to synchronize data, and then votes after completion;
[0067] S209: If the term in the request is less than the term of the current node, proceed to step S211 to vote against; if the term in the request is greater than the term of the current node, update the term of the current node to the term in the request, and then compare the data update date. If the data update date of the candidate node is newer than that of the current node, proceed to step S210 to vote in favor, otherwise proceed to step S211 to vote against; each node is only allowed to vote once in one election cycle;
[0068] S210: Other nodes vote in favor;
[0069] S211: Other nodes vote against it;
[0070] S212: Perform vote aggregation. If the candidate node obtains votes from more than half of the nodes before the election times out, it proceeds to step S213 and becomes the new master node; otherwise, it proceeds to step S214;
[0071] S213: Becomes the new master node. The new master node will take over the work of the original master node and notify other nodes of its status;
[0072] S214: Reset the election timeout timer, return to the slave node state, and wait for the next election opportunity;
[0073] S215: The new master node copies the data to other nodes through the data synchronization mechanism;
[0074] S216: The new master node will be responsible for processing client requests again;
[0075] S217: This round of elections ends.
[0076] The voting process adopts a decentralized arbitration mechanism, and any slave node / arbitration node can initiate a vote.
[0077] When a slave node / arbitration node detects a master node failure, it will immediately initiate an arbitration voting mechanism, and the slave node will immediately transform itself into a candidate node. The candidate will increase the current term by one to identify this election round, and then initiate an election vote and send a voting request message to other slave nodes / arbitration nodes to ask whether they support it to become the new master node.
[0078] When other nodes receive a voting request, if the data synchronization from the node is not completed, they need to wait until the data synchronization is completed before deciding whether to vote. Whether to vote is based on whether the data update date of the candidate node is the latest and whether the term is the largest. If the term in the request is less than the term of the current node, then vote against it; if the term in the request is greater than the term of the current node, then update the term of the current node to the term in the request, and then compare the data update date. If the data update date of the candidate node is newer than the date of the current node, then vote in favor, otherwise vote against it. Each node is only allowed to vote once in an election cycle.
[0079] Each node has an election timeout timer. After the candidate node sends a voting request, it will wait for responses from other nodes to aggregate votes. If the candidate node obtains more than half of the votes from the nodes before the election timeout, it will become the new master node. Otherwise, it will reset the election timeout timer, return to the slave node state, and wait for the next election opportunity.
[0080] Through the majority voting principle, the candidate node that obtains more than half of the votes will become the new master node. This rule ensures the fairness of the election and avoids the situation where multiple candidates receive votes at the same time.
[0081] After the new master node is successfully elected, it will take over the work of the original master node and notify other nodes of its status. Other nodes will use the new master node as their own master node and synchronize data. The new master node will be responsible for processing client requests again and copy data to other nodes through the data synchronization mechanism.
[0082] The data synchronization process adopts a service bus mechanism based on a message queue. The message queues used by the service bus include ActiveMQ, RabbitMQ, ZeroMQ, Kafka, and RocketMQ. Preferably, ZeroMQ is used.
[0083] If an abnormality occurs in the new master node or the transfer fails during the failover process, the system will automatically re-trigger the voting mechanism according to the above method and re-elect a new master node until the fault is completely eliminated to ensure the continuous and stable operation of the monitoring system.
[0084] The system can also repair and recover the failed node so that it can rejoin the database cluster at the appropriate time. The repair of the failed node can include hardware repair, software upgrade, data recovery and other steps. After the repair is completed, the failed node can be rejoined to the database cluster, and data synchronization and status check can be performed to ensure that it can operate normally.
[0085] If the failed node recovers, it will automatically rejoin the cluster as a slave node and accept data replication and synchronization from the master node.
[0086] In a distributed system, network partition may cause a cluster to have multiple master nodes. To avoid this, the present invention uses a majority principle based on an arbitration voting mechanism to ensure that only one node can be elected as the master node at any time.
[0087] Assume that a network partition causes the cluster to be split into two parts, with nodes A, B, and C in one part of the network and D and E in the other part. When the failure occurs, node A cannot communicate with nodes D and E due to the network partition. Node A initiates an election, and nodes B and C vote for A in the same network partition. Nodes D and E cannot elect a new master node because they do not receive enough votes. The system automatically prohibits nodes D and E from initiating a new master node election to avoid multiple master nodes. When the network is restored, node A synchronizes the latest data to nodes D and E to ensure data consistency across the cluster.
[0088] An embodiment of the present invention also provides another computer device, including one or more processors and a memory for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors execute the aforementioned method of fault tolerance and automatic fault transfer of a campus-level integrated energy monitoring system timing database based on a voting arbitration mechanism.
[0089] In practical applications, the processor may be a server chip, a desktop chip, a mobile chip, an embedded microprocessor, or a DSP chip. It is understandable that for different devices, the electronic device used to implement the processor function may be other, which is not specifically limited in the embodiments of the present invention.
[0090] The above-mentioned memory can be a server, a workstation, a minicomputer, a laptop, an industrial computer, an embedded development board; or a combination of the above-mentioned types of memory, and provides instructions and data to the processor.
[0091] In an exemplary embodiment, an embodiment of the present invention further provides a computer-readable storage medium for storing a computer program.
[0092] Optionally, the computer-readable storage medium can be applied to any one of the methods in the embodiments of the present invention, and the computer program enables the computer to execute the corresponding processes implemented by the processor in each method in the embodiments of the present invention. For the sake of brevity, they are not described here.
[0093] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0094] In summary, the present invention provides a method for fault tolerance and automatic fault transfer of a time series database of a campus-level integrated energy monitoring system based on a voting arbitration mechanism. The nodes in the cluster are divided into master nodes, slave nodes, candidate nodes and arbitration nodes. A heartbeat detection mechanism based on a sliding window is used to detect the health of the database and the validity of the connection. When a failure of the master node is detected, the slave node and the arbitration node will start the arbitration voting mechanism. The voting process adopts a decentralized arbitration mechanism. Any slave node / arbitration node can initiate a vote. Other nodes vote based on the term of the candidate node and the data update date. The new node uses a service bus mechanism based on a message queue to synchronize data between multiple data copies in the cluster, which can ensure data consistency of all nodes in the cluster. Fast and efficient automatic fault transfer is achieved in the time series database cluster, reducing dependence on manual labor and improving the stability and fault tolerance of the system.
[0095] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes. The schemes in the embodiments of the present invention may be implemented in various computer languages, for example, object-oriented programming language Java and literal scripting language JavaScript, etc.
[0096] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0097] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0098] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0099] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0100] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A method for automatic transfer of a time series database failure in a park-level integrated energy monitoring system, wherein the time series database cluster nodes of the park-level integrated energy monitoring system form a distributed network topology; characterized in that: include, When the master node fails, all slave nodes serve as candidate nodes and participate in the arbitration vote of the master node; Each slave node has one vote to select a candidate node to vote for or against failover; The votes of each candidate node are counted. If a candidate node obtains more than half of the total votes in favor of failover before the election times out, the candidate node will serve as the new master node.
2. The method according to claim 1, characterized in that: The number of nodes in the time series database cluster is counted. If the number of nodes is an even number, an arbitration node is added to the time series database cluster. The arbitration node does not contain a data set, cannot be used as a candidate node, and has only one voting right.
3. The method according to claim 1, characterized in that: The method to determine whether the master node fails is: A heartbeat connection is established between each node in the time series database cluster and other nodes, and the source node periodically sends heartbeat information to the target node at a time interval T1; wherein the target node is the master node; The target node processes the heartbeat information and sends a response to the source node; If the source node fails to receive a heartbeat response from the target node for n2 consecutive times within the set time interval T2, the target node is considered to be faulty, where T2 is the time length of the sliding time window, satisfying T2≥n1T1.
4. The method according to claim 1, characterized in that: The arbitration vote of the master node includes the following specific contents: The master node initializes the current term; When the master node fails, all slave nodes act as candidate nodes, increase their current terms by one, and send voting request messages to other nodes; When each of the other nodes receives the voting request, if the term in the voting request is less than the term of the current node, it votes against it; otherwise, it updates the term of the current node to the term in the voting request, and then compares the data update date; if the data update date of the candidate node is later than that of the current node, it votes in favor; otherwise, it votes against it.
5. The method according to claim 1, characterized in that: When the master node fails, any slave node can initiate an arbitration vote.
6. The method according to claim 1, characterized in that: The votes of each candidate node are counted. If no candidate node obtains more than half of the total votes in favor of failover before the election times out, each candidate node returns to the slave node state and waits for the next election.
7. The method according to claim 1, characterized in that: It also includes that when it is determined that the number of normally surviving nodes in the time series database cluster of the park-level comprehensive energy monitoring system is less than 1 / 2 of the total number of clusters, the cluster is unavailable, that is, it only supports read operations but not write operations.
8. The method according to claim 1, characterized in that: It also includes repairing and restoring the failed master node, and automatically rejoining the cluster as a slave node after recovery to accept data replication and synchronization from the master node.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor; characterized in that: When the processor executes the computer program, the steps of the method for automatic transfer of time series database failure of a park integrated energy monitoring system as described in any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium storing a computer program; characterized in that: When the computer program is executed by the processor, the steps of the method for automatic transfer of time series database failure of a park integrated energy monitoring system as described in any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Cluster voting arbitration method and system
CN106789193A
Information processing method and device and electronic equipment
CN113886129A
Raft master selection method and system in heterogeneous environment
CN116319277A
Computer cluster with adaptive quorum rules
US11210187B1
System and method for augmenting consensus election in a distributed database
US20170032007A1
Cited By
Cluster node data processing method and system based on edge side
CN120110889A
Distributed single-chip microcomputer state synchronization management method
CN120812071A