High-availability hierarchical distributed storage system with strong data consistency
By designing a shared storage architecture and multi-node identity hierarchy mechanism in a distributed storage system, combined with the optimal conversion method of cold backup nodes, the problem of difficult to achieve high availability in distributed storage systems under the requirements of strong consistency and low latency is solved, and high availability and data consistency of the system in the event of failure are achieved.
Patent Information
- Application Number
- CN202411792091.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-12-06
AI Technical Summary
While ensuring strong consistency and low latency, existing distributed storage systems are difficult to achieve high availability, especially when node failures, the availability and data consistency of the system are difficult to meet at the same time.
A highly available hierarchical distributed storage system with strong data consistency is designed, adopting a shared storage architecture, and the node identities are divided into main nodes, hot standby nodes, cold standby nodes and offline nodes. The best choice method is used to convert the hot standby node by preset cold standby nodes, and efficient conversion of the cold standby nodes is achieved to ensure that the system can still maintain high availability and strong consistency in the event of failure.
A distributed storage system with high availability nodes larger than two nodes under strong consistency and low latency requirements is realized. It can maintain data consistency and system availability when system failures are minimized, and ensure low latency of read and write requests.
Smart Images

Figure CN120091028A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a highly available hierarchical distributed storage system with strong data consistency and a method for preferentially selecting a cold standby node to be converted into a hot standby node based on the system. Background Art
[0002] The technology of distributed system consistency and availability is a relatively mature technology, and there are many different solutions in the industry. However, due to the limitations of the CAP basic principle (Consistency, Availability, and Partition tolerance), all solutions cannot avoid the trade-off between consistency and availability. The CAP principle states that a distributed system cannot simultaneously satisfy consistency, availability, and partition tolerance, and partition tolerance is the basis of a distributed system. Therefore, in fact, only a choice can be made between consistency and availability. The more refined PACELC theorem (Consistency, Availability, Partition Tolerance, Eventual Consistency, Low Latency) further considers the CAP principle. When a network partition occurs, as described by the CAP principle, it is necessary to balance consistency and availability. When there is no network partition, the relationship between consistency and latency also needs to be considered. For example, if every step of a distributed system is synchronized to all nodes in the cluster, strong consistency can definitely be guaranteed, but the latency will be so high that it is unacceptable for practical applications. On the contrary, if there is no synchronization at all, although the latency is low, the consistency cannot be guaranteed, and it is also difficult to apply to actual scenarios.
[0003] Currently, the solutions to the CAP problem in the industry are generally divided into AP and CP, that is, on the basis of ensuring partition tolerance, they tend more towards availability or consistency. Products based on the CP direction, such as the well-known distributed consistency system ZooKeeper, which is based on the variant Zab of the classic distributed algorithm Paxos, ensures partition tolerance and strong consistency of write operations. However, for read operations, it only guarantees a slightly weaker sequential consistency. At the same time, it supports write operations on the leader node and read operations on all nodes, which ensure partition tolerance and strong consistency of read and write operations. However, the cost is that both read and write operations must pass through the leader node, and strict strong consistency read operations cannot be performed on other nodes, let alone write operations. There are also some products that further relax the requirements for consistency, becoming eventual consistency, causal consistency, read-your-own-writes consistency, etc. The effects of these consistencies are strongly related to their respective application scenarios, such as some implementations of Memcached and MySQL master-slave replication. Products based on the AP direction are generally not called distributed consistency systems because their design goals are to give up consistency to a certain extent and are often used in scenarios that are not sensitive to consistency, such as the comment and like system of social websites. A well-known product is the distributed database Redis, which gives up consistency and is more inclined to ensure that all nodes are available at all times.
[0004] In the above examples, the strictness of consistency is usually strongly related to the business logic and cannot be covered by a single standard. Existing solutions based on the Paxos protocol and its variant algorithms are difficult to be directly applied to distributed storage software that requires strong consistency of read and write and low latency, because their read and write performance will be directly affected by the majority consistency requirement of Paxos, resulting in each read and write operation having to go through the consensus of the distributed cluster, which will involve several network communications and local IOs. For distributed storage software, such latency is unacceptable. And other relatively low-latency solutions cannot be used for strict distributed storage because they cannot meet strong consistency, otherwise it may cause serious problems such as data loss or inconsistency.
[0005] In the field of distributed storage, a more common solution is to use the distributed locks or distributed leases implemented by these CP systems to select a master node (Active) in the cluster. The master node provides read and write services externally and synchronizes data with the standby node. Since only communication between two nodes is required and there are caches inside both the master and standby nodes, it is not necessary to perform data persistence after each write. In this way, a two-node highly available and strongly consistent distributed storage system with strong consistency between the master and standby and low read and write latency is formed. Such a distributed storage system can tolerate at most one node failure without affecting service availability, while ensuring data is not lost. Only when both nodes fail simultaneously, the service becomes unavailable and there is a risk of data loss. Here, "simultaneously" means that the time difference between the failures is less than the time it takes for the master or standby to detect the failure and persist the cached data.
[0006] If the requirement for high availability continues to increase, simply increasing the number of standby nodes and replicating the synchronization link between the master and standby is not feasible. This will increase the amount of synchronized data for a single write request and also increase the probability of failure. When the synchronization fails, in order to ensure strong consistency, data persistence must be performed, which in turn increases the pressure on the storage medium. This simple solution will significantly increase the network traffic of the distributed system and the write pressure on the storage medium, consume a large amount of resources, and may even cause flooding. At the same time, the read and write latency will also increase significantly, seriously affecting performance and stability in actual application scenarios. Therefore, it is necessary to propose a design solution for a distributed storage system with more than two highly available nodes under the requirements of strong consistency and low latency. Summary of the Invention
[0007] This application discloses a highly available hierarchical distributed storage system with strong data consistency and a method for preferentially selecting nodes when cold standby nodes are converted to hot standby nodes based on the system.
[0008] In a first aspect, this application discloses a highly available hierarchical distributed storage system with strong data consistency, the system includes:
[0009] The hierarchical distributed storage system is based on a shared storage architecture;
[0010] Multiple nodes of the hierarchical distributed storage system share a memory to access data;
[0011] The nodes of the hierarchical distributed storage system include four identities: master node, hot standby node, cold standby node, and offline node. When a cold standby node is converted to a hot standby node, based on a preset method for preferentially selecting nodes when cold standby nodes are converted to hot standby nodes, the cold standby node is converted to a hot standby node.
[0012] In an exemplary embodiment of the present disclosure, the system includes:
[0013] When the node is the master node, it is used to provide read and write services;
[0014] When the node is a hot standby node, it is used to synchronize data with the master node and be promoted to the master node based on a preset selection method for converting the hot standby node to the master node;
[0015] When the node is a cold standby node, it has no data storage and is used to convert to a hot standby node by a preferred selection method when preset cold standby nodes are converted to hot standby nodes;
[0016] When the node is an offline node, it is in a fault state.
[0017] In an exemplary embodiment of the present disclosure, the system includes:
[0018] When the node is the master node, it is used to provide read and write services;
[0019] When the node is a hot standby node, it maintains the internal service running state to keep the data consistent with the master node.
[0020] In an exemplary embodiment of the present disclosure, the system includes:
[0021] Between the master node and the hot standby node, only the master node is used to provide read and write services through components that meet distributed mutual exclusion.
[0022] In an exemplary embodiment of the present disclosure, the system includes:
[0023] There is also a preset cache between the master node and the hot standby node, and the cache is used to save the data written by the client in time sequence in the cache.
[0024] In an exemplary embodiment of the present disclosure, the system includes:
[0025] The data written by the client in time sequence saved in the cache ensures strong consistency of the cache data through strong synchronization;
[0026] The data written by the client in time sequence saved in the cache is saved to the shared storage asynchronously.
[0027] In an exemplary embodiment of the present disclosure, the system includes:
[0028] When the master node or the hot standby node fails, the hot standby node is converted based on a preset selection method for converting cold standby nodes to hot standby nodes.
[0029] In a second aspect, the present application shows a preferred selection method for converting cold standby nodes to hot standby nodes based on the system, and the method includes:
[0030] Manually configure weights for each operation dimension index of each server in the distributed cluster, and continuously monitor the operation dimension indexes during operation;
[0031] When there is a need to convert a cold standby node to a hot standby node, calculate the coefficient of variation of each operation dimension index in the cluster respectively;
[0032] Based on the coefficient of variation, determine the candidate cold standby node to be converted, and calculate the expected value of the index load when converting the candidate cold standby node to a hot standby node;
[0033] Based on the expected value of the index load, calculate the coefficient of variation of each operation dimension index when converting the candidate cold standby node to a hot standby node, and obtain the coefficient of variation of the expected value of the load of each operation dimension index when the candidate cold standby node is converted to a hot standby node;
[0034] Perform weighted average processing on the coefficient of variation of the expected value of the load;
[0035] Select the cold standby node corresponding to the minimum value of the weighted average as the optimal choice for converting to a hot standby node.
[0036] In a third aspect, the present application discloses an electronic device, which includes: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the method described in any of the above aspects.
[0037] In a fourth aspect, the present application discloses a non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by the processor of the electronic device, the electronic device can execute the method described in any of the above aspects.
[0038] In a fifth aspect, the present application discloses a computer program product, when the instructions in the computer program product are executed by the processor of the electronic device, the electronic device can execute the method described in any of the above aspects.
[0039] The present application provides a highly available hierarchical distributed storage system with strong data consistency and a method for preferentially selecting cold standby nodes to convert to hot standby nodes based on the system. The hierarchical distributed storage system is based on a shared storage architecture; multiple nodes of the hierarchical distributed storage system share a memory to access data; the nodes of the hierarchical distributed storage system include four identities: master node, hot standby node, cold standby node, and offline node. When converting a cold standby node to a hot standby node, the cold standby node is converted to a hot standby node based on a preset method for preferentially selecting cold standby nodes to convert to hot standby nodes. The system of the present disclosure provides a high-availability and strong-consistency implementation solution for a distributed storage system based on shared storage. The system enables the consistency of an N-node distributed storage system based on shared storage to meet strong consistency, and the availability can satisfy at most N-1 faulty nodes without affecting the system's service provision, while ensuring low latency for read and write requests. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a schematic diagram of node conversion of a highly available hierarchical distributed storage system with strong data consistency according to the present application.
[0041] Figure 2 It is a schematic diagram of node relationships in an embodiment based on ZooKeeper of a highly available hierarchical distributed storage system with strong data consistency according to the present application.
[0042] Figure 3 It is a flowchart of steps of a method for preferentially selecting cold standby nodes to convert to hot standby nodes based on a highly available hierarchical distributed storage system with strong data consistency according to the present application.
[0043] Figure 4 It is a block diagram of an electronic device according to the present application.
[0044] Figure 5 It is a block diagram of a computer-readable storage medium according to the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0046] The present application provides a highly available hierarchical distributed storage system with strong data consistency and a method for preferentially selecting a cold standby node to be converted into a hot standby node based on the system. The hierarchical distributed storage system is based on a shared storage architecture; multiple nodes of the hierarchical distributed storage system share a memory to access data; the nodes of the hierarchical distributed storage system include four identities: a primary node, a hot standby node, a cold standby node, and an offline node. When the cold standby node is converted into a hot standby node, the cold standby node is converted into a hot standby node based on a preset method for preferentially selecting a cold standby node to be converted into a hot standby node. The system of the present disclosure provides a solution for achieving high availability and strong consistency of a distributed storage system based on shared storage. The system enables the consistency of an N-node distributed storage system based on shared storage to meet strong consistency, and the availability can satisfy at most N-1 faulty nodes without affecting the system's service provision, while ensuring low latency for read and write requests.
[0047] Embodiment 1:
[0048] Referring to Figure 1 , a schematic diagram of node conversion of a highly available hierarchical distributed storage system with strong data consistency according to the present application is shown. Specifically, the system may include:
[0049] The hierarchical distributed storage system is based on a shared storage architecture;
[0050] Multiple nodes of the hierarchical distributed storage system share a memory to access data;
[0051] The nodes of the hierarchical distributed storage system include four identities: a primary node, a hot standby node, a cold standby node, and an offline node. When the cold standby node is converted into a hot standby node, the cold standby node is converted into a hot standby node based on a preset method for preferentially selecting a cold standby node to be converted into a hot standby node.
[0052] In the embodiment of this example, the system includes:
[0053] When the node is a primary node, it is used to provide read and write services;
[0054] When the node is a hot standby node, it is used to synchronize data with the primary node and is promoted to the primary node based on a preset method for selecting a hot standby node to be converted into a primary node;
[0055] When the node is a cold standby node, it has no data storage and is used to be converted into a hot standby node by a preset method for preferentially selecting a cold standby node to be converted into a hot standby node;
[0056] When the node is an offline node, it is in a faulty state.
[0057] In the embodiment of this example, the system includes:
[0058] When the node is a primary node, it is used to provide read and write services;
[0059] When the node is a hot standby node, it maintains the internal service running state to keep the data consistent with that of the master node.
[0060] In the embodiment of this example, the system includes:
[0061] Between the master node and the hot standby node, a component that satisfies distributed mutual exclusion is used to provide read and write services only when the master node is used.
[0062] In the embodiment of this example, the system includes:
[0063] A preset cache is also included between the master node and the hot standby node, and the cache is used to save the data written by the client in time sequence in the cache.
[0064] In the embodiment of this example, the system includes:
[0065] The data written by the client in time sequence saved in the cache ensures strong consistency of the cache data through strong synchronization;
[0066] The data written by the client in time sequence saved in the cache is saved to the shared storage asynchronously.
[0067] In the embodiment of this example, the system includes:
[0068] When the master node or the hot standby node fails, the hot standby node is converted based on the method of preferentially selecting when converting the cold standby node to the hot standby node preset.
[0069] Embodiment 2:
[0070] In the embodiment of this example, the basis of the system of the present disclosure is a distributed storage system using a shared storage architecture. Multiple nodes of this distributed storage system share a memory to access data, which can ensure that the data persisted in the memory is strongly consistent.
[0071] There are five necessary features of the technical solution proposed by the present invention:
[0072] 1. The node identities include master, hot standby, and cold standby. The master is responsible for providing services externally; the hot standby does not provide services externally, but maintains the internal service running state and keeps the data consistent with that of the master. The hot standby can be switched to the master when the master fails; the cold standby does not provide services externally, but maintains the internal service running state and does not keep the data consistent with that of the master. The cold standby can be switched to the hot standby when the hot standby fails.
[0073] 2. Under normal operating conditions, there are at least two nodes, one master and one hot standby. Between them, a component that satisfies distributed mutual exclusion is used to meet the requirement that only the master can provide read and write services externally at any time.
[0074] 3. There are caches for the primary and hot standby. The data written by the client is stored in the cache in chronological order. The strong consistency of the cache data is ensured between the primary and hot standby through a strong synchronization method. The data in the cache will eventually be persisted to the shared storage asynchronously.
[0075] 4. Each node in the system has a mechanism to automatically detect faults and can automatically convert the node identity after a fault is detected.
[0076] 5. When the number of nodes maintaining consistent cache data is insufficient due to a fault in the primary or hot standby, the system will select at least one cold standby to be converted to a hot standby through an optimization algorithm to make up the quantity.
[0077] Next, specifically describe the node identity functions in the technical solution and the state transition process of the state machine.
[0078] There are four types of node identities: Active (primary), Hot Standby, Cold Standby, and Offline. Their functions are as follows:
[0079] Active: Provide read and write services.
[0080] Hot Standby: Synchronize data with the Active and can be promoted to Active.
[0081] Cold Standby: Has no data but is in a healthy state and can be promoted to Hot Standby.
[0082] Offline: Fault.
[0083] In the embodiment of this example, the state transition process of the node state machine is as follows:
[0084] 1. Before the state machine starts and after a fault, the state is Offline;
[0085] 2. After the state machine starts, necessary initialization work is carried out. If it fails, it remains in the Offline state and retries continuously until successful. After success, the state of the state machine becomes Cold Standby;
[0086] 3. When the state machine is in the Cold Standby state, it remains in this state until it discovers its own "upgrade" in some way and the state becomes Hot Standby;
[0087] 4. When the state machine is in Hot Standby, a component for distributed mutual exclusion is used to participate in the competition. The state machine that wins the competition will change its state to Active, and the state machine that does not win the competition will remain in Hot Standby. If it loses the ability to participate in the competition, its state will change to Cold Standby;
[0088] 5. When the state machine is in Active, as long as it holds the above-mentioned mutual exclusion component and the component is valid, it will remain in Active; otherwise, its state will change to Hot Standby.
[0089] In the embodiments of this example, the cached data between the primary and the hot standby is sorted in chronological order to determine the order of writes by the client to the same location, and no overwrite write concurrency errors will occur. When the primary or the hot standby fails, there is a risk of losing the unpersisted data in the cache. After the new primary is generated, the cached data needs to be persisted before processing the client I / O; otherwise, data will be lost. Then, the caches of the new primary and the new hot standby are cleared to achieve strong consistency again. This method has similarities with Write-Ahead Logging, and the purpose of both is to reduce the amount of disk write operations and reduce the latency of client I / O. The difference is that Write-Ahead Logging usually needs to persist the log content and still requires disk I / O, while this method only stores it in the cache. The advantage is that the latency is further reduced, and the cost is an increased risk of data loss. This risk can be mitigated by means such as increasing the number of hot standbys and increasing the cache persistence frequency.
[0090] The cold standby node is a standby node that does not communicate with the primary and the hot standby for data synchronization. Therefore, even if any number of cold standby nodes are added, the network load between clusters will not increase significantly. The cold standby is different from the common "observer" in the field of distributed technology. The latter may be able to provide read services, but can never provide write services, and its identity is immutable; the cold standby cannot provide read and write services, but has the ability to provide read and write services after a state transition.
[0091] If the number of hot standbys is insufficient, multiple cold standbys can be converted into hot standbys, and the priority of the selected cold standby will be determined according to an optimization algorithm. The dimensions considered by the optimization algorithm include but are not limited to system load, custom weights, user preference configurations, etc.
[0092] In the embodiments of this example, the availability, consistency, and latency levels of the technical solution proposed by the present invention are as follows:
[0093] Strong consistency: The strong synchronization of the cache between the primary and the hot standby ensures the strong consistency of the unpersisted data, and the shared storage architecture ensures the strong consistency of the persisted data.
[0094] High availability: For a distributed storage system with one master, N hot spares, and M cold spares, at most N + M node failures are allowed simultaneously without affecting system availability. When and only when the master and N hot spares fail simultaneously, it is possible to lose the data written to the client cache. Here, "simultaneously" means that the time interval between failures is less than the time taken for the master or hot spare to detect the failure of the other party and persist the cache data.
[0095] Low latency: The write request of the client only needs to be written to the cache of the master and complete strong synchronization of a limited number of links to be considered successful, without the need to complete disk I / O. The number of synchronization links depends on the number of hot spares. And due to the existence of an optimization algorithm, the hot spares are usually the nodes with the lowest network communication load in the system, further reducing the impact of data synchronization on latency.
[0096] Embodiment Three:
[0097] In the embodiment of this example, as Figure 2 shown, the open-source software ZooKeeper for distributed consistency is used as the basis of this embodiment. ZooKeeper has several features that can effectively reduce the implementation difficulty of the proposed solution of the present invention. First, there are mature solutions in the industry to use ZooKeeper to implement distributed leases, and distributed leases can be used as a tool for mutual exclusion competition between the master and hot spares; second, ZooKeeper provides a type of ephemeral node, which requires the client session to remain active to exist. Once the client session loses activity for various reasons, the ephemeral node will be immediately deleted, and this feature can be used as a way for each node in the cluster to detect failures; third, ZooKeeper itself can also be used as a lightweight distributed storage system, and the data it maintains satisfies sequential consistency. Some distributed data read and write requirements of this embodiment can be directly implemented using ZooKeeper.
[0098] In the embodiment of this example, the state transition process of the state machine is combined to specifically describe how this embodiment implements the technical solution of the invention.
[0099] 1. Before the state machine is started and after a failure, it loses connection with ZooKeeper, the session expires, and the created ephemeral node is automatically deleted. At this time, the external state is Offline;
[0100] 2. After the state machine is started, it creates a session with ZooKeeper. After success, it creates an ephemeral node, which is called the alive node. The main function of the alive node is to identify the activity of its own state machine. Then it creates a persistent node, which is called the authority node. The main function of the authority node is to identify whether its own state machine is eligible to participate in the competition to become Active. After completing the above steps, the state becomes Cold Standby;
[0101] When the state machine is in Cold Standby, if it finds that the authority is marked with a specific value, it believes that it has the right to participate in the competition for Active, and the state changes to Hot Standby.
[0102] When the state machine is in Hot Standby, it participates in the competition for Active using a distributed lease. The state machine that wins the competition becomes Active, and the state machine that does not win the competition remains in Hot Standby. If the lease loses its activity during this period (usually due to losing connection with ZooKeeper), the state changes to Cold Standby.
[0103] When the state machine is in Active, as long as it holds the lease and the lease is valid, it remains Active; otherwise, the state changes to Hot Standby.
[0104] Both the alive node and the distributed lease have similar "volatile" characteristics: when the state machine loses connection with the ZooKeeper cluster, the alive node is deleted and the lease becomes invalid, which can be considered as a state machine failure.
[0105] Summarizing the above information stage by stage, in this embodiment, distributed mutual exclusion between the master and the hot standby is achieved through the distributed lease based on ZooKeeper, strong synchronization of cached data is implemented using Socket Channel, a mechanism for detecting node failures is achieved through the "volatility" of the alive node and the distributed lease, and the cold standby can be controlled to switch between the states of not providing services and providing services through the content information of the authority node. The specific solution will be described below.
[0106] In the embodiment of this example, three issues are further explained in detail: when and how the content of the authority node is set; when concurrent failures occur, how to recover from possible error states to the normal state; and how to design the optimization algorithm for the cold standby to hot standby conversion.
[0107] When the authority node is created, the default value of its content is the boolean value false, which means that it currently does not have the qualification to participate in the Active competition. When none of the state machines have the qualification, the state transition of the highly available cluster will be in a stalled state, and there will be no primary or hot standby. There are two possible scenarios for this situation. One is that all state machines are starting for the first time, and the content of all authority nodes is the default value false. The other is that all state machines with the authority node content of true fail simultaneously, leaving only several state machines with the authority node content of false. Once this situation occurs, a third party needs to intervene to authorize at least one state machine to have the qualification to compete and restart the automatic state transition process of the highly available cluster.
[0108] The third party is a service independent of this highly available system, and it is also highly available itself. In this embodiment, it can be a primary-backup highly available system based on distributed leases. The third party monitors the changes of the alive nodes and authority nodes of these state machines through the method of ZooKeeper listening and callback. When any alive node change is detected, the third party conducts an availability check on the highly available cluster. By checking the existence of all alive nodes and the content of the authority nodes, it can determine whether the highly available system still has the ability to perform state transitions. When the content of the authority nodes of all state machines with existing alive nodes is false, that is, when all non-failed state machines do not have the qualification to participate in the Active competition, the highly available system no longer has the ability to perform state transitions. When the third party discovers this scenario, it needs to intervene to recover. The specific method is to set the content of the authority node of one of the state machines with an existing alive node to true, so that this state machine can automatically change from cold standby to hot standby and then become the primary to resume service.
[0109] When the primary and the hot standby do not fail simultaneously, error recovery can be performed without the participation of a third party.
[0110] In the process where the main state machine is located, there is a communication thread that synchronizes data with the hot standby. This thread also maintains the heartbeat between the main and the hot standby. When this thread detects that the heartbeat times out, the main considers the hot standby to have failed. At this time, the main needs to persist the data first. If the persistence is successful, the synchronization thread will select a cold standby with an alive node and set the content of its authority node to true. Then the state machine of this cold standby will automatically convert to a hot standby, and the system will return to the stable state of one main and one hot standby. If the persistence or other steps fail, it indicates that the main has a fault. The state machine of the main releases the distributed lease and then converts to a hot standby. Depending on the impact of the fault, it may continue to convert to a cold standby or go offline. At this time, there are two possibilities for this highly available system: one is that both the main and the hot standby fail simultaneously, and the recovery method for this scenario has been described above; the other is that the hot standby has no fault or has recovered from a fault. Then the hot standby can win the lease competition and become the main, and at this time, there is at least one main in the system, which is the same as the situation where the main recovers through the synchronization thread in the above text.
[0111] In the embodiment of this example, the heartbeat maintained by the above-mentioned main synchronization thread between the main and the hot standby is also a way to implement a fault discovery mechanism. Compared with the method of judging based on the "volatile" nodes of ZooKeeper, the difference lies in whether the relationship between the thread for discovering faults and the thread for data synchronization is synchronous or asynchronous. The synchronous or asynchronous implementation method affects the performance of data consistency, and distributed storage systems with different requirements for consistency can choose different implementation methods.
[0112] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions involved are not necessarily required by the present application.
[0113] Refer to Figure 3 , which shows the step flowchart of a method for preferentially selecting a cold standby node to convert to a hot standby node in a highly available hierarchical distributed storage system based on strong data consistency according to the present application. The method includes:
[0114] Step S110: Manually configure weights for each running dimension index of each server in the distributed cluster and continuously monitor each running dimension index during operation;
[0115] Step S120: When there is a need to convert a cold standby node to a hot standby node, calculate the coefficient of variation of each running dimension index in the cluster respectively;
[0116] Step S130: Determine candidate conversion cold standby nodes based on the coefficient of variation, and calculate the expected value of the metric load when converting the candidate conversion cold standby nodes into hot standby nodes.
[0117] Step S140: Calculate the coefficient of variation of each running dimension metric when converting the candidate conversion cold standby nodes into hot standby nodes based on the expected value of the metric load, and obtain the coefficient of variation of the load expected value of each running dimension metric when the candidate conversion cold standby nodes are converted into hot standby nodes.
[0118] Step S150: Perform weighted average processing on the coefficient of variation of the load expected value.
[0119] Step S160: Select the cold standby node corresponding to the minimum weighted average value as the optimal choice for converting to a hot standby node.
[0120] In the embodiment of this example, consider several running dimension metrics of each server in the distributed cluster, including but not limited to the utilization rate of the incoming network bandwidth, the utilization rate of the outgoing network bandwidth, the remaining memory, and the CPU utilization rate. Manually configure their weights and continuously monitor the above metrics during operation. When a cold standby to hot standby scenario occurs, calculate the coefficient of variation of these metrics in the cluster respectively. The coefficient of variation is a dimensionless statistical value, and its mathematical meaning is the discrete level of the distribution of a certain metric in the cluster. The larger the value, the higher the degree of dispersion, that is, the more uneven the current load of this metric in the cluster. Then simulate the conversion of a candidate cold standby to a hot standby and calculate the expected value of the metric load in this case. In this way, the expected values of the loads after simulating the conversion of all candidate cold standbys can be obtained. Calculate the coefficient of variation of these load expected values again. The mathematical meaning is the impact of converting different cold standbys on the discrete situation of the distribution of this metric in the cluster. The larger the value, the greater the impact of the distribution of this cold standby on the discrete situation of this statistical value, that is, how much impact the selection of the cold standby will have on the load balancing of this metric. In this way, the coefficient of variation of the load expected value of all metrics can be calculated. Since the coefficient of variation is dimensionless, weighted averaging can be performed on it, and thus the overall load balancing situation of the cluster after converting a certain cold standby to a hot standby can be obtained. Select the cold standby corresponding to the minimum weighted average value, that is, select the cold standby that can make the load distribution of each metric in the cluster relatively the most balanced after converting it to a hot standby.
[0121] Optionally, the embodiment of the present application also provides an electronic device, including: a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements each process of the above method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0122] Embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium includes, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0123] Figure 4 FIG. 4 is a block diagram of an electronic device 800 shown in the present application. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0124] Referring to Figure 4 FIG. 4, the electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0125] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0126] The memory 804 is configured to store various types of data to support the operation of the device 800. Examples of these data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, images, videos, etc. The memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0127] The power supply component 806 provides power for various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0128] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0129] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0130] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power-on button, and a lock button.
[0131] The sensor assembly 814 includes one or more sensors for providing a status assessment of various aspects of the electronic device 800. For example, the sensor assembly 814 can detect the on / off state of the device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also detect a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and a change in the temperature of the electronic device 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0132] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 can access a wireless network based on communication standards, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast operation information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0133] In an exemplary embodiment, the electronic device 800 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0134] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the above instructions can be executed by a processor 820 of the electronic device 800 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0135] Figure 5FIG. 0 is a block diagram of a computer-readable storage medium 1900 shown in the present application. For example, the computer-readable storage medium 1900 can be provided as a server.
[0136] Referring to Figure 5 , the computer-readable storage medium 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above-described method.
[0137] The computer-readable storage medium 1900 may further include a power component 1926 configured to perform power management of the computer-readable storage medium 1900, a wired or wireless network interface 1950 configured to connect the computer-readable storage medium 1900 to a network, and an input / output (I / O) interface 1958. The computer-readable storage medium 1900 may operate based on an operating system stored in the memory 1932, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.
[0138] It should be noted that, in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0139] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0140] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.
[0141] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the embodiments of the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0142] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.
[0143] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0144] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0145] In addition, the functional units in the various embodiments of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0146] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0147] As described above, the foregoing are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.
Claims
1. A highly available hierarchical distributed storage system with strong data consistency, characterized in that: The system comprises: The hierarchical distributed storage system is based on a shared storage architecture; Multiple nodes of the hierarchical distributed storage system share a memory to access data; The nodes of the hierarchical distributed storage system include four identities: master node, hot standby node, cold standby node, and offline node. When the cold standby node is converted to a hot standby node, the cold standby node is converted to a hot standby node based on a preset cold standby node conversion method.
2. The system according to claim 1, characterized in that The system comprises: The node is used to provide read and write services when it is a master node; When the node is a hot standby node, it is used to synchronize data with the master node, and is promoted to the master node based on a method selected when the hot standby node is converted to the master node; When the node is a cold standby node, there is no data storage, and when the preset cold standby node is converted to a hot standby node, a preferred method is selected to convert the hot standby node; The node is in a fault state when it is an offline node.
3. The system according to claim 2, characterized in that The system comprises: The node is used to provide read and write services when it is a master node; When the node is a hot standby node, it maintains an internal service running state to keep data consistent with that of the master node.
4. The system according to claim 2, characterized in that The system comprises: The master node and the hot standby node provide read and write services when only the master node is used through components that meet distributed mutual exclusion.
5. The system according to claim 2, characterized in that The system comprises: A preset cache is also included between the master node and the hot standby node, and the cache is used to save the data written by the client in the cache in time sequence.
6. The system according to claim 5, characterized in that The system comprises: The time-series data written by the client stored in the cache is synchronized to ensure the strong consistency of the cached data; The time-sequential data written by the client stored in the cache is asynchronously saved to the shared storage.
7. The system according to claim 1, characterized in that The system comprises: When the master node or the hot standby node fails, the hot standby node is converted by selecting a preferred method based on the preset cold standby node conversion to the hot standby node.
8. A method for selecting the best cold standby node when converting to a hot standby node in a highly available hierarchical distributed storage system based on strong data consistency, characterized in that: The method comprises: Manually configure weights for each operational dimension indicator of each server in the distributed cluster, and continuously monitor the operational dimension indicators during operation; When there is a need to convert a cold standby node to a hot standby node, the coefficient of variation of each operating dimension indicator in the cluster is calculated respectively; Determine a candidate cold standby node for conversion based on the coefficient of variation, and calculate an expected value of an indicator load when converting the candidate cold standby node for conversion to a hot standby node; Calculate the coefficient of variation of each operating dimension indicator when the candidate cold standby node is converted to a hot standby node based on the expected value of the indicator load, and obtain the coefficient of variation of the expected value of the load of each operating dimension indicator when the candidate cold standby node is converted to a hot standby node; Performing weighted average processing on the coefficient of variation of the load expectation value; The cold standby node corresponding to the minimum weighted average value is selected as the preferred choice for converting to the hot standby node.
9. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the method according to claim 8 when executed by the processor.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to claim 8 is implemented.
Citation Information
Patent Citations
Method for realizing node standby and system
CN101958782A
Hypervisor i / o staging on external cache devices
CN103823786A
Cache cluster-based cache method and system
CN105739924A
State machine copying method, device and system and storage medium
CN111240899A
Strong consistency query method, device and system for read-write separation architecture service system
CN111797121A
Cited By
High-availability, tiered distributed storage system with strong data consistency
WO2026118862A1