Data processing method and device based on cluster network, electronic equipment and medium
By introducing a distributed lock mechanism and log caching mechanism into the cluster network, data inconsistency and loss caused by misjudgment of sentinel nodes in the cluster network are solved, and the accuracy and reliability of data processing are improved.
Patent Information
- Application Number
- CN202510015467.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-06
AI Technical Summary
In the cluster network, after the sentry mechanism is introduced, data inconsistency or data loss is prone to problems, especially in the case of network failure or instability, the sentry node may misjudgment the operating status of the main node, resulting in two main nodes in the cluster.
By introducing a distributed lock mechanism in the cluster network, it is ensured that when the master node fails, transaction requests can only be processed by one node, and transaction requests and operation statements are stored in the log cache, and data is updated after the new master node is elected.
It effectively avoids data inconsistency and data loss, and improves the accuracy and reliability of data processing in cluster networks.
Smart Images

Figure CN119946065A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communication technology, and in particular to a data processing method and device, electronic equipment, and medium based on a cluster network. Background Art
[0002] A cluster network refers to the network infrastructure used for communication and data exchange between nodes (or servers) in a computer cluster. For example, in the Redis service provided by a distributed system (each Redis service can be called a "node"), a Redis cluster can be formed to provide services to the terminal. In order to avoid multiple nodes being able to perform transaction operations on data at the same time (that is, adding, deleting, and modifying data), the master-slave mode is used to form the cluster (that is, the master node provides write operations, and the slave node provides read operations). However, in this mode, once the master node fails, one of the slave nodes needs to be manually promoted to the master node, otherwise the sub-cluster will be unavailable. Therefore, it is particularly important to reasonably and effectively manage each node within the cluster so that the cluster as a whole can provide high-availability and high-performance services to the outside world.
[0003] At present, the related technology usually introduces a sentinel mechanism for cluster networks to process node data, that is, the master and slave nodes are monitored by the set sentinel nodes, and when there is a problem with the operation of the master node, a new master node is selected to ensure the correctness of data processing in network communication. However, in the case of network failure or instability, the sentinel node is prone to misjudgment of the operating status of the master node, resulting in two master nodes in the cluster network, which is prone to data inconsistency (that is, the master nodes of different clusters may write the same data differently, resulting in data inconsistency) or data loss (that is, the new master node clears the write command executed on the old master node during the master-slave switch). In addition, the method proposed by the related technology to avoid the above problems by adjusting the parameters of the configuration file cannot solve the above problems well. Summary of the invention
[0004] The main purpose of the embodiments of the present application is to propose a data processing method and device, electronic device, and medium based on a cluster network, which can better solve the problem of data inconsistency or data loss when processing data in a cluster network that introduces a sentinel mechanism, and improve the accuracy of data processing in the cluster network.
[0005] To achieve the above-mentioned purpose, a first aspect of an embodiment of the present application proposes a data processing method based on a cluster network, which is applied to a cluster network, wherein the cluster network includes a first master node, at least one first slave node and a sentinel cluster, wherein the first master node and a first client are located in a first network partition of the cluster network, the first master node is communicatively connected to the first client, the sentinel cluster is located in a second network partition of the cluster network, the sentinel cluster is communicatively connected to the first client, the first slave node is located in the first network partition or the second network partition, and the total number of nodes contained in the first network partition is less than or equal to half of the total number of nodes contained in the cluster network, and the method includes:
[0006] When the sentinel cluster detects that the running state of the first master node at the first time point is a fault state, the first slave node of the third network partition in the cluster network is selected based on the sentinel cluster to determine the second master node; wherein the fault state is used to indicate that the communication between the first master node and the sentinel cluster is abnormal, the total number of nodes included in the third network partition is greater than or equal to half of the total number of nodes included in the cluster network, and the second network partition is the same as or different from the third network partition;
[0007] Use the other first slave nodes in the cluster network as second slave nodes, and complete the master-slave switching of the second master node at a second time point; wherein the second time point is later than the first time point;
[0008] When the first master node responds to the transaction request sent by the client at a third time point, the transaction request is executed based on a preset distributed lock; wherein the distributed lock is used to lock the data associated with the execution of the operation statement corresponding to the transaction request, and the third time point is later than the first time point;
[0009] The distributed lock is stored in a preset first log cache, and the operation statement corresponding to the transaction request is stored in a preset second log cache;
[0010] When the Sentinel cluster detects that the operating status of the first master node at the fourth time point is normal, the preset database of the second master node is updated based on the first log cache and the second log cache; wherein, the normal status is used to indicate that the communication between the first master node and the Sentinel cluster is normal, the fourth time point is later than the third time point, and the fourth time point is later than the second time point.
[0011] In some embodiments, updating data on a preset database of the second master node based on the first log cache and the second log cache includes:
[0012] Release the distributed lock in the first log cache through the first master node, and send the second log cache to the second master node;
[0013] Executing, by the second master node, an operation statement corresponding to the transaction request, to obtain the first transaction operation data at the third time point;
[0014] The preset database is updated based on the first transaction operation data.
[0015] In some embodiments, the method further comprises:
[0016] Using the first master node as the second slave node, and sending a full synchronization signal to all the second slave nodes based on the updated preset database;
[0017] Each of the second slave nodes responds to the full synchronization signal, and updates the data stored in each of the second slave nodes based on the preset database.
[0018] In some embodiments, when the sentinel cluster detects that the running state of the first master node at the first time point is a fault state, before performing node selection on the first slave node of the third network partition in the cluster network based on the sentinel cluster and determining the second master node, the method further includes:
[0019] When the first sentinel node in the sentinel cluster detects that the running state of the first master node at the first time point is a fault state, sending a master node fault voting request to other sentinel nodes in the sentinel cluster;
[0020] receiving, through the first sentinel node, a master node failure voting response returned by the other sentinel nodes, wherein the master node failure voting response is used to indicate a detection result of the other sentinel nodes on the running status of the first master node at the first time point;
[0021] When the number of the master node fault voting responses indicating that the operating state of the first master node is the fault state is greater than a preset number, it is determined that the operating state of the first master node at the first time point is the fault state.
[0022] In some embodiments, the step of performing node selection on the first slave node of the third network partition in the cluster network based on the sentinel cluster to determine the second master node includes:
[0023] Determine a sentinel master node from the sentinel nodes of the sentinel cluster based on a preset consensus algorithm;
[0024] The second master node is determined by performing node selection on the first slave node of the third network partition in the cluster network based on the sentinel master node.
[0025] In some embodiments, the step of performing node selection on the first slave node of the third network partition in the cluster network based on the sentinel master node to determine the second master node includes:
[0026] Obtaining a node operation status and node attribute parameters of each of the first slave nodes in the third network partition;
[0027] The first slave nodes in the third network partition are screened based on the node running status and the node attribute parameters to determine the second master node.
[0028] In some embodiments, when the sentinel cluster detects that the running state of the first master node at the fourth time point is normal, before updating the data of the preset database of the second master node based on the first log cache and the second log cache, the method further includes:
[0029] Each of the sentinel nodes of the sentinel cluster sends a PING command to the first master node based on preset frequency data, and determines a plurality of response data returned by the master node according to the PING command within a preset time period;
[0030] Detecting the operating status of the first master node according to the multiple response data;
[0031] When the second sentinel node in the sentinel cluster detects that the running status of the first master node at the fourth time point is normal, it is determined that the running status of the first master node at the first time point is normal.
[0032] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application proposes a data processing device based on a cluster network, which is applied to a cluster network, wherein the cluster network includes a first master node, at least one first slave node and a sentinel cluster, wherein the first master node and a first client are located in a first network partition of the cluster network, the first master node is communicatively connected to the first client, the sentinel cluster is located in a second network partition of the cluster network, the sentinel cluster is communicatively connected to the first client, the first slave node is located in the first network partition or the second network partition, and the total number of nodes contained in the first network partition is less than or equal to half of the total number of nodes contained in the cluster network, and the device includes:
[0033] A master node determination module, used for, when the sentinel cluster detects that the running state of the first master node at a first time point is a fault state, performing node selection on the first slave node of the third network partition in the cluster network based on the sentinel cluster to determine a second master node; wherein the fault state is used to indicate that the communication between the first master node and the sentinel cluster is abnormal, the total number of nodes included in the third network partition is greater than or equal to half of the total number of nodes included in the cluster network, and the second network partition is the same as or different from the third network partition;
[0034] A switching module, used to use the other first slave nodes in the cluster network as second slave nodes, and complete the master-slave switching of the second master node at a second time point; wherein the second time point is later than the first time point;
[0035] a request execution module, configured to execute the transaction request based on a preset distributed lock when the first master node responds to the transaction request sent by the client at a third time point; wherein the distributed lock is used to lock data associated with the execution of the operation statement corresponding to the transaction request, and the third time point is later than the first time point;
[0036] A storage module, used to store the distributed lock in a preset first log cache, and store the operation statement corresponding to the transaction request in a preset second log cache;
[0037] An update module is used to update data in a preset database of the second master node based on the first log cache and the second log cache when the Sentinel cluster detects that the operating status of the first master node at a fourth time point is normal; wherein the normal status is used to indicate that the communication between the first master node and the Sentinel cluster is normal, the fourth time point is later than the third time point, and the fourth time point is later than the second time point.
[0038] To achieve the above objectives, a third aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect is implemented.
[0039] To achieve the above-mentioned purpose, the fourth aspect of an embodiment of the present application proposes a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect is implemented.
[0040] The data processing method and device, electronic device, and medium based on the cluster network proposed in the embodiment of the present application, by using a distributed lock to execute a transaction request on the first master node in a faulty state, can make the transaction request executed on the old first master node after a brain split (i.e., data inconsistency or data loss problem) only be operated by one node, thereby avoiding data inconsistency caused by multiple nodes operating the same data at the same time. In addition, the present application stores the distributed lock of the first master node in a preset first log cache, stores the operation statement corresponding to the transaction request in a preset second log cache, and synchronizes the first log cache and the second log cache to the newly determined second master node to update the data of the preset database of the second master node, so that data loss can be effectively avoided. Therefore, the present application can better solve the problem of data inconsistency or data loss when processing data on a cluster network that introduces a sentinel mechanism, and improve the accuracy of data processing in the cluster network. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 This is a cluster network deployment diagram using a sentinel mechanism provided in an embodiment of the present application;
[0042] Figure 2A This is a schematic diagram of a cluster network provided in an embodiment of the present application after a split-brain occurs;
[0043] Figure 2B This is another schematic diagram of a cluster network provided in an embodiment of the present application after a brain split occurs;
[0044] Figure 3 is a flow chart of a data processing method based on a cluster network provided in an embodiment of the present application;
[0045] Figure 4 It is a schematic diagram of network partitioning in a cluster network provided in an embodiment of the present application;
[0046] Figure 5A This is a schematic diagram of the principle of subjective offline of cluster nodes in sentinel mode;
[0047] Figure 5B This is a schematic diagram of the principle of objective offline of cluster nodes in sentinel mode;
[0048] Figure 6 yes Figure 3 A flowchart of step S310 in FIG.
[0049] Figure 7 It is a logical diagram of the sentinel in the cluster in sentinel mode notifying the master node of changes;
[0050] Figure 8 yes Figure 6 A flowchart of step S620 in FIG.
[0051] Fig. 9 It is a timeline diagram of a post-split-brain failover provided in an embodiment of the present application;
[0052] Fig.10 yes Figure 3 A flowchart of step S350 in FIG.
[0053] Fig.11 It is a specific logical schematic diagram of the data processing method based on the cluster network provided in the embodiment of the present application;
[0054] Fig.12 is a structural diagram of a data processing device based on a cluster network provided in an embodiment of the present application;
[0055] Fig.13 It is a hardware structure diagram of the electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION
[0056] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0057] It should be noted that, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the specification, claims and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0059] First, some nouns involved in this application are analyzed:
[0060] Split-Brain: In the field of communication, it is a vivid metaphor, like "brain splitting", that is, one "brain" is split into two or more "brains". In a distributed system, the communication between nodes is disconnected due to network partition or node failure, which leads to data consistency problems. In a high-availability cluster, when multiple servers cannot detect each other due to network reasons within a specified time, they each form a new small-scale cluster, and in each small cluster, a new master node will be elected to provide independent services to the outside world.
[0061] Cluster deployment: refers to the collection of multiple servers to provide the same service, so that the client sees only one server. This deployment method uses multiple computers to perform parallel computing, which improves the computing speed and increases the reliability of the system through redundant deployment. Even if a server fails, the entire system can still operate normally. Among them, the physical machines in the cluster deployment may be deployed across regions. Once the network between regions fails, the brain split problem may occur.
[0062] Redis: It is a memory-based key-value pair storage database, which is widely used in scenarios such as database, cache or message middleware.
[0063] Redis cluster: refers to a cluster formed by multiple Redis services through the election of master and slave node services. In a Redis cluster, once the master node fails, one of the slave nodes needs to be upgraded to the new master node as soon as possible to ensure the normal operation of the Redis cluster. Since the master-slave mode of the Redis cluster does not have the function of actively electing and switching nodes, manual switching is required, which not only affects business operations but also increases operation and maintenance costs. Therefore, the Sentinel mode is introduced to automatically switch when the master node fails to ensure the high availability of Redis.
[0064] Redis split-brain: Redis split-brain refers to an abnormal situation that occurs during the operation of the Redis server, that is, a Redis cluster is split into two or more independently running Redis clusters, and communication and data synchronization are lost between them.
[0065] Redis master-slave mode: In the master-slave architecture of Redis, the master-slave mode is read-write separated. The master node provides write operations, while the slave node only provides read operations, and the master node will let the slave node synchronize its own data.
[0066] Sentinel mechanism: Usually deployed in a separate server or container to avoid sharing hardware resources with the main server, thereby reducing the risk of service interruption due to hardware failure. The Sentinel mechanism can use the Raft protocol, and the sentinel nodes should not be deployed on the same physical machine to prevent hardware failures from affecting all sentinel nodes. For high availability, sentinel nodes can form a cluster and monitor each other to ensure the robustness of the system. Through the above deployment strategies and configurations, it can be ensured that the Redis sentinel mechanism can respond quickly when a failure occurs and realize automatic failover, thereby improving the stability and availability of the Redis cluster.
[0067] Raft protocol: is a highly consistent, decentralized, and highly available distributed protocol used to solve consistency problems in distributed systems. The Raft protocol stipulates that nodes have three roles: Leader, Follower, and Candidate. Under normal circumstances, most nodes are in the Follower state, and the Leader is responsible for processing all client transaction requests. If the Leader goes down, the Follower will become a Candidate and elect a new Leader through voting.
[0068] Full synchronization: refers to the process of completely synchronizing all data in the master node to the slave node.
[0069] Incremental synchronization: The process of synchronizing only the newly added or changed parts of the business data to the target node.
[0070] Subjective offline: A single sentinel node periodically sends a PING command to the Redis master and slave nodes. If no reply is received within a certain period of time, the Redis node is considered to be subjectively offline.
[0071] Objective offline: Multiple sentinel nodes in the sentinel cluster regularly send PING commands to the Redis master and slave nodes. When most of the sentinel nodes do not receive a reply within a certain period of time, a collective decision (that is, combined with the majority principle) is made to determine that the Redis node is objectively offline.
[0072] Distributed lock: used to ensure mutually exclusive access between different processes or node services in a distributed system. This means that the scope of the lock can span multiple processes or node services, ensuring that only one client can access shared resources (such as table data in a database) at any time.
[0073] A cluster network refers to the network infrastructure used for communication and data exchange between nodes (or servers) in a computer cluster. For example, in the Redis service (each Redis service can be called a "node") provided by a distributed system, in order to solve single point failures, improve system availability, achieve load balancing and horizontal expansion, the same Redis service is often deployed on multiple physical machines in different regions (each Redis service is called a "node" in this article). In this way, a Redis cluster can be formed to provide services to the terminal. Since each node is equal and provides the same services, it often only takes one Redis service to process the client request of the terminal. In addition, in order to avoid multiple nodes being able to perform transaction operations on data at the same time (i.e., add, delete, and modify data), the master-slave mode is used to form a cluster (i.e., the master node provides write operations and the slave node provides read operations). However, in this mode, once the master node fails, one of the slave nodes needs to be manually promoted to the master node, otherwise the sub-cluster will be unavailable. This defect greatly reduces the availability of the system. Therefore, it is particularly important to reasonably and effectively manage each node within the cluster so that the cluster as a whole can provide high-availability and high-performance services to the outside world.
[0074] At present, the relevant technology usually introduces a sentinel mechanism for cluster networks to process node data, that is, the master and slave nodes are monitored by the set sentinel nodes. When there is a problem with the operation of the master node, a new master node is selected to ensure the correctness of data processing in network communication. Figure 1 As shown, Figure 1 The diagram of the cluster network deployment using the sentinel mechanism is shown. Among them, client 1 and client 2 can communicate with the master node in the cluster network, and the master node connects slave node 1 to slave node 3 in a master-slave mode. At the same time, at least one sentinel node (such as Figure 1 Sentinel nodes 1 to 3 in the cluster monitor each node in the cluster network in real time.
[0075] However, the Redis cluster is prone to the following two types of split-brain behaviors: (1) In the event of a network failure or instability, the communication between the Redis master node and the sentinel or slave node may be interrupted. At this time, the sentinel may mistakenly believe that the Redis master node has crashed. Then the sentinel starts the election process and selects a new master node from the existing slave nodes. If the network is restored at this time, the old Redis master node is still running, which will result in two master nodes. (2) If the Redis master node "pseudo-dies", that is, the Redis master node process is still there, but it is not responding. At this time, because the sentinel node's failover strategy is too sensitive to network anomalies, it is easy to elect a master node under the wrong circumstances, that is, because of the false failure, the Redis cluster will select another master node. In this way, the sentinel node is prone to misjudging the running status of the Redis master node, resulting in two master nodes in the cluster network, which is prone to data inconsistency (that is, master nodes in different clusters may write the same data differently, resulting in data inconsistency) or data loss (that is, the new master node will send a full synchronization (that is, master-slave replication) command to all slave nodes, asking all slave nodes to re-synchronize the full amount. Before this, the data on the instance will be cleared first, so during the master-slave switch, the write commands executed on the old master node will also be cleared).
[0076] Based on this, in order to solve the data inconsistency and data loss problems caused by the split-brain of the Redis cluster in sentinel mode, the solution of the related technology is to avoid the split-brain problem by adjusting the parameters of the configuration file. For example, the adjustment of the parameters of the configuration file may include: (1) min-slaves-to-write: the number of slave nodes communicating with the master node must be greater than or equal to the value of the master node, otherwise the master node refuses to write; (2) min-slaves-max-lag: the ACK message delay between the master node and the slave node must be less than this value, otherwise the master node refuses to write. For the parameter min-slaves-to-write, when the sentinel cluster believes that the master node has failed, it will re-initiate the election of a new node, and the election of the master node is based on the "majority principle", that is, more than half of the slave nodes must elect the same node for that node to become the master node. Therefore, min-slaves-to-write is generally set to be greater than half of the number of cluster nodes (such as setting min-slaves-to-write to 3 in a cluster with 5 nodes). Figure 2A and Figure 2B , two cases are listed respectively after the occurrence of brain split in a cluster with 5 clusters. Figure 2AAfter the split-brain occurs, since the number of nodes in cluster 2 is 3, cluster 2 will generate a new node. At this time, the number of slave nodes in cluster 1 that communicate with the master node is 2, which is less than the value min-slaves-to-write, so the master node is denied writes. Figure 2B After the brain split, since the number of nodes in cluster 2 is 2, which is less than half, no new nodes can be generated, while the number of nodes in cluster 1 is 3 (including the master node itself), indicating that the number of slave nodes in cluster 1 that communicate with the master node is equal to min-slaves-to-write, so the master node can write. For the parameter min-slaves-max-lag, because the data synchronization between the master node and the slave node is synchronized once every period of time. Assuming that the data synchronization time interval is 8 seconds, when min-slaves-max-lag is set to 10, that is, when the ACK message delay between the master node and the slave node is greater than or equal to 10 seconds, the master node refuses the write operation. At this time, if the min-slaves-max-lag requirement is met, the following scenarios will cause defects in this solution: assuming that the master node has a brief network failure for 9 seconds, after 8 seconds, the sentinel node will consider the master node to be offline and start failover; assuming that the failover time takes 4 seconds, when the master node recovers at the 9th second, the old master node can still provide write operations within the 9th to 12th seconds, causing the data generated during this period to be discarded when the full data is synchronized after the new master node is elected.
[0077] Thus, although the Redis split-brain problem can be avoided as much as possible through reasonable configuration of configuration items, it is difficult to solve it completely. In addition, it is difficult to weigh and dynamically adjust the parameters of the configuration file in real time according to the actual scenario. These two configuration items will also affect the availability of the cluster to a certain extent. For example, once a slave node hangs up and the currently available slave nodes are less than the value of the configuration item, even if the master node of the cluster is normal, it cannot provide services to the terminal.
[0078] Based on this, the embodiments of the present application provide a data processing method and device, an electronic device, and a medium based on a cluster network, which can better solve the problem of data inconsistency or data loss when processing data in a cluster network that introduces a sentinel mechanism, and improve the accuracy of data processing in the cluster network.
[0079] The data processing method based on cluster network provided in the embodiment of the present application relates to the field of communication technology, and can be applied to the terminal, can also be applied to the server side, and can also be software running in the terminal or the server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or a server cluster or distributed system composed of multiple physical servers, and can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN) and big data and artificial intelligence platforms; the software can be an application that implements the data processing method based on cluster network, etc., but is not limited to the above forms.
[0080] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network personal computers (Personal Computer, PC), minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0081] See also Figure 3 , Figure 3 It is an optional flow chart of the data processing method based on the cluster network provided in the embodiment of the present application. Figure 3 The method may specifically include but is not limited to steps S310 to S350. Figure 3 These five steps are introduced in detail.
[0082] Step S310, when the sentinel cluster detects that the running state of the first master node at the first time point is a fault state, the sentinel cluster performs node selection on the first slave node of the third network partition in the cluster network to determine the second master node;
[0083] Step S320, using other first slave nodes in the cluster network as second slave nodes, and completing master-slave switching of the second master node at a second time point;
[0084] Step S330, when the first master node responds to the transaction request sent by the client at the third time point, executes the transaction request based on the preset distributed lock;
[0085] Step S340, storing the distributed lock in a preset first log cache, and storing the operation statement corresponding to the transaction request in a preset second log cache;
[0086] Step S350, when the sentinel cluster detects that the operating status of the first master node at the fourth time point is normal, the preset database of the second master node is updated based on the first log cache and the second log cache.
[0087] It should be noted that the data processing method based on the cluster network provided in the present application can be applied to a cluster network (such as a Redis cluster), the cluster network includes a first master node, at least one first slave node and a sentinel cluster, the first master node and the first client are located in the first network partition of the cluster network, the first master node is communicatively connected to the first client, the sentinel cluster is located in the second network partition of the cluster network, the sentinel cluster is communicatively connected to the first client, the first slave node is located in the first network partition or the second network partition, and the total number of nodes contained in the first network partition is less than or equal to half of the total number of nodes contained in the cluster network.
[0088] It should be noted that, because there are many clients (which can be understood as APP terminals) and they are scattered in different regions, the clients mentioned in this application are not limited. In addition, if the sentinel cluster and the first master node are in the same network partition, the first master node will not be considered to be objectively offline, because the sentinel node determines whether the first master node is offline by periodically sending a PING heartbeat command to the first master node, and there will be no brain split phenomenon. The so-called brain split phenomenon refers to the fact that a cluster originally generated two or more clusters due to a short network failure. If the sentinel node is in the same network partition as the first master node, then the other slave nodes cannot communicate with the sentinel node, and the sentinel node cannot elect a new master node for these slave nodes. After this happens, the system can still provide normal transaction requests and read operations. The only problem is that the cluster becomes a single point for providing services. Therefore, the first master node and the sentinel cluster of this application are located in different network partitions.
[0089] It should be noted that if Figure 4As shown, the first master node and the first client are located in the first network partition of the cluster network. This is because the first master node that provides the transaction request must be in the same network partition as the first client, while the sentinel node and other slave nodes must be in another network partition. In this way, during the period when the first master node has a brain split, the first client cannot communicate normally with the sentinel node, and the first client cannot know who the new master node is, and will always access its default first master node, and cannot connect to the new master node, thereby generating the scenario problem to be solved by this application.
[0090] It should be noted that, combined with Figure 4 , since the first slave node (such as slave node 1 and slave node 2) may be in the same network partition as the old first master node, it may also be in the same network partition as the new second master node. In order to better apply this application, one condition must be guaranteed, that is, the total number of nodes in the partition where the first master node is located cannot exceed half of the total number of nodes contained in the cluster network. If it exceeds half, the network partition where the second master node is located cannot elect a new master node (this is because the principle of the sentinel mechanism to elect the master node is "more than half").
[0091] In step S310 of some embodiments, the first master node refers to the master node that has been currently selected. The running state is used to refer to the communication status of the first master node. The fault state is used to indicate that the communication between the first master node and the sentinel cluster is abnormal. The first network partition where the first master node is located is different from the third network partition, and the total number of nodes contained in the third network partition is greater than or equal to half of the total number of nodes contained in the cluster network, and the second network partition is the same as or different from the third network partition. The third network partition is used to select one from the first slave nodes contained therein as the second master node when a brain split occurs in the first master node. The second master node refers to the master node re-determined by the sentinel cluster through the automatic failover method.
[0092] It should be noted that when the sentinel cluster detects that the running state of the first master node at the first time point is a fault state, it can be determined that the first master node is objectively offline at the first time point, thereby generating a split brain situation. In order to avoid the problems caused by the split brain situation, the method provided by this application is adopted.
[0093] In some embodiments of the present application, before step S310, the method provided in the embodiment of the present application may further include but is not limited to the following steps:
[0094] When the first sentinel node in the sentinel cluster detects that the running state of the first master node at the first time point is a fault state, it sends a master node fault voting request to other sentinel nodes in the sentinel cluster;
[0095] Receiving, through the first sentinel node, a master node failure voting response returned by other sentinel nodes, where the master node failure voting response is used to indicate a detection result of other sentinel nodes on the running status of the first master node at a first time point;
[0096] When the number of master node fault voting responses indicating that the operating state of the first master node is a fault state is greater than a preset number, it is determined that the operating state of the first master node at the first time point is a fault state.
[0097] In the above embodiment, the present application determines the operating status of the first master node based on the sentinel cluster. Specifically, by default, the sentinel nodes in the sentinel cluster will periodically send PING commands to the first master node and the first slave node at a preset frequency (such as once per second or twice per second, etc.) to detect the online status of these nodes. Figure 5A As shown, when the first sentinel node in the sentinel cluster (Sentinel node 1 in the figure) detects that the running status of the first master node (the master node in the figure) at the first time point is in a faulty state, that is, there is no heartbeat reply to the PING command, the first master node is marked as subjectively offline (if the Redis master node is really offline, it is necessary to start the master node switch, including re-election, synchronization of data, these operations will consume a lot of resources, so the offline judgment of the master node should be treated with caution. It is possible that the network failure of the sentinel node causes the heartbeat detection of the Redis master node to fail, so the sentinel cluster needs to be introduced to make decisions). Then, Figure 5B As shown, a sentinel cluster is introduced to make a judgment together, that is, the first sentinel node (such as sentinel node 1 in the figure) can send a master node fault voting request to other sentinel nodes in the sentinel cluster, and then other sentinel nodes will send a PING command to the first master node (such as the master node in the figure). The master node fault voting response returned by other sentinel nodes is received by the first sentinel node, and the master node fault voting response is used to indicate the detection result of the other sentinel nodes on the running state of the first master node at the first time point. If more than half of the sentinel nodes judge that the first master node has been subjectively offline, that is, the number of master node fault voting responses indicating that the running state of the first master node is a fault state is greater than the preset number (the preset number at this time can be a preset value, or it can be half of the number of sentinel nodes contained in the sentinel cluster, without limitation), then the first master node is marked as objectively offline, that is, it is determined that the running state of the first master node at the first time point is a fault state.
[0098] In other words, the sentinel node can detect whether the current master node has failed through regular monitoring, and reach a consensus with other sentinel nodes that the master node has failed through negotiation. In this way, it can be arranged that the failure is not caused by the master node, but by the sentinel node that discovered the failure. This situation often occurs when the network of the sentinel node is isolated.
[0099] See also Figure 6 , Figure 6 It is a specific flow chart of step S310 provided in an embodiment of the present application. In some embodiments of the present application, step S310 specifically includes but is not limited to steps S610 to S620.
[0100] Step S610, determining a sentinel master node from the sentinel nodes of the sentinel cluster based on a preset consensus algorithm;
[0101] Step S620: Perform node selection on the first slave node of the third network partition in the cluster network based on the sentinel master node to determine the second master node.
[0102] In step S610 and step S620 of some embodiments, when it is determined that the first master node is in a faulty state, that is, it is marked as objectively offline and cannot work normally, the sentinel cluster can start automatic failover operations. Specifically, when the first master node fails, the corresponding first slave node synchronization connection is interrupted, and the master-slave replication stops. Furthermore, the sentinel nodes can use a preset consensus algorithm (such as the Raft algorithm) to elect a leader role as the sentinel master node. Furthermore, the sentinel master node is responsible for subsequent failover work. Figure 7 As shown, the failover work may specifically include: the sentinel master node selects a first slave node from the first slave nodes of the third network partition in the cluster network as the second master node (slave node 3 in the figure). The sentinel node can send the connection information of the new second master node to other slave nodes, allowing the slave nodes to execute the master-slave replication command, establish a connection with the new second master node, and perform data replication of all data of the new second master node. In addition, the sentinel node will also notify the client program to transfer to the new master node, allowing the client to find the new master node (slave node 3 in the figure) for writing.
[0103] It should be noted that when the client is initialized, it can obtain the master node address of the current Redis service by connecting to the sentinel node. Among them, since the first client and the first master node are in a network partition and the first master node fails, the sentinel node cannot connect to the first client, so the first client cannot know the current new master node address.
[0104] See also Figure 8 , Figure 8 It is a specific flow chart of step S620 provided in an embodiment of the present application. In some embodiments of the present application, step S620 specifically includes but is not limited to steps S810 to S820.
[0105] Step S810, obtaining the node operation status and node attribute parameters of each first slave node in the third network partition;
[0106] Step S820: Screen the first slave nodes in the third network partition based on the node operation status and node attribute parameters to determine the second master node.
[0107] In step S810 of some embodiments, the node operation state is used to indicate whether the corresponding first slave node is in a fault state or a normal state. The node attribute parameters are used to indicate parameters related to the attributes of the first slave node, and the node attribute parameters may include slave node priority (slave-priority), slave node replication offset (replication offset, used to indicate the progress of the corresponding slave node copying data, the larger the offset, the newer the data of the slave node, the closer to the data state of the original master node), node operation identification (i.e., runid, which is a random identifier automatically generated when each Redis node is started, and the runid corresponding to each Redis node is different), etc.
[0108] In step S820 of some embodiments, because the sentinel cluster will exclude those nodes that are already considered to be faulty, these slave nodes will not be considered as candidates for the new master node. Therefore, the sentinel master node can first screen the first slave nodes in the third network partition based on the node running status, filter out the first slave nodes in the faulty state, and screen the remaining first slave nodes based on the node attribute parameters. Specifically, the first slave node with the highest slave node priority setting can be selected as the second master node. If there are multiple first slave nodes with the same slave node priority, the sentinel master node can select the first slave node with the largest slave node replication offset as the second master node. If there are at least two first slave nodes with the same slave node replication offset, the sentinel master node can select the first slave node with the smallest or largest runid as the new second master node.
[0109] It should be noted that the present application can also flexibly select node attribute parameters, so that the method of determining the second master node can be flexibly adjusted without limitation.
[0110] It should be noted that in order to better understand the data processing of different nodes in the cluster when a split-brain occurs, and why a Redis cluster split-brain occurs and causes data loss, refer to Fig. 9 , Fig. 9 The figure shows the timeline of failover after split-brain. Node 1 is the first master node, and Node 2 is the newly determined second master node. After the second master node is determined, a full synchronization (or master-slave replication) command will be sent to all slave nodes to let all slave nodes re-synchronize. Full synchronization will first clear the data on the node. Therefore, during the period from when the old master node is determined to be offline to when the old master node resumes communication with the cluster where the new master node is located (i.e. Fig. 9The data executed during the time period [t3, t8] in the figure will be lost, so this is the specific reason for the data loss. During the period [t5, t7], because the client and the sentinel node cannot communicate normally, the client can only access its default first master node, that is, the first master node can receive the client's transaction request for data writing.
[0111] In addition, Fig. 9 In the time period [t3, t8] in the old master node, the client's transaction request (i.e., the client's request to add, delete, or modify database data) will be executed. The result of the execution is that the cached data in the Redis service will be modified first, and then Redis will persist the data, that is, Redis will write the data to the database (e.g., MySQL). This is because the client that is in a network partition with the first master node can only connect to node 1 but not node 2 in the time period [t5, t7]. From the analysis of data loss in the previous article, it can be seen that when the new master node allows all slave nodes (the old master node has been downgraded to a slave node) to re-synchronize the full data, the data executed by the old master node in the time period [t3, t8] will be cleared, while the data that has been persisted to the database (e.g., MySQL) in the time period [t3, t8] will be retained. Therefore, when the new master node persists the data in the Redis cache to the database (e.g., MySQL), data inconsistency will occur. In this way, if the clients in other regions can communicate normally with the cluster where Node 2 is located (that is, the cluster where the Sentinel is located) at the beginning, then when Node 2 is elected as the new master node, the Sentinel will immediately notify these clients that Node 2 is the new master node. Then, during the period [t5, t7], these clients will not connect to Node 1 but will connect to Node 2.
[0112] Based on this, the purpose of the solution of this application is to save the data in the time period [t3, t8] and lock the data executed in the time period [t3, t8] to prevent tampering by the newly elected master node of other network partitions. After the communication between the network partitions is restored to normal, the lock is released and the new master node is allowed to update the data in the time period [t3, t8] first, and then all slave nodes are allowed to perform a full synchronization operation.
[0113] In step S320 of some embodiments, after the second master node is determined, other first slave nodes in the cluster network can be used as second slave nodes, and the second master node completes the master-slave switch at a second time point. The second time point is used to represent the time point when the second master node completes the master-slave switch, and the second time point is later than the first time point. Fig. 9 As shown, the first time point may be t1, and the second time point may be t5.
[0114] It should be noted that, in some other embodiments, the present application can also divide the Redis master node into a separate network partition, and the remaining Redis slave nodes and the sentinel cluster form another network partition, and the data between the two network partitions cannot be communicated.
[0115] In step S330 of some embodiments, the third time point is any time point in the time period from when the sentinel cluster determines that the first master node is objectively offline to when the first master node reestablishes a connection with the sentinel cluster. Fig. 9 As shown, the third time point can be any time point in the time period [t3, t8]. At this time, when the first master node still responds to the transaction request sent by the client at the third time point, each time a transaction request is received, the first master node applies for a distributed lock, and does not release the lock after the transaction request is executed. Among them, the distributed lock is used to lock the data associated with the execution of the operation statement corresponding to the transaction request, and the third time point is later than the first time point.
[0116] It should be noted that the distributed lock applied by the first master node can lock the data after the operation statement corresponding to the transaction request is executed, so that other nodes except the first master node cannot perform addition, deletion, or modification operations on it. Among them, the distributed lock can be provided by Redis, and in the Redis cluster, no matter which Redis node generates the distributed lock, its application scope is the entire cluster.
[0117] In step S340 of some embodiments, further, the first master node can store the distributed lock to the preset first log cache (such as log cache S1), and store the operation statement corresponding to the transaction request to the preset second log cache (such as log cache F1), that is, the first master node can write the transaction request into the log cache F1 in the form of an SQL statement. Among them, log cache F1 stores SQL transaction statements (i.e., transaction operations), and log cache S1 stores the distributed locks of each statement in F1. When the first master node performs a transaction operation on the database, such as the operation statement corresponding to the transaction request is "update table1 set name="zhangsan"where id="7", at this time, the row of data "id="7" in the corresponding table table1 is locked, resulting in that other nodes except the first master node cannot perform addition, deletion, and modification operations on it. When the first master node sends the log cache F1 to the new Redis master node and releases the distributed lock, the new Redis master node can update the transaction operation of F1 (such as: update table1 set name="zhangsan"where id="7") to its own database. In addition, the new Redis master node no longer needs to save the log cache after updating F1's transaction operation.
[0118] In step S350 of some embodiments, when the sentinel cluster detects that the operating state of the first master node at the fourth time point is normal, that is, the sentinel node resumes communication with the first master node, then the second master node can also communicate normally with the first master node, so that the second master node can update the data of the preset database of the second master node based on the first log cache and the second log cache. The normal state is used to indicate that the communication between the first master node and the sentinel cluster is normal, the fourth time point is later than the third time point, and the fourth time point is later than the second time point. Fig. 9 As shown, the fourth time point may be t8.
[0119] In some embodiments, before step S350, the method provided by the present application may further include:
[0120] Each sentinel node of the sentinel cluster sends a PING command to the first master node based on the preset frequency data, and determines a plurality of response data returned by the master node according to the PING command within a preset time period;
[0121] Detecting the operating status of the first master node according to the multiple response data;
[0122] When the second sentinel node in the sentinel cluster detects that the running status of the first master node at the fourth time point is normal, it is determined that the running status of the first master node at the first time point is normal.
[0123] In the above embodiment, the first sentinel node and the second sentinel node both refer to any node in the sentinel cluster, which may be the same or different. Each sentinel node of the sentinel cluster will regularly send a PING command to the first master node at a preset frequency (such as once per second or twice per second, etc.). When the sentinel node detects that the first master node is online, the first master node can be returned to the Redis cluster as a slave node.
[0124] See also Fig.10 , Fig.10 It is a specific flow chart of step S350 provided in an embodiment of the present application. In some embodiments of the present application, step S350 specifically includes but is not limited to steps S1010 to S1030.
[0125] Step S1010, releasing the distributed lock in the first log cache through the first master node, and sending the second log cache to the second master node;
[0126] Step S1020, executing the operation statement corresponding to the transaction request through the second master node to obtain the first transaction operation data at the third time point;
[0127] Step S1030: updating a preset database based on the first transaction operation data.
[0128] In steps S1010 to S1030 of some embodiments, after determining that the first master node has returned to the Redis cluster as a slave node, the first master node may release the distributed lock in the first log cache and send the second log cache to the second master node. Further, the second master node may execute the operation statement corresponding to the transaction request in the second log cache, obtain the first transaction operation data at the third time point, and update the first transaction operation data to the preset database of the second master node.
[0129] In some embodiments, after step S350, the method provided in the embodiment of the present application may further include:
[0130] The first master node is used as the second slave node, and a full synchronization signal is sent to all second slave nodes based on the updated preset database;
[0131] Each second slave node responds to the full synchronization signal and updates the data stored in each second slave node based on a preset database.
[0132] It is understandable that after updating these first transaction operation data to the preset database of the second master node, the sentinel node can notify all current slave nodes to copy all data of the second master node in the preset database.
[0133] In a specific embodiment, if Fig.11 As shown, Fig.11 A complete logical diagram of the data processing method based on the cluster network provided by the present application is shown. In the process of the Redis cluster (i.e., the cluster network of the present application) providing services to the client, the Redis node (i.e., the master node and slave node included in the present application) and the sentinel node are respectively monitoring the network online status of the node. The data processing logic is described below according to the Redis node end and the sentinel node end.
[0134] Among them, for the Redis node end (i.e., the first master node of this application, referred to as Redis master node R1 below), the following steps may be specifically included: First, the heartbeats of the Redis master node R1 and the sentinel cluster are stopped, and it is determined that Redis brain split occurs. However, at this time, the Redis master node R1 can still communicate with the client. Then, the Redis master node R1 continues to respond to the client's transaction request. Each time a transaction request is received, the Redis master node R1 applies for a distributed lock, does not release the lock after executing the transaction request, and saves the distributed lock in the log cache S1, and writes the transaction request in the form of an SQL statement to the log cache F1. Furthermore, when the Redis master node R1 continues to respond to the client's transaction request, the sentinel node will also regularly send a PING command to the Redis master node R1 at a frequency of once per second to determine whether it is online. The Redis master node R1 will simultaneously monitor the heartbeat changes with the sentinel node, waiting to reconnect to the sentinel node.
[0135] Among them, for the sentinel node side, the following steps may be specifically included: First, when the first sentinel node detects that the heartbeat with the Redis master node R1 has stopped, the first sentinel node negotiates with other sentinel nodes to reach a consensus that the majority agrees on the failure of the Redis master node, that is, the sentinel cluster determines that the Redis master node R1 is objectively offline, then the sentinel cluster elects a sentinel master node through the Raft algorithm to be responsible for the failover of the Redis cluster. Then, the sentinel master node elects a slave node as the new Redis master node R2 (that is, the second master node of this application, referred to as Redis master node R2 below) in the Redis network partition with more than half of the nodes. After that, the sentinel node will notify all slave nodes to copy all data of the new Redis master node R2. Finally, under the notification of the sentinel node, the new Redis master node R2 takes over the position of the master node of the old Redis master node R1 and continues to respond to the client's transaction requests.
[0136] Furthermore, the sentinel master node completes the failover of the Redis cluster, and then waits for the network communication of the two network partitions to resume interconnection. Specifically, it includes: First, step 1: the sentinel node will regularly send a PING command to the old Redis master node R1 at a frequency of once per second. When the sentinel node detects that the old Redis master node R1 is online, the old Redis master node R1 will return to the Redis cluster as a slave node (the old Redis master node R1 is renamed slave node R1). Then, slave node R1 releases the distributed lock of log cache S1, and sends log cache F1 to the new Redis master node R2. The new Redis master node R2 updates the transaction operation of F1 to this node service (ie, Redis master node R2). After that, the sentinel node notifies all slave nodes to copy all data of the new Redis master node R2.
[0137] The data processing method based on cluster network provided by the embodiment of the present application can effectively improve the problems caused by the brain split of the cluster. The beneficial effects brought by the improvement may include: (1) Improving data consistency: By using distributed locks, the transaction request executed on the old master node after the brain split can only be operated by one node, avoiding data inconsistency caused by multiple nodes operating the same data at the same time. (2) Avoiding data loss: After the brain split occurs, the transaction request executed on the old master node will be saved in the log cache. After the new master node is elected, the old master node will synchronize this part of the cache log to the new master node. (3) Improving the availability of the system: After the improvement of the present application, even if the brain split occurs, the ability of Redis to provide services is not affected from the outside world. (4) Improving the robustness of the system: In the original solution to cluster brain split, if the slave node crashes, it may affect the availability of the master node to provide services to the client, but the solution of the present application can avoid the negative impact of the slave node crash on the master node. Based on this, when a brain split occurs, by adding a distributed lock to each transaction operation, it is ensured that the same data can only be operated by one master node during the cluster split, avoiding the master nodes of different clusters from performing different writes on the same data, which ultimately leads to data inconsistency. In addition, when a brain split occurs, the transaction operation log executed by the old master node is saved, so that the new master node can synchronously update the part of the data of the transaction operation executed by the old master node, avoiding data loss. Therefore, the present application can better solve the problem of data inconsistency or data loss when processing data in a cluster network that introduces a sentinel mechanism, and improve the accuracy of data processing in the cluster network.
[0138] See also Fig.12 , Fig.12 : is a schematic diagram of the structure of a data processing device based on a cluster network provided in an embodiment of the present application, and the device specifically includes:
[0139] The master node determination module 1210 is used to select a first slave node of a third network partition in the cluster network based on the sentinel cluster to determine a second master node when the sentinel cluster detects that the running state of the first master node at a first time point is a fault state; wherein the fault state is used to indicate that the communication between the first master node and the sentinel cluster is abnormal, the total number of nodes included in the third network partition is greater than or equal to half of the total number of nodes included in the cluster network, and the second network partition is the same as or different from the third network partition;
[0140] The switching module 1220 is used to use other first slave nodes in the cluster network as second slave nodes, and complete the master-slave switching of the second master node at a second time point; wherein the second time point is later than the first time point;
[0141] The request execution module 1230 is configured to execute the transaction request based on a preset distributed lock when the first master node responds to the transaction request sent by the client at a third time point; wherein the distributed lock is used to lock the data associated with the execution of the operation statement corresponding to the transaction request, and the third time point is later than the first time point;
[0142] The storage module 1240 is used to store the distributed lock in a preset first log cache, and store the operation statement corresponding to the transaction request in a preset second log cache;
[0143] Update module 1250 is used to update data in the preset database of the second master node based on the first log cache and the second log cache when the Sentinel cluster detects that the operating status of the first master node at the fourth time point is normal; wherein the normal status is used to indicate that the communication between the first master node and the Sentinel cluster is normal, the fourth time point is later than the third time point, and the fourth time point is later than the second time point.
[0144] It should be noted that the cluster network-based data processing device of the embodiment of the present application is used to implement the cluster network-based data processing method of the above-mentioned embodiment. The cluster network-based data processing device of the embodiment of the present application corresponds to the aforementioned cluster network-based data processing method. For the specific processing process, please refer to the aforementioned cluster network-based data processing method, which will not be repeated here.
[0145] The embodiment of the present application also provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the cluster network-based data processing method of the embodiment of the present application when executing the computer program. The electronic device can be any intelligent terminal including a tablet computer, a car computer, etc.
[0146] See also Fig.13 , Fig.13The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:
[0147] The processor 1310 may be implemented by a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0148] The memory 1320 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1320 can store an operating system and other application programs. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 1320, and the processor 1310 calls and executes the cluster network-based data processing method of the embodiment of the present application;
[0149] Input / output interface 1330, used to implement information input and output;
[0150] The communication interface 1340 is used to realize the communication interaction between the device and other devices. The communication can be realized through a wired manner (such as USB, network cable, etc.) or a wireless manner (such as mobile network, WIFI, Bluetooth, etc.);
[0151] bus 1350 , which transmits information between the various components of the device (e.g., processor 1310 , memory 1320 , input / output interface 1330 , and communication interface 1340 );
[0152] The processor 1310 , the memory 1320 , the input / output interface 1330 , and the communication interface 1340 are connected to each other in communication within the device via a bus 1350 .
[0153] The embodiment of the present application further provides a computer-readable storage medium, which stores a computer program, and the computer program is used to enable a computer to execute the data processing method based on the cluster network in the above embodiment.
[0154] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0155] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0156] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0157] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0158] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.
[0159] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0160] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0161] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0162] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0163] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0164] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.
[0165] The preferred embodiments of the present invention are described with reference to the accompanying drawings, but the scope of the present invention is not limited by the above. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A data processing method based on a cluster network, characterized in that: Applied to a cluster network, the cluster network includes a first master node, at least one first slave node and a sentinel cluster, the first master node and a first client are located in a first network partition of the cluster network, the first master node is in communication connection with the first client, the sentinel cluster is located in a second network partition of the cluster network, the sentinel cluster is in communication connection with the first client, the first slave node is located in the first network partition or the second network partition, and the total number of nodes included in the first network partition is less than or equal to half of the total number of nodes included in the cluster network, the method includes: When the sentinel cluster detects that the running state of the first master node at the first time point is a fault state, the first slave node of the third network partition in the cluster network is selected based on the sentinel cluster to determine the second master node; wherein the fault state is used to indicate that the communication between the first master node and the sentinel cluster is abnormal, the total number of nodes included in the third network partition is greater than or equal to half of the total number of nodes included in the cluster network, and the second network partition is the same as or different from the third network partition; Use the other first slave nodes in the cluster network as second slave nodes, and complete the master-slave switching of the second master node at a second time point; wherein the second time point is later than the first time point; When the first master node responds to the transaction request sent by the client at a third time point, the transaction request is executed based on a preset distributed lock; wherein the distributed lock is used to lock the data associated with the execution of the operation statement corresponding to the transaction request, and the third time point is later than the first time point; The distributed lock is stored in a preset first log cache, and the operation statement corresponding to the transaction request is stored in a preset second log cache; When the Sentinel cluster detects that the operating status of the first master node at the fourth time point is normal, the preset database of the second master node is updated based on the first log cache and the second log cache; wherein, the normal status is used to indicate that the communication between the first master node and the Sentinel cluster is normal, the fourth time point is later than the third time point, and the fourth time point is later than the second time point.
2. The method according to claim 1, characterized in that: The updating of data in a preset database of the second master node based on the first log cache and the second log cache includes: Release the distributed lock in the first log cache through the first master node, and send the second log cache to the second master node; Executing, by the second master node, an operation statement corresponding to the transaction request, to obtain the first transaction operation data at the third time point; The preset database is updated based on the first transaction operation data.
3. The method according to claim 2, characterized in that The method further comprises: Using the first master node as the second slave node, and sending a full synchronization signal to all the second slave nodes based on the updated preset database; Each of the second slave nodes responds to the full synchronization signal, and updates the data stored in each of the second slave nodes based on the preset database.
4. The method according to claim 1, characterized in that: When the sentinel cluster detects that the running state of the first master node at the first time point is a fault state, before performing node selection on the first slave node of the third network partition in the cluster network based on the sentinel cluster and determining the second master node, the method further includes: When the first sentinel node in the sentinel cluster detects that the running state of the first master node at the first time point is a fault state, sending a master node fault voting request to other sentinel nodes in the sentinel cluster; receiving, through the first sentinel node, a master node failure voting response returned by the other sentinel nodes, wherein the master node failure voting response is used to indicate a detection result of the other sentinel nodes on the running status of the first master node at the first time point; When the number of the master node fault voting responses indicating that the operating state of the first master node is the fault state is greater than a preset number, it is determined that the operating state of the first master node at the first time point is the fault state.
5. The method according to claim 1, characterized in that: The step of selecting a node from the first slave node of the third network partition in the cluster network based on the sentinel cluster to determine a second master node includes: Determine a sentinel master node from the sentinel nodes of the sentinel cluster based on a preset consensus algorithm; The second master node is determined by performing node selection on the first slave node of the third network partition in the cluster network based on the sentinel master node.
6. The method according to claim 5, characterized in that The step of performing node selection on the first slave node of the third network partition in the cluster network based on the sentinel master node to determine the second master node includes: Obtaining a node operation status and node attribute parameters of each of the first slave nodes in the third network partition; The first slave nodes in the third network partition are screened based on the node running status and the node attribute parameters to determine the second master node.
7. The method according to any one of claims 1 to 6, characterized in that: When the sentinel cluster detects that the running state of the first master node at the fourth time point is normal, before updating data in the preset database of the second master node based on the first log cache and the second log cache, the method further includes: Each of the sentinel nodes of the sentinel cluster sends a PING command to the first master node based on preset frequency data, and determines a plurality of response data returned by the master node according to the PING command within a preset time period; Detecting the operating status of the first master node according to the multiple response data; When the second sentinel node in the sentinel cluster detects that the running status of the first master node at the fourth time point is normal, it is determined that the running status of the first master node at the first time point is normal.
8. A data processing device based on a cluster network, characterized in that: Applied to a cluster network, the cluster network includes a first master node, at least one first slave node and a sentinel cluster, the first master node and a first client are located in a first network partition of the cluster network, the first master node is in communication connection with the first client, the sentinel cluster is located in a second network partition of the cluster network, the sentinel cluster is in communication connection with the first client, the first slave node is located in the first network partition or the second network partition, and the total number of nodes included in the first network partition is less than or equal to half of the total number of nodes included in the cluster network, the device includes: A master node determination module, used for, when the sentinel cluster detects that the running state of the first master node at a first time point is a fault state, performing node selection on the first slave node of the third network partition in the cluster network based on the sentinel cluster to determine a second master node; wherein the fault state is used to indicate that the communication between the first master node and the sentinel cluster is abnormal, the total number of nodes included in the third network partition is greater than or equal to half of the total number of nodes included in the cluster network, and the second network partition is the same as or different from the third network partition; A switching module, used to use the other first slave nodes in the cluster network as second slave nodes, and complete the master-slave switching of the second master node at a second time point; wherein the second time point is later than the first time point; a request execution module, configured to execute the transaction request based on a preset distributed lock when the first master node responds to the transaction request sent by the client at a third time point; wherein the distributed lock is used to lock data associated with the execution of the operation statement corresponding to the transaction request, and the third time point is later than the first time point; A storage module, used to store the distributed lock in a preset first log cache, and store the operation statement corresponding to the transaction request in a preset second log cache; An update module is used to update data in a preset database of the second master node based on the first log cache and the second log cache when the Sentinel cluster detects that the operating status of the first master node at a fourth time point is normal; wherein the normal status is used to indicate that the communication between the first master node and the Sentinel cluster is normal, the fourth time point is later than the third time point, and the fourth time point is later than the second time point.
9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Trading method and device of hotspot account based on Redis, medium and equipment
CN118365452A
Recovery and flush of endurant cache
US9058326B1
Cited By
Java background service deployment method and system based on double-node dynamic routing
CN120416016A
Method for determining master node of cluster, storage medium, electronic equipment and program product
CN120950343A