Database node management method and device, computer equipment and storage medium
By autonomously detecting and updating the database node topology, the database proxy can accurately perceive the master-slave relationship without relying on an external platform, solving the problem of request routing errors caused by changes in node relationships in existing technologies and reducing operation and maintenance costs.
Patent Information
- Application Number
- CN202410558139.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-07
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2044-05-07
AI Technical Summary
In existing technologies, database proxies cannot independently detect changes in the master-slave relationship of the underlying database nodes, leading to incorrect request routing. Furthermore, relying on external platforms or introducing registry center components increases operational costs.
The database agent obtains the configuration file, traverses the nodes in the topology cluster, generates the node topology structure, detects the target master node when the failover awareness function is enabled, updates the topology structure, updates the topology relationship based on the node status detection data, and autonomously perceives the master-slave relationship.
It can accurately perceive the topology of database nodes without relying on external platforms and additional components, reducing operation and maintenance costs and improving the availability and flexibility of database services.
Smart Images

Figure CN120910041A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of database management, and particularly relates to a database node management method and device, computer equipment and a storage medium. BACKGROUND
[0002] With the development of the Internet, the use of databases is becoming higher and higher. Now, a database proxy is usually used to route a database operation request initiated by a client to a corresponding database node to perform corresponding operations. However, the database proxy cannot perceive the master-slave relationship between different database nodes at the bottom layer. When the master-slave relationship of the bottom layer nodes changes, the request routing will be incorrect due to the inability to perceive the change. To avoid this phenomenon, the existing solution is to rely on an external platform or introduce a special registration center component to ensure that the database proxy knows the correct database node topology relationship. However, the first way belongs to a customized way, depends on the implementation of a specific platform, needs the database to be coupled with the external platform, and is not conducive to reuse. Moreover, the timeliness of the detection of this way often depends on the probe of the external platform, and the accuracy of the detection may be affected by factors such as hardware, clock, and network environment. The second way needs to maintain the introduced registration center component, which will increase the additional operation and maintenance cost. SUMMARY
[0003] The present application provides a database node management method and device, computer equipment and a storage medium to solve how to ensure that the database proxy knows the correct database node topology relationship without relying on an external platform and introducing additional components.
[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present application provides a database node management method and device, computer equipment and a storage medium.
[0005] In one aspect, the present application provides a database node management method, which comprises:
[0006] obtaining and loading a current configuration file to obtain a first configuration object, wherein the first configuration object comprises a topology cluster corresponding to each data shard, and each topology cluster contains a plurality of database nodes related to the corresponding data shard;
[0007] traversing each database node in a current topology cluster according to a data collection command to generate a first node topology structure corresponding to a current data shard, wherein the current topology cluster is a topology cluster corresponding to any one data shard in the first configuration object, and the first node topology structure contains the topology relationship between a plurality of database nodes related to the current data shard;
[0008] detect a target master node in the current topology cluster when a failover awareness function of a current database agent is enabled, and update the first node topology structure based on a detection result of the target master node to obtain a second node topology structure;
[0009] update the second node topology structure based on state detection data of each database node in the second node topology structure to obtain a target topology structure, wherein the target topology structure comprises topology relationships between each database node related to the current data shard after master node awareness and node state detection.
[0010] In another aspect, the present application provides a database node management apparatus, the apparatus comprising:
[0011] an obtaining module configured to obtain and load a current configuration file to obtain a first configuration object, wherein the first configuration object comprises topology clusters corresponding to each data shard, and each topology cluster comprises a plurality of database nodes related to the corresponding data shard;
[0012] a generating module configured to traverse each database node in a current topology cluster according to a data collection command to generate a first node topology structure corresponding to a current data shard, wherein the current topology cluster is a topology cluster corresponding to any data shard in the first configuration object, and the first node topology structure comprises topology relationships between a plurality of database nodes related to the current data shard;
[0013] a detecting module configured to detect a target master node in the current topology cluster when a failover awareness function of a current database agent is enabled, and update the first node topology structure based on a detection result of the target master node to obtain a second node topology structure;
[0014] an updating module configured to update the second node topology structure based on state detection data of each database node in the second node topology structure to obtain a target topology structure, wherein the target topology structure comprises topology relationships between each database node related to the current data shard after master node awareness and node state detection.
[0015] In another aspect, a computer device is provided, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory complete communication with each other through the communication bus;
[0016] the memory is configured to store a computer program;
[0017] the processor is configured to execute the program stored on the memory to implement the steps of the database node management method of any one of the embodiments of the first aspect.
[0018] In another aspect, a computer readable storage medium is provided, having stored thereon a computer program which, when executed by a processor, implements the steps of the database node management method according to any of the first aspect.
[0019] Compared with the prior art, the above technical solutions provided by the embodiments of the present application have the following advantages:
[0020] The method provided by the embodiments of the present application, the database agent can traverse the first configuration object to obtain the first node topology structure, and construct the first node topology structure corresponding to the current data shard according to the data collection command, and by starting the failover awareness function of the current database agent, the target master node in the current topology cluster can be automatically perceived, and the detection result of the target master node can be used to determine whether the first node topology structure has changed, if changed, the first node topology structure is updated according to the detection result of the target master node to obtain the second node topology structure, and the node state detection is performed on the second node topology structure, and the second node topology structure is updated based on the state detection data of each database node in the second node topology structure, so as to obtain the target topology structure with correct master-slave relationship and normal node state. Based on the above method, the coupling between the database agent and the external platform can be reduced without relying on the external platform, and additional components do not need to be added to ensure that the database agent knows the correct database node topology relationship, thereby saving the operation and maintenance cost brought by the additional components. BRIEF DESCRIPTION OF DRAWINGS
[0021] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and serve to explain the principles of the present application together with the specification.
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, those skilled in the art can obtain other drawings from these drawings without creative labor.
[0023] Figure 1 An application environment schematic diagram of a database node management method provided by the embodiments of the present application;
[0024] Figure 2 A flowchart of a database node management method provided by the embodiments of the present application;
[0025] Figure 3 An effect schematic diagram of the database agent acquiring the session lock provided by the embodiments of the present application;
[0026] Figure 4An effect diagram of a database agent acquiring a session lock according to an embodiment of the present application;
[0027] Figure 5 A flow diagram of a database node management method according to an embodiment of the present application;
[0028] Figure 6 A flow diagram of a database node management method according to an embodiment of the present application;
[0029] Figure 7 An effect diagram of a database agent scanning a target master node according to an embodiment of the present application;
[0030] Figure 8 A structure diagram of a database node management apparatus according to an embodiment of the present application;
[0031] Figure 9 A structure diagram of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION
[0032] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0033] Figure 1 An application environment diagram of a database node management method in an embodiment. Refer to Figure 1The database node management method is applied to a database node management system. The database node management system includes a client 110, a database agent 120, and a database server 130. The client 110, the database agent 120, and the database server 130 are connected through a network. The client 110 can be implemented by a desktop terminal or a mobile terminal, and the mobile terminal can be at least one of a mobile phone, a tablet computer, a notebook computer, and the like. The database agent 120 can be implemented by an independent server or a server cluster composed of multiple servers. The database server 130 is configured to store data resources. The database server includes multiple database nodes. A complete data resource is split into multiple data shards. Each data shard is distributed on a topology cluster. A topology cluster includes multiple database nodes of different node types. The node types include a master node, a backup node, a read-only node, and a combined node. The master node is a master node of the database and can bear read and write traffic. The backup node is a slave node of the master node and is a candidate node when the master node fails over or switches over. The read-only node is a slave node of the master node, can bear read traffic, and is an extended node for read-write separation but is not a candidate node when the master node fails over or switches over. The combined node (backup node + read-only node) can bear read traffic and is a candidate node when the master node fails over or switches over.
[0034] For a scenario in which a backup node and a read-only node are not distinguished, a node can be a backup node or a read-only node and bear read traffic. For a cluster mode of a public cloud RDS, there is one master node, one backup node, and multiple read-only instances. The backup node does not bear read traffic and is only a candidate node for failover. The backup node does not share traffic and thus does not need to maintain a connection pool and only needs to detect its state.
[0035] In one embodiment, Figure 2 FIG. 1 is a flowchart of a database node management method according to an embodiment of the present disclosure. The database node management method is applied to a database node management system. The database node management system includes a client 110, a database agent 120, and a database server 130. The client 110, the database agent 120, and the database server 130 are connected through a network. The client 110 can be implemented by a desktop terminal or a mobile terminal, and the mobile terminal can be at least one of a mobile phone, a tablet computer, a notebook computer, and the like. The database agent 120 can be implemented by an independent server or a server cluster composed of multiple servers. The database server 130 is configured to store data resources. The database server includes multiple database nodes. A complete data resource is split into multiple data shards. Each data shard is distributed on a topology cluster. A topology cluster includes multiple database nodes of different node types. The node types include a master node, a backup node, a read-only node, and a combined node. The master node is a master node of the database and can bear read and write traffic. The backup node is a slave node of the master node and is a candidate node when the master node fails over or switches over. The read-only node is a slave node of the master node, can bear read traffic, and is an extended node for read-write separation but is not a candidate node when the master node fails over or switches over. The combined node (backup node + read-only node) can bear read traffic and is a candidate node when the master node fails over or switches over. Figure 2 Figure 1 The database node management method specifically includes the following steps:
[0036] In step S2, a current configuration file is obtained and loaded to obtain a first configuration object. The first configuration object includes topology clusters corresponding to each data shard. Each topology cluster includes multiple database nodes related to the corresponding data shard.
[0037] Specifically, the current configuration file refers to the latest configuration file that can be obtained by the current database agent, and the update of the configuration file is triggered by an externally initiated command, and the current database agent obtains a first configuration object (Config) by loading the latest version of the configuration file. The first configuration object contains the global configuration and topology structure of all database nodes. After loading the configuration file, the Config object is modified by comparing the configuration files before and after loading to update the database node cluster topology boundary corresponding to each data shard. The database node cluster topology boundary refers to all database nodes related to the data shard and having a master-slave relationship, that is, all database nodes within the cluster topology boundary have a direct or indirect master-slave relationship. Different data shards correspond to different topology clusters, and a topology cluster is regarded as a Group. Different Groups use different threads to realize the perception update of the topology structure to prevent mutual blocking between Groups.
[0038] Step S4: According to the data collection command, each database node in the current topology cluster is traversed to generate a first node topology structure corresponding to the current data shard, wherein the current topology cluster is a topology cluster corresponding to any data shard in the first configuration object, and the first node topology structure contains the topology relationship between a plurality of database nodes related to the current data shard.
[0039] Specifically, the data collection command is denoted as show slave status. Each database node in the current topology cluster is traversed, that is, the current database agent establishes a communication connection with each database node in turn and obtains node information. The node information at least includes the node type and the master-slave relationship of the database node. According to the node type and the master-slave relationship, the first node topology structure can be constructed, and the node identifiers of different node types in the first node topology structure are added to the list of the same node type, that is, the node identifier of the node with the node type of the standby node is added to the candidate node list (candidate), the node identifier of the node with the node type of the read-only node is added to the read-only node list, and the master node is separately identified. Therefore, the node topology relationship indicated by the current configuration file can be reflected through the first node topology structure.
[0040] Step S6: When the failover perception function of the current database agent is started, the target master node in the current topology cluster is detected, and the first node topology structure is updated based on the detection result of the target master node to obtain a second node topology structure.
[0041] Specifically, the current database agent can be any database agent, the failover awareness function is a Failover awareness function, and the start state of the failover awareness function is used to determine whether the current database agent has a node master-slave relationship awareness function. In the case of starting the failover awareness function, the current database node will automatically perceive the target master node in the current topology cluster, the target master node is the Master node, and the current topology cluster is any topology cluster corresponding to a data shard. The master-slave mode corresponding to the topology cluster includes a single master mode, a master-backup mode (one master and multiple backups), a master-slave mode (one master and multiple slaves), and a one-master-multiple-backup-multiple-slave mode. That is, no matter which mode the topology cluster is in, there will be a Master node in the topology cluster, and other database nodes are backup nodes or read-only nodes belonging to the master node. Therefore, by detecting the target master node and comparing the target master node with the master node in the first node topology structure, it can be determined whether the master-slave relationship of the first node topology structure has changed, and then the first node topology structure can be updated according to the detection result of the target master node to obtain a second node topology structure after the topology structure is updated.
[0042] When the second topology structure is obtained by updating the first node topology structure, the master node, the read-only node list, and the candidate node list are also updated synchronously.
[0043] Step S8, updating the second node topology structure according to the state detection data of each database node in the second node topology structure to obtain a target topology structure, wherein the target topology structure contains the topology relationship between each database node related to the current data shard after the master node perception and node state detection.
[0044] Specifically, the state detection data of each database node in the second node topology structure is obtained, and the state detection data is used to indicate whether the state of the database node is normal. The second node topology structure is updated based on the state detection data of each database node in the second node topology structure, so as to obtain a target topology structure with correct node master-slave relationship and normal node state. Based on the above method, the coupling between the database agent and the external platform can be reduced without relying on the external platform, and additional components do not need to be added to ensure that the database agent knows the correct database node topology relationship, thereby saving the operation and maintenance cost brought by adding additional components.
[0045] In one embodiment, when the failover awareness function of the current database agent is started, the target master node in the current topology cluster is detected, and the first node topology structure is updated based on the detection result of the target master node to obtain a second node topology structure, comprising:
[0046] When a failover awareness function of a current database agent is enabled, a database agent configuration parameter is acquired;
[0047] When the database agent configuration parameter indicates multiple database agents and a timing election time of the current database agent is reached, a timing election task is executed to acquire session locks of each database node in the current topology cluster;
[0048] When a number of session locks acquired by the current database agent is greater than or equal to a preset number, the current database agent is determined as a first database agent;
[0049] A target master node in the current topology cluster is detected according to a node detection scheme corresponding to the first database agent, and a detection result of the target master node is obtained, wherein the detection result includes the detection result;
[0050] The first node topology structure is updated based on the detection result of the target master node to obtain a second node topology structure.
[0051] Specifically, the failover awareness function being enabled indicates that the current database agent can detect topology changes of data shards, and each database agent corresponds to a different timing election time, so that different database agents participate in election of the first database agent at different times, that is, the probability of simultaneously grabbing session locks is reduced. The first database agent is referred to as a Primary database agent, and the election method of the Primary database agent is to acquire session locks of each database node in the current topology cluster through a GET_LOCK() method. The session lock is used to indicate a network connection state between the current database agent and the database node, and a voting result of the database node for multiple database agents participating in election at the same time, that is, the database node can only send a session lock to one of multiple database agents participating in election at the same time, so that whether the current database agent is elected as the Primary database agent can be determined according to a number of session locks acquired by the current database agent. The preset number is greater than half of a number of all database nodes in the current topology cluster. If the number of session locks acquired by the current database agent is greater than or equal to the preset number, it indicates that the number of session locks acquired by the current database agent exceeds half of the number of nodes in the topology cluster, and the current database agent is taken as the Primary database agent. Only the Primary database agent is allowed to detect a target master node in the current topology cluster, so as to ensure that only one database agent performs an awareness operation at the same time, and also ensure that the awareness operation can be performed only when the database agent can access most of the MySQL database nodes, so as to prevent a split-brain phenomenon.
[0052] As Figure 3As shown, Proxy1 and Proxy2 indicate different database proxies, Master, Standby1, Standby2, ReadOnly1, and ReadOnly2 indicate different types of database nodes, Proxy1 obtains 3 session locks, and Proxy2 obtains 2 session locks, so Proxy1 obtains more than half of the number of session locks of the nodes and is elected as the Primary database proxy, and Proxy2 is the Secondary database proxy.
[0053] In one embodiment, after obtaining the database proxy configuration parameter, the method further comprises:
[0054] When the database proxy configuration parameter indicates a single database proxy, probing a target master node in the current topology cluster to obtain a probing result of the target master node;
[0055] Updating the first node topology structure based on the probing result of the target master node to obtain a second node topology structure.
[0056] Specifically, if the database proxy configuration parameter indicates a single database proxy deployment, the first database proxy election is not performed, and the target master node in the current topology cluster is directly probed to obtain a probing result of the target master node. Since the master node is no longer specified by an external platform, the database proxy needs to automatically perceive which node is the master node according to the topology relationship, rather than being explicitly specified by a configuration file.
[0057] In one embodiment, when the number of session locks obtained by the current database proxy is greater than or equal to a preset number, the current database proxy is determined to be the first database proxy, including:
[0058] When the number of session locks obtained by the current database proxy for all database nodes in the current topology cluster is greater than or equal to a preset number, the current database proxy is determined to be the first database proxy; or,
[0059] When the number of session locks obtained by the current database proxy for all online database nodes in the current topology cluster is greater than or equal to a preset number, the current database proxy is determined to be the first database proxy.
[0060] Specifically, the current database agent is elected as the first database agent only when it acquires more than half of the session locks of all database nodes under the current data shard. However, if there are many offline database nodes, there may be a scenario where the database agent is always without a master, which can easily lead to poor availability of the database service. The participation of all database nodes in the current topology cluster can ensure the consistency of the election results and prevent the split-brain phenomenon, which is the election of at least two database agents as the first database agent.
[0061] If the current database agent is elected as the first database agent only after acquiring more than half of the session locks of all online database nodes under the current data shard, it can avoid the scenario of a database agent being leaderless due to offline database nodes participating in the election, thus improving the availability of the database service, but it is very likely to cause a split-brain phenomenon.
[0062] The two methods described above are used to limit the types of database nodes participating in the election. The specific method chosen to elect the first database agent can be selected based on the actual application scenario.
[0063] In one embodiment, after acquiring the session locks of each database node within the current topology cluster, the method further includes:
[0064] When the number of session locks acquired by the current database agent is less than a preset number but not zero, and no first database agent has been elected, the session locks acquired by the current database agent are released, and a random sleep duration is randomly generated.
[0065] A timed election task is performed when the random sleep duration ends.
[0066] Specifically, such as Figure 4 As shown, if the session locks of all online database nodes are acquired to elect the first database agent, it may result in multiple database agents not acquiring more than half of the locks. In this case, a primary database agent cannot be elected. For database agents that have acquired session locks but do not have more than half of the session locks, the previously acquired locks are released, and a random sleep duration is generated so that multiple different database agents can sleep according to different random sleep durations. This staggers the execution of timed election tasks by different database agents and reduces the probability of multiple database agents competing for locks at the same time.
[0067] In one embodiment, after acquiring the session locks of each database node within the current topology cluster, the method further includes:
[0068] When the number of session locks acquired by the current database agent is zero and no first database agent has been elected, the agent shall hibernate for a preset hibernation duration.
[0069] The timing election task is executed when the preset sleep duration ends.
[0070] Specifically, in the case where the first database agent is not elected, no operation is performed on the database agent that does not obtain the session lock, and the next timing election task is executed after the preset sleep duration ends.
[0071] In one embodiment, the target master node in the current topology cluster is detected according to the node detection scheme corresponding to the first database agent, and a detection result of the target master node is obtained, including:
[0072] Each database node in the current topology cluster is traversed according to the node detection scheme corresponding to the first database agent, and the node identifier of an online database node that has no master-slave relationship and is a standby node is added to the relationship list;
[0073] When the database node indicated by the replication address of the database node that has a master-slave relationship and is a standby node is online, the first node identifier is added to the relationship list, where the first node identifier is the node identifier of the database node indicated by the replication address of the database node that has a master-slave relationship and is a standby node;
[0074] When the database node indicated by the replication address of the database node that has a master-slave relationship and is a read-only node is online, the second node identifier is added to the relationship list, where the second node identifier is the node identifier of the database node indicated by the replication address of the database node that has a master-slave relationship and is a read-only node;
[0075] The detection result of the target master node is determined according to the relationship list.
[0076] Specifically, the node detection scheme includes a data collection command, that is, the data collection command is executed to traverse all database nodes in the current topology cluster and collect node information, the node information including the online state, the node type, and the replication address of each database node, the replication address indicating the address of the database node to which the database node belongs. That is, the current database agent individually establishes a communication connection with each database node in the current topology cluster in turn, and the communication connection is released after being used, so that the communication connection in the connection pool is not occupied. If the network between the current database agent and the database node is not connected, the database node is directly ignored, and the next unconnected database node is selected for network connection and node information acquisition.
[0077] The current database agent adds a node key of a node without master-slave relationship and of a backup node type to a relationship list according to the collected node information, the relationship list is used to cache node keys of candidate nodes that can be target master nodes, the node without master-slave relationship means that the backup node has no associated master node and read-only node, and the backup node is an independent node and can be a target master node.
[0078] The node key of a database node with master-slave relationship and of a backup node type is added to the relationship list, that is, the backup node with master-slave relationship means that the backup node has an associated master node, and the replication address of the backup node refers to the address of the master node to which the backup node belongs, so that the node key of the master node can be found according to the replication address of the backup node, and the node key of the master node is added to the relationship list, that is, the master node is regarded as a candidate node that can be a target master node.
[0079] All read-only nodes in the current topology cluster are traversed, and the corresponding node key is found according to the replication address of the read-only node and added to the relationship list, that is, the replication address of the read-only node refers to the address of the master node to which the read-only node belongs, so that the node key of the master node can be found according to the replication address of the read-only node, and the node key of the master node is added to the relationship list.
[0080] It should be noted that each node key in the relationship list is saved only once, that is, when the replication address of the backup node and the replication address of the read-only node are the same, only the node key corresponding to one of the replication addresses is saved in the relationship list, that is, each node key in the relationship list is a de-duplicated result. The detection result of the target master node is determined according to each node key in the relationship list. The node key is denoted as host:port (IP address; port).
[0081] In one embodiment, the detection result of the target master node is determined according to the relationship list, including at least one of the following:
[0082] When the relationship list is empty, the detection result of the target master node is determined to be no target master node, and the current data shard is marked as unavailable;
[0083] When there is a node key with the largest number of slave nodes in the relationship list, the database node corresponding to the node key with the largest number of slave nodes is taken as the target master node, and a mutual exclusion lock is marked on the target master node;
[0084] When there are multiple node keys with the largest number of slave nodes in the relationship list, the detection result of the target master node is determined to be no target master node, and the current data shard is marked as unavailable.
[0085] Specifically, the relationship list not only contains the node identifier of the master node, but also records the number of slave nodes of each master node. If the relationship list is empty, it indicates that there is no target master node sensed by the current database agent, which means that the master-slave relationship of the current topology cluster is incorrect, the detection result of the target master node is no target master node, an error log is printed, the current data shard is marked as unavailable, and the current round of sensing detection is terminated in advance.
[0086] If there is a node identifier with the largest number of slave nodes in the relationship list, it indicates that the database node corresponding to the node identifier has the largest number of slave nodes, and the database node is determined as the target master node. The target master node is marked with a mutex, as shown in the following formula: Figure 5 The mutex is used to inform other database agents that the database node marked with the mutex is the target master node perceived by the database agent. In this way, the determination results of the target master node by multiple database agents are consistent. The detection result of the target master node includes not only the node identifier but also the master-slave relationship associated with the node identifier.
[0087] If there are multiple node identifiers with the largest number of slave nodes in the relationship list, it indicates that there are multiple database nodes with the same number of slave nodes, and the target master node cannot be selected according to the largest number of slave nodes. The detection result of the target master node is no target master node, an error log is printed, the current data shard is marked as unavailable, and the current round of sensing detection is terminated in advance.
[0088] In one embodiment, after the session lock of each database node in the current topology cluster is acquired, the method further includes:
[0089] When the number of session locks acquired by the current database agent is less than the preset number, and the first database agent is elected, the current database agent is determined as the second database agent.
[0090] The target master node marked with the mutex in the current topology cluster is scanned according to the node scanning scheme corresponding to the second database agent to obtain a scanning result of the target master node. The detection result includes the scanning result.
[0091] When the scanning result is scanning success, it is determined whether the scanned target master node is the same as the first master node in the first node topology structure.
[0092] When the target master node is not the same as the first master node, the first node topology structure is updated based on the scanning result of the target master node to obtain a second node topology structure.
[0093] The availability status of the current master node is obtained based on the second node topology, wherein the current master node includes the target master node;
[0094] When the current master node is in an available state, mark the current master node with a mutex lock.
[0095] Specifically, such as Figure 6 As shown, if the number of session locks acquired by the current database agent is less than the preset number, it means that the current database agent has failed to obtain election votes from most of the database nodes in the current topology cluster, or has failed to establish communication connections with most of the database nodes, and another database agent has been elected as the first database agent. In this case, the current database agent is regarded as the second database agent, which is the Secondary database agent. The Secondary database agent cannot initiate awareness operations on the current topology cluster. Only the Primary database agent has the right to initiate awareness operations on the current topology cluster.
[0096] Therefore, the second database agent periodically scans the target master nodes marked with mutex locks in the current topology cluster according to the node scanning scheme, obtaining the scan results of the target master nodes. The node scanning scheme includes timed scanning tasks and timed scanning cycles. The timed scanning cycle refers to the time interval between two executions of the timed scanning task, that is, the timed scanning task is executed periodically according to the timed scanning cycle to realize the timed scanning of target master nodes marked with mutex locks. Figure 7 As shown, Proxy1, acting as the first database proxy, marks the Master node, while Proxy2, acting as the second database proxy, only scans the Master node marked with a mutex lock.
[0097] If the scan result is successful, the target master node detected by the first database agent can be identified. Then, the scanned target master node is compared with the first master node in the first node topology to determine whether the topology of the current topology cluster has changed. If the target master node is different from the first master node, it means that the topology between the database nodes in the current topology cluster has changed. The first node topology is then updated according to the scan result of the target master node. If the scan result of the target master node is successful, it includes the master-slave relationship corresponding to the target master node. Therefore, the first node topology can be updated based on the master-slave relationship corresponding to the target master node to obtain the second node topology.
[0098] After the first node topology structure is updated, it is judged whether the current master node is available, the current master node being the master node in the second node topology structure, that is, the target master node, if the current master node is available, the current master node is marked with a mutual exclusion lock, that is, the target master node is continuously taken as the Master node, if the current master node is not available, the current data shard is marked as unavailable.
[0099] In one embodiment, the method further comprises:
[0100] When the scanning result is a scanning failure, or the target master node is the same as the first master node, the available state of a current master node is obtained based on the first node topology structure, wherein the current master node comprises the first master node;
[0101] When the available state of the current master node is available, the current master node is marked with a mutual exclusion lock;
[0102] When the available state of the current master node is not available, the current data shard is marked as unavailable.
[0103] Specifically, if the second database agent fails to scan the target master node, that is, the second database agent does not scan the Master node, it may be that the second database agent itself is faulty and cannot access the Master node, or it may be that the first database agent is faulty or the network is faulty, causing the mutual exclusion lock of the Master node to be marked as disconnected, that is, the mutual exclusion lock of the Master node is lost. Regardless of which case causes the second database agent to fail to scan the target master node, it is judged whether the current master node is available. Since the target master node is not scanned, the current master node is still the master node in the first node topology structure. If the current master node is available, the current master node is continuously marked with a mutual exclusion lock and is regarded as the Master node. If the current master node is not available, the current data shard is marked as unavailable.
[0104] The current database agent does not perform the timing election task until it reaches the timing election time corresponding to the current database agent, and then re-joins the election to strive to perform the sensing operation as the first database agent to detect the Master node.
[0105] In one embodiment, the first node topology structure is updated based on the detection result of the target master node to obtain a second node topology structure, comprising:
[0106] The detection result of the target master node is taken as a second configuration object;
[0107] When the second configuration object and the first configuration object fail to match, a second storage object is newly created;
[0108] In the second storage object, the first node topology structure is deleted or added with nodes according to the comparison result between the second configuration object corresponding node topology structure and the first node topology structure, to obtain a second node topology structure, wherein the second storage object is different from the first storage object where the first configuration object is located, and the reference relationship between the storage object of the first configuration object and the database node is replaced by the reference relationship between the second storage object and the database node through an atomic variable.
[0109] Specifically, the detection result of the target master node contains not only the node identifier of the detected target master node, but also the master-slave relationship of the target master node, and therefore the second configuration object also contains these contents. The second configuration object is compared with the first configuration object, so as to determine whether the topology cluster boundary and the master-slave relationship are changed. If the second configuration object fails to match the first configuration object, it indicates that the topology structure is changed, and a second storage object is newly created. The second storage object refers to a storage object that is different from the storage object storing the first node topology structure. The first node topology structure is updated in the second storage object, that is, the first node topology structure is not modified in the original storage object, but is modified in the second storage object. In this way, the database node in use still uses the first node topology structure before the update, and the second storage object uses the update thread to reference the second node topology structure after the update, so that the request received after the reference is updated uses the second node topology structure for request routing, that is, the reference is replaced by an atomic variable.
[0110] The second configuration object corresponding node topology structure is compared with the first node topology structure, that is, the master nodes in the first node topology structure are compared with the master nodes corresponding to the second configuration object, the read-only node list corresponding to the first node topology structure is compared with the read-only node list corresponding to the second configuration object, and the candidate node list corresponding to the first node topology structure is compared with the candidate node list corresponding to the second configuration object. Therefore, the comparison result contains a first sub-result, a second sub-result and a third sub-result. The first sub-result indicates the difference between the master nodes, the second sub-result indicates the difference between the read-only node lists, and the third sub-result indicates the difference between the candidate node lists. Therefore, based on the comparison result, the difference between the first node topology structure and the second configuration object corresponding node topology structure can be known. According to the comparison result, the first node topology structure is deleted or added with nodes, which is to refresh the topology structure.
[0111] For example, deleting a read-only node: closing the read-only node, releasing connection resources, disconnecting the database node from the client 110, and removing all read-only nodes that have not established a master-slave relationship with the new master node (i.e., the target master node) from the read-only node list. Deleting a standby node: removing all standby nodes that have not established a master-slave relationship with the new master node from the candidate node list. Adding a read-only node: initializing the read-only node (at this time, resources such as connection pools need to be created), and placing the new read-only node in the read-only node list. Adding a standby node: placing the new standby node in the candidate node list. Whether it is updating the configuration object or updating the read-only node list or the candidate node list, no modification is made to the original storage object, but after the second storage object, the reference is replaced by an atomic variable.
[0112] If the second configuration object matches the first configuration object successfully, it indicates that the topology structure has not changed, i.e., the master node, the standby node list, and the candidate node list have not changed, and no subsequent operation is required.
[0113] In one embodiment, the updating of the second node topology structure according to the state detection data of each database node in the second node topology structure to obtain a target topology structure comprises:
[0114] When the failover awareness function of the current database agent is turned on, the state detection data of all node types of database nodes in the second node topology structure is obtained.
[0115] The online state of the corresponding database node is updated according to the state detection data of all node types of database nodes in the second node topology structure to obtain a target topology structure.
[0116] Specifically, the database nodes are divided into master nodes, standby nodes, and read-only nodes. The state detection data of the master node includes a connection state parameter (max_failure_count), a read-only state parameter (read_only), and a master-slave relationship parameter. The state detection data of the read-only node includes a connection state parameter, a read-only state parameter, a master-slave relationship parameter, a thread state parameter, and a synchronization progress parameter. The state detection data of the standby node includes a read-only state parameter, a master-slave relationship parameter, a thread state parameter, and a synchronization progress parameter.
[0117] The connection state parameter is used to indicate the connection state between the current database agent and the database node. The connection state is detected every time a connection or online state check timing task is acquired. When the number of consecutive detection failures exceeds the threshold corresponding to the connection state parameter, the database node is marked as offline. If the read-only state parameter is false, the database node is marked as offline. If the master-slave relationship parameter of the master node indicates that it is a slave of other nodes (i.e., a slave node), the master node is marked as offline. If the master-slave relationship parameter of the read-only node indicates that it does not establish a master-slave relationship with the master node, the read-only node is marked as offline. If the master-slave relationship parameter of the standby node indicates that it does not establish a master-slave relationship with the master node, the standby node is marked as offline. If the thread state parameter indicates that the Slave IO thread or the Slave SQL thread is incorrect (i.e., the parameter is not yes), the database node is marked as offline. The synchronization progress parameter is recorded as Second_behind_master. When the synchronization progress parameter exceeds the corresponding threshold max_slave_lag parameter, the database node is marked as offline.
[0118] The state of the database node of different node types is updated by the state detection data. The state detection is performed on the standby node only when the failover awareness function of the current database agent is turned on. Otherwise, if there are more than two standby nodes in the current data shard, the current data shard is marked as unavailable. The state detection of the database node does not change the topology structure, but only updates the online state of the database node to offline or online.
[0119] In one embodiment, after generating the first node topology structure corresponding to the current data shard, the method further comprises:
[0120] When the failover awareness function of the current database agent is turned off, the state detection data of the database node of the node type of the master node or the read-only node in the first node topology structure is acquired.
[0121] The online state of the corresponding database node is updated according to the state detection data of the database node of the node type of the master node or the read-only node in the first node topology structure, and a target topology structure is obtained.
[0122] Specifically, as Figure 5 and Figure 6As shown, after obtaining the first node topology structure, it is found that the failover awareness function of the current database agent is closed, so that the awareness operation of the master node cannot be performed, and the second node topology structure cannot be obtained through the awareness operation. At this time, only the first node topology structure indicated by the current configuration file is available, and node state detection is performed on each database node in the first node topology structure. Since the failover awareness function of the current database agent is closed, only the database nodes of the master node type or the read-only node type are subjected to state detection, and there is no need to perform state detection on the standby nodes. Therefore, the state detection data of the nodes of the master node type or the read-only node type in the first node topology structure is obtained, and the online state of the corresponding master node or read-only node is updated according to the state detection data of the master node or read-only node, so as to obtain the target topology structure after the node online state is updated.
[0123] When the failover awareness function of the current database agent is closed, only node state detection can be implemented, and at this time, an external operation and maintenance platform needs to be used to perceive the change of the node topology structure. However, when the failover awareness function of the current database agent is opened, it can be independently deployed to perceive the change of the node topology structure. Therefore, the database agent in the present application can be independently deployed or can be deployed by relying on an external operation and maintenance platform, and has better usability and a wider applicable scenario.
[0124] In one embodiment, after obtaining and loading the current configuration file to obtain the first configuration object, the method further comprises:
[0125] When the configuration hot loading command or the cluster change command is received, the configuration hot loading command or the cluster change command is executed to obtain an updated configuration file;
[0126] The updated configuration file is loaded to obtain a third configuration object, and a third storage object is newly created for the third configuration object, wherein the third configuration object is different from the first configuration object;
[0127] In the third storage object, the first node topology structure is subjected to node deletion or node addition according to the comparison result between the node topology structure corresponding to the third configuration object and the first node topology structure, to obtain a third node topology structure, wherein the third node topology structure is used as the first node topology structure to perform the step of detecting the target master node in the current topology cluster when the failover awareness function of the current database agent is opened, and updating the first node topology structure based on the detection result of the target master node to obtain the second node topology structure.
[0128] Specifically, the configuration hot loading command is a command for requesting to change the current configuration file, and the cluster change command is a command for changing the topology cluster. Both commands change the topology structure corresponding to the topology cluster. Therefore, an updated configuration file is obtained after receiving and executing the configuration hot loading command or the cluster change command, so that a third configuration object different from the first configuration object is loaded. On the basis of the first node topology structure, the topology structure is refreshed according to the difference between the third configuration object and the first configuration object. The specific operation includes at least one of adding a node, deleting a node, and updating a node.
[0129] Adding a node is to add a node identifier that does not exist in the original topology but exists in the new topology. For example, adding a standby node is to add the node identifier of the standby node to the candidate node list, and adding a read-only node is to add the node identifier of the read-only node to the read-only node list. Deleting a node is to delete a node identifier that does not exist in the new topology but exists in the original topology. If a master node is deleted, it means that failover or switchover occurs, the master related client 110 connection is disconnected, master node detection is initiated, and the topology is updated. Deleting a standby node removes the node from the Candidate list. (Only when enable_auto_refresh_group=true, execute) Deleting a read-only node: close the read-only node (close the connection pool, disconnect the client 110 connection), and remove the node from the read-only node list. Updating a node: the node identifier exists in both the new and old topologies, but the node configuration has changed. This configuration is divided into two categories: role configuration and normal configuration. Changing the role configuration will trigger topology change. For example, changing a master node to a read-only node means that failover or switchover occurs, the master node detection is initiated, and the topology is updated. If a read-only node is changed to a standby node, it is equivalent to executing "delete read-only node + add standby node". If a standby node is changed to a read-only node, it is equivalent to executing "delete standby node + add read-only node".
[0130] Then, the third node topology structure is used as the first node topology structure to re-execute the master node detection operation and the node state detection operation in the case where the failover awareness function of the current database is enabled. That is, the perception operation and the state detection operation are cyclically executed again using the updated configuration object, so that the change of the topology structure between the nodes can be perceived in a timely manner according to the changed configuration file.
[0131] Whether updating the configuration object or updating the read-only node list or the candidate node list, no modification is made on the original storage object, but a new storage object is created, and then the reference is replaced by the atomic variable. That is, each time the configuration hot loading command or the cluster change command is executed, a new configuration object is generated; and each time the read-only node list and the candidate node list are changed, a new list is generated. Thus, it is ensured that the thread using the object holds the version before the object is updated, and the updating thread refreshes the new version to the reference, so as to ensure the reference data consistency and avoid unnecessary locking.
[0132] Figure 2 、 5 , 6 is a flowchart of a database node management method in an embodiment. It should be understood that although the steps in the flowchart of 6 are displayed in sequence according to the arrows, these steps are not necessarily executed in sequence according to the arrows. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, Figure 2 、 5 , the steps in the flowchart of 6 can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these sub-steps or stages is not necessarily sequential, but can be alternately executed with other steps or sub-steps or stages of other steps. Figure 2 、 5 , at least part of the steps in 6 can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these sub-steps or stages is not necessarily sequential, but can be alternately executed with other steps or sub-steps or stages of other steps.
[0133] In an embodiment, as shown in Figure 8 , a database node management apparatus is provided, comprising:
[0134] The acquisition module 310 is configured to acquire and load a current configuration file to obtain a first configuration object, wherein the first configuration object includes topology clusters corresponding to each data shard, and each topology cluster contains a plurality of database nodes related to the corresponding data shard.
[0135] The generation module 320 is configured to traverse each database node in a current topology cluster according to a data collection command to generate a first node topology structure corresponding to a current data shard, wherein the current topology cluster is a topology cluster corresponding to any data shard in the first configuration object, and the first node topology structure contains a topology relationship between a plurality of database nodes related to the current data shard.
[0136] The detection module 330 is configured to detect a target master node in the current topology cluster when the failover awareness function of the current database agent is started, and update the first node topology structure to obtain a second node topology structure based on a detection result of the target master node.
[0137] The updating module 340 is configured to update the second node topology structure according to state detection data of each database node in the second node topology structure, to obtain a target topology structure, wherein the target topology structure contains a topology relationship between each database node related to the current data shard after master node awareness and node state detection.
[0138] In an embodiment, the detection module 330 is further configured to:
[0139] When the failover awareness function of the current database agent is started, obtain a database agent configuration parameter;
[0140] When the database agent configuration parameter indicates multiple database agents, and a timing election time of the current database agent is reached, perform a timing election task to obtain a session lock of each database node in the current topology cluster;
[0141] When the number of session locks obtained by the current database agent is greater than or equal to a preset number, determine that the current database agent is a first database agent;
[0142] Detect a target master node in the current topology cluster according to a node detection scheme corresponding to the first database agent, to obtain a detection result of the target master node, wherein the detection result includes the detection result;
[0143] Update the first node topology structure based on the detection result of the target master node, to obtain a second node topology structure.
[0144] In an embodiment, the detection module 330 is further configured to:
[0145] When the database agent configuration parameter indicates a single database agent, detect a target master node in the current topology cluster, to obtain a detection result of the target master node;
[0146] Update the first node topology structure based on the detection result of the target master node, to obtain a second node topology structure.
[0147] In an embodiment, the detection module 330 is further configured to:
[0148] determining that the current database agent is a first database agent when a number of session locks acquired by the current database agent for all database nodes in the current topology cluster is greater than or equal to a preset number; or
[0149] determining that the current database agent is a first database agent when a number of session locks acquired by the current database agent for all online database nodes in the current topology cluster is greater than or equal to a preset number.
[0150] In one embodiment, the detection module 330 is further configured to:
[0151] traverse each database node in the current topology cluster according to a node detection scheme corresponding to the first database agent, and add a node identifier of an online database node without a master-slave relationship and of a standby node type to a relationship list;
[0152] add a first node identifier to the relationship list when a database node indicated by a replication address of a database node with a master-slave relationship and of a standby node type is online, where the first node identifier is a node identifier of a database node indicated by the replication address of the database node with the master-slave relationship and of the standby node type;
[0153] add a second node identifier to the relationship list when a database node indicated by a replication address of a database node with a master-slave relationship and of a read-only node type is online, where the second node identifier is a node identifier of a database node indicated by the replication address of the database node with the master-slave relationship and of the read-only node type;
[0154] determine a detection result of the target master node according to the relationship list.
[0155] In one embodiment, the detection module 330 is further configured to:
[0156] when the relationship list is empty, determine that the detection result of the target master node is no target master node, and mark the current data shard as unavailable;
[0157] when there is a node identifier with the largest number of slave nodes in the relationship list, take a database node corresponding to the node identifier with the largest number of slave nodes as the target master node, and mark the target master node with a mutual exclusion lock;
[0158] when there are multiple node identifiers with the largest number of slave nodes in the relationship list, determine that the detection result of the target master node is no target master node, and mark the current data shard as unavailable.
[0159] In one embodiment, the detection module 330 is further configured to:
[0160] when the number of session locks acquired by the current database agent is less than the preset number and the first database agent is elected, determining that the current database agent is a second database agent;
[0161] scanning the target master node marked with the mutual exclusion lock in the current topology cluster according to a node scanning scheme corresponding to the second database agent to obtain a scanning result of the target master node, wherein the detection result includes the scanning result;
[0162] when the scanning result is scanning success, determining whether the scanned target master node is the same as a first master node in the first node topology structure;
[0163] when the target master node is not the same as the first master node, updating the first node topology structure based on the scanning result of the target master node to obtain a second node topology structure;
[0164] obtaining an available state of a current master node based on the second node topology structure, wherein the current master node includes the target master node;
[0165] when the available state of the current master node is available, marking the current master node with the mutual exclusion lock.
[0166] In an embodiment, the detection module 330 is further configured to:
[0167] when the number of session locks acquired by the current database agent is less than the preset number but not zero and the first database agent is not elected, releasing the session locks acquired by the current database agent, and randomly generating a random sleep duration;
[0168] when the random sleep duration ends, performing a timing election task.
[0169] In an embodiment, the detection module 330 is further configured to:
[0170] when the number of session locks acquired by the current database agent is zero and the first database agent is not elected, sleeping for a preset sleep duration;
[0171] when the preset sleep duration ends, performing a timing election task.
[0172] In an embodiment, the detection module 330 is further configured to:
[0173] when the scanning result is scanning failure or the target master node is the same as the first master node, obtaining an available state of a current master node based on the first node topology structure, wherein the current master node includes the first master node.
[0174] when the availability status of the current master node is available, marking a mutual exclusion lock for the current master node;
[0175] when the availability status of the current master node is unavailable, marking that the current data shard is unavailable.
[0176] In an embodiment, the detection module 330 is further configured to:
[0177] taking the detection result of the target master node as a second configuration object;
[0178] when the second configuration object fails to match the first configuration object, creating a second storage object;
[0179] in the second storage object, performing node deletion or node addition on the first node topology structure according to a comparison result between a node topology structure corresponding to the second configuration object and the first node topology structure, to obtain a second node topology structure, wherein the second storage object is different from a first storage object in which the first configuration object is located, and a reference relationship between the storage object of the first configuration object and a database node is replaced by a reference relationship between the second storage object and the database node through an atomic variable.
[0180] In an embodiment, the update module 340 is further configured to:
[0181] when a failover awareness function of a current database agent is turned on, obtaining state detection data of database nodes of all node types in the second node topology structure;
[0182] updating an online state of a corresponding database node according to the state detection data of the database nodes of all node types in the second node topology structure, to obtain a target topology structure.
[0183] In an embodiment, the update module 340 is further configured to:
[0184] when the failover awareness function of the current database agent is turned off, obtaining state detection data of database nodes of node types of master nodes or read-only nodes in the first node topology structure;
[0185] updating an online state of a corresponding database node according to the state detection data of the database nodes of node types of master nodes or read-only nodes in the first node topology structure, to obtain a target topology structure.
[0186] In an embodiment, the generation module 320 is further configured to:
[0187] When the configuration hot loading command or the cluster change command is received, the configuration hot loading command or the cluster change command is executed to obtain an updated configuration file;
[0188] The updated configuration file is loaded to obtain a third configuration object, and a second storage object is newly created, wherein the third configuration object is different from the first configuration object;
[0189] In the third storage object, the first node topology structure is deleted or added according to a comparison result between a node topology structure corresponding to the third configuration object and the first node topology structure, to obtain a third node topology structure, wherein the third node topology structure is used as the first node topology structure to perform the step of detecting a target master node in the current topology cluster when the failover awareness function of the current database agent is started, and updating the first node topology structure based on a detection result of the target master node to obtain a second node topology structure.
[0190] As shown in Figure 9 The embodiments of the present application provide a computer device, which comprises a processor 711, a communication interface 712, a memory 713 and a communication bus 714, wherein the processor 711, the communication interface 712 and the memory 713 complete mutual communication through the communication bus 714;
[0191] The memory 713 is used for storing a computer program.
[0192] The processor 711 is used for executing the program stored in the memory 713, and realizes the database node management method provided by any one of the preceding method embodiments.
[0193] Those skilled in the art can understand that Figure 9 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0194] In one embodiment, the database node management apparatus provided by the present application can be realized in the form of a computer program, which can run on the computer device as shown in Figure 9 The memory of the computer device can store various program modules constituting the database node management apparatus, such as the obtaining module 310, the generating module 320, the detecting module 330 and the updating module 340 shown in Figure 8 The computer program constituted by various program modules makes the processor execute the database node management method of each embodiment of the present application described in the specification.
[0195] Figure 9 The computer device shown can perform the obtaining and loading of the current configuration file by the obtaining module 310 in the database node management apparatus as shown Figure 8 The computer device can perform the obtaining and loading of the current configuration file by the obtaining module 310 in the database node management apparatus as shown The computer device can perform the obtaining and loading of the current configuration file by the obtaining module 310 in the database node management apparatus as shown The computer device can perform the obtaining and loading of the current configuration file by the obtaining module 310 in the database node management apparatus as shown The computer device can perform the obtaining and loading of the current configuration file by the obtaining module 310 in the database node management apparatus as shown
[0196] The computer device can perform the obtaining and loading of the current configuration file by the obtaining module 310 in the database node management apparatus as shown
[0197] It should be noted that, in this document, the relationship terms such as "first" and "second" are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or sequence between these entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article or device including the element.
[0198] The foregoing detailed description of the application has been presented for purposes of illustration and description. Various modifications and changes can be made to these embodiments without departing from the spirit and scope of the application. It is intended that the scope of the application should not be limited by the particular representative embodiments described above.
Claims
1. A database node management method, characterized by, The method comprises: obtaining and loading a current configuration file to obtain a first configuration object, wherein the first configuration object comprises a plurality of topology clusters corresponding to respective data shards, and each topology cluster contains a plurality of database nodes related to the corresponding data shard; traversing each database node in a current topology cluster according to a data collection command to generate a first node topology structure corresponding to a current data shard, wherein the current topology cluster is a topology cluster corresponding to any data shard in the first configuration object, and the first node topology structure contains a topology relationship between a plurality of database nodes related to the current data shard; when a failover awareness function of a current database agent is turned on, detecting a target master node in the current topology cluster, and updating the first node topology structure based on a detection result of the target master node to obtain a second node topology structure; updating the second node topology structure according to state detection data of each database node in the second node topology structure to obtain a target topology structure, wherein the target topology structure contains a topology relationship between each database node related to the current data shard after master node awareness and node state detection.
2. The method of claim 1, wherein, The method further comprises: when the failover awareness function of the current database agent is turned on, obtaining a database agent configuration parameter; when the database agent configuration parameter indicates multiple database agents and a timing election time of the current database agent is reached, performing a timing election task to obtain a session lock of each database node in the current topology cluster; when a number of session locks obtained by the current database agent is greater than or equal to a preset number, determining that the current database agent is a first database agent; detecting a target master node in the current topology cluster according to a node detection scheme corresponding to the first database agent to obtain a detection result of the target master node, wherein the detection result comprises the detection result; updating the first node topology structure based on the detection result of the target master node to obtain a second node topology structure.
3. The method of claim 2, wherein, After obtaining the database agent configuration parameter, the method further comprises: when the database agent configuration parameter indicates a single database agent, detecting a target master node in the current topology cluster to obtain a detection result of the target master node; updating the first node topology structure based on the detection result of the target master node to obtain a second node topology structure.
4. The method of claim 2, wherein, When the number of session locks obtained by the current database agent is greater than or equal to the preset number, determining that the current database agent is a first database agent, comprises: when the number of session locks obtained by the current database agent for all database nodes in the current topology cluster is greater than or equal to the preset number, determining that the current database agent is a first database agent; or, The current database agent is determined as the first database agent when the number of session locks obtained by the current database agent for all online database nodes in the current topology cluster is greater than or equal to a preset number.
5. The method of claim 2, wherein, The target master node in the current topology cluster is detected according to the node detection scheme corresponding to the first database agent, and a detection result of the target master node is obtained. Each database node in the current topology cluster is traversed according to the node detection scheme corresponding to the first database agent, and a node identifier of an online database node without a master-slave relationship and of a standby node type is added to a relationship list. When a database node indicated by a replication address of a database node with a master-slave relationship and of a standby node type is online, a first node identifier is added to the relationship list, where the first node identifier is a node identifier of a database node indicated by the replication address of the database node with the master-slave relationship and of the standby node type. When a database node indicated by a replication address of a database node with a master-slave relationship and of a read-only node type is online, a second node identifier is added to the relationship list, where the second node identifier is a node identifier of a database node indicated by the replication address of the database node with the master-slave relationship and of the read-only node type. The detection result of the target master node is determined according to the relationship list.
6. The method of claim 5, wherein, The detection result of the target master node is determined according to the relationship list, including at least one of the following: When the relationship list is empty, it is determined that the detection result of the target master node is no target master node, and the current data shard is marked as unavailable; When there is a node identifier with the largest number of slave nodes in the relationship list, a database node corresponding to the node identifier with the largest number of slave nodes is taken as the target master node, and the target master node is marked with a mutual exclusion lock; When there are multiple node identifiers with the largest number of slave nodes in the relationship list, it is determined that the detection result of the target master node is no target master node, and the current data shard is marked as unavailable.
7. The method of claim 6, wherein, After obtaining the session locks of each database node in the current topology cluster, the method further includes: When the number of session locks obtained by the current database agent is less than the preset number and the first database agent is elected, the current database agent is determined as a second database agent; The target master node marked with the mutual exclusion lock in the current topology cluster is scanned according to the node scanning scheme corresponding to the second database agent, and a scanning result of the target master node is obtained, where the detection result includes the scanning result; When the scanning result is scanning success, it is determined whether the scanned target master node is the same as the first master node in the first node topology structure; When the target master node is not the same as the first master node, the first node topology structure is updated based on the scanning result of the target master node, and a second node topology structure is obtained; The availability state of a current master node is obtained based on the second node topology structure, where the current master node includes the target master node. When the availability state of the current master node is available, mark the current master node with a mutual exclusion lock.
8. The method of claim 7, wherein, After the session locks of the database nodes in the current topology cluster are acquired, the method further comprises: When the number of session locks acquired by the current database agent is less than the preset number and is not zero, and the first database agent is not elected, release the session locks acquired by the current database agent, and randomly generate a random sleep duration; When the random sleep duration ends, perform a timing election task.
9. The method of claim 7, wherein, After the session locks of the database nodes in the current topology cluster are acquired, the method further comprises: When the number of session locks acquired by the current database agent is zero, and the first database agent is not elected, sleep for a preset sleep duration; When the preset sleep duration ends, perform a timing election task.
10. The method of claim 7, wherein, The method further comprises: When the scanning result is a scanning failure, or the target master node is the same as the first master node, acquire the availability state of the current master node based on the first node topology structure, wherein the current master node includes the first master node; When the availability state of the current master node is available, mark the current master node with a mutual exclusion lock; When the availability state of the current master node is not available, mark the current data shard as unavailable.
11. The method of claim 2, wherein, Update the first node topology structure based on the detection result of the target master node to obtain a second node topology structure, comprising: Take the detection result of the target master node as a second configuration object; When the second configuration object does not match the first configuration object, create a second storage object; In the second storage object, perform node deletion or node addition on the first node topology structure according to the comparison result between the node topology structure corresponding to the second configuration object and the first node topology structure, to obtain a second node topology structure, wherein the second storage object is different from the first storage object where the first configuration object is located, and the reference relationship between the storage object of the first configuration object and the database node is replaced by the reference relationship between the second storage object and the database node through an atomic variable.
12. The method of claim 1, wherein, The second node topology structure is updated based on the state detection data of each database node in the second node topology structure to obtain a target topology structure, comprising: When the failover awareness function of the current database agent is turned on, acquire the state detection data of the database nodes of all node types in the second node topology structure; Update the online state of the corresponding database node based on the state detection data of the database nodes of all node types in the second node topology structure to obtain a target topology structure.
13. The method of claim 1, wherein, After the first node topology structure corresponding to the current data shard is generated, the method further comprises: When the failover awareness function of the current database agent is turned off, acquire the state detection data of the database nodes of the node types of master nodes or read-only nodes in the first node topology structure; According to state detection data of database nodes of which node types are master nodes or read-only nodes in the first node topology structure, an online state of a corresponding database node is updated to obtain a target topology structure.
14. The method of claim 1, wherein, After the current configuration file is acquired and loaded to obtain the first configuration object, the method further comprises: Upon receiving a configuration hot loading command or a cluster change command, the configuration hot loading command or the cluster change command is executed to obtain an updated configuration file; The updated configuration file is loaded to obtain a third configuration object, and a third storage object is newly created, wherein the third configuration object is different from the first configuration object; In the third storage object, the first node topology structure is deleted or added with nodes according to a comparison result between a node topology structure corresponding to the third configuration object and the first node topology structure, to obtain a third node topology structure, wherein the third node topology structure is used as the first node topology structure to execute the step of detecting a target master node in the current topology cluster when the failover awareness function of the current database agent is enabled, and updating the first node topology structure based on a detection result of the target master node to obtain a second node topology structure.
15. A database node management apparatus, characterized by comprising: The apparatus comprises: an acquisition module configured to acquire and load a current configuration file to obtain a first configuration object, wherein the first configuration object comprises topology clusters corresponding to respective data shards, and each topology cluster contains a plurality of database nodes related to the corresponding data shard; a generation module configured to traverse each database node in a current topology cluster according to a data collection command to generate a first node topology structure corresponding to a current data shard, wherein the current topology cluster is a topology cluster corresponding to any data shard in the first configuration object, and the first node topology structure contains a topology relationship between a plurality of database nodes related to the current data shard; a detection module configured to detect a target master node in the current topology cluster when a failover awareness function of a current database agent is enabled, and update the first node topology structure based on a detection result of the target master node to obtain a second node topology structure; an update module configured to update the second node topology structure according to state detection data of each database node in the second node topology structure to obtain a target topology structure, wherein the target topology structure contains a topology relationship between each database node related to the current data shard after master node awareness and node state detection.
16. A computer device, comprising: The apparatus comprises a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; The memory is configured to store a computer program; The processor is configured to execute the program stored in the memory to implement the database node management method in any one of claims 1-14.
17. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the database node management method in any one of claims 1-14.
Citation Information
Patent Citations
System topological structure maintenance method and device, computer equipment and storage medium
CN113497737A
Node management method and device for database cluster and electronic equipment
CN115878361A
Data security management method and device, equipment and storage medium
CN116361073A
Database management method and system based on container platform
CN116561096A
Network topology sensing method and device, electronic equipment and storage medium
CN117768331A