A database management method, a database management apparatus, and a computing device cluster

CN122838499APending Publication Date: 2026-09-29HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510372498.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

然而,这种集中式管理存在一个潜在风险,一旦集中式管理平台中的某个节点发生故障,集中式管理平台即失去作用

Benefits of technology

[0040]上述提供的任一种装置、计算设备集群、计算机存储介质、或者计算机程序产品,均用于执行上文所提供的方法,因此,其所能达到的有益效果可参考上文提供的对应方法中的对应方案的有益效果,此处不再赘述。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838499A_ABST
    Figure CN122838499A_ABST
Patent Text Reader

Abstract

Provided are a database management method and device based on distributed technology and a computing device cluster. The method is applied to a first management platform and includes: sending probe information to a second database, obtaining a first probe result of the second database, sending the first probe result to a first storage system, and determining whether the second database is faulty according to the fact that the first storage system stores all probe results related to the second database; and in the case where the second database is faulty, switching the first database to a primary database. The first management platform, a second management platform for managing the second database, and the first storage system are deployed on an infrastructure, and the first database is a backup database of the second database.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a database management method, a database management device, and a computing device cluster. Background Technology

[0002] In database management scenarios, to ensure that database failures do not affect normal user operation, a primary and a standby database are typically configured. This configuration allows for rapid failover to the standby database in the event of a primary database outage or unexpected interruption, thus ensuring uninterrupted business operations. Currently, centralized management is commonly used for primary and standby databases. However, this centralized management has a potential risk: if any node in the centralized management platform fails, the platform becomes ineffective. In this situation, the high availability capability of the primary and standby databases is lost, failover becomes impossible, and consequently, normal business operations are affected. Summary of the Invention

[0003] This application provides a database cluster management method, management platform, and computing device cluster based on distributed technology, which can improve the availability of the database cluster.

[0004] Firstly, this application provides a database management method based on distributed technology. This method is applied to a first management platform, which runs on an infrastructure. The first management platform manages a first database, which serves as a backup database for a second database. A second management platform is also deployed on the infrastructure, and this second management platform manages the second database.

[0005] The method includes: sending probe information to a second database to obtain a first probe result of the second database; sending the first probe result to a first storage system, wherein the first storage system is deployed on infrastructure and stores all probe results related to the second database, including the first probe result and the second probe result, the second probe result being obtained by a second management platform sending probe information to the second database; determining whether the second database is faulty based on all probe results; and switching the first database to the primary database in the event of a fault in the second database.

[0006] In the above scheme, the database cluster is configured with a distributed management platform, which includes multiple management platforms. Specifically, each database in the primary database and at least one backup database is configured with one management platform. This avoids the problem of the database cluster being unable to perform primary / backup failover if a single management platform fails. With distributed management, a first storage system is added as a communication medium between the various management platforms. The management platform determines whether the primary database is faulty based on all probe results related to the primary database stored in the first storage system. This avoids a situation where one management platform mistakenly believes the primary database is faulty and resolves the split-brain problem between the management platforms.

[0007] In one possible implementation, the first storage system includes multiple storage nodes, each of which stores all the detection results. A first database is located in the first node, and a second database is located in the second node. The infrastructure includes multiple nodes, including the first node and the second node.

[0008] In one possible implementation, a first management platform is located in a first node, a second management platform is located in a second node, the first node and the second node are located in different availability zones or different regions of the infrastructure, and multiple storage nodes are located in at least two availability zones or at least two regions of the infrastructure.

[0009] In one possible implementation, a third management platform is also deployed on the infrastructure. The third management platform is used to manage a third database, which is a backup database for the second database. All detection results also include the third detection results, which are obtained by the third management platform sending detection information to the second database.

[0010] In one possible implementation, the first storage system further stores indicator data of a first database and indicator data of a third database. The method further includes: determining to switch the first database to the primary database based on the indicator data of the first database and the indicator data of the third database stored in the first storage system.

[0011] In the above scheme, the management platform of each backup database determines which backup database needs to be upgraded to the primary database based on the indicator data of each backup database. This can avoid the split-brain problem that occurs when multiple backup database management platforms perform the operation of upgrading the primary database.

[0012] In one possible implementation, the first storage system also stores the value of a failover key, which includes a first value and a second value, wherein the first value indicates that the failover state is not in progress and the second value indicates that the failover state is in progress.

[0013] Before switching the first database to the primary database, the method further includes: locking the value of the fault switching key stored in the first storage system if the value of the fault switching key is a first value and the value of the fault switching key is not locked, and modifying the value of the fault switching key stored in the first storage system to a second value.

[0014] After switching the first database to the primary database, the method further includes: modifying the value of the failover key stored in the first storage system to the first value, and releasing the lock on the value of the failover key.

[0015] In the above solution, the first storage system stores the failover key and adds a locking mechanism, which can further prevent the split-brain problem from occurring on the management platform of multiple backup databases.

[0016] In one possible implementation, the first storage system also stores the value of a normal switching key, which includes a third value, a fourth value, and a fifth value. The third value indicates that the system is not in a normal switching state, the fourth value indicates that the system is in a normal switching state during the downgrade phase, and the fifth value indicates that the system is in a normal switching state during the master-slave phase.

[0017] The method further includes: if the value of the normal switching key stored in the first storage system is a third value, modifying the value of the normal switching key stored in the first storage system to a fourth value; the second management platform is used to switch the second database to a backup database if the value of the normal switching key stored in the first storage system is a fourth value, and to modify the value of the normal switching key stored in the first storage system to a fifth value.

[0018] The method further includes: if the value of the normal switching key stored in the first storage system is the fifth value, switching the first database to the primary database and modifying the value of the normal switching key stored in the first storage system to the third value.

[0019] In the above scheme, each management platform achieves normal switching between the primary and backup databases through the value of the normal switching key on the first storage system.

[0020] In one possible implementation, the infrastructure also includes a second storage system for providing the Internet Protocol (IP) address of the primary database to the client. Switching the primary database to the primary database also includes: modifying the IP address of the primary database stored in the second storage system to the IP address of the primary database, so that the client can establish a connection with the primary database through the IP address of the primary database.

[0021] In the above solution, deploying a secondary storage system within the infrastructure allows clients to obtain the primary database's IP address through this system. For example, when the primary and backup databases are deployed across availability zones or regions, clients can promptly obtain the primary database's IP address from the secondary storage system.

[0022] Secondly, this application also provides a database management device based on distributed technology. This database management device is applied to a first management platform, which runs on an infrastructure. The first management platform manages a first database, which serves as a backup database for a second database. A second management platform is also deployed on the infrastructure, and this second management platform manages the second database.

[0023] The database management device includes a detection module and a switching module.

[0024] The detection module is used to send detection information to the second database, obtain the first detection result of the second database, and send the first detection result to the first storage system. The first storage system is deployed on the infrastructure and stores all detection results related to the second database. All detection results include the first detection result and the second detection result. The second detection result is obtained by the second management platform sending detection information to the second database.

[0025] The switching module is used to determine whether the second database is faulty based on all detection results, and, if the second database is faulty, to switch the first database to the main database.

[0026] In one possible implementation, the first storage system includes multiple storage nodes, each of which stores all the detection results. A first database is located in the first node, and a second database is located in the second node. The infrastructure includes multiple nodes, including the first node and the second node.

[0027] In one possible implementation, a first management platform is located in a first node, a second management platform is located in a second node, the first node and the second node are located in different availability zones or different regions of the infrastructure, and multiple storage nodes are located in at least two availability zones or at least two regions of the infrastructure.

[0028] In one possible implementation, a third management platform is also deployed on the infrastructure. The third management platform is used to manage a third database, which is a backup database for the second database. All detection results also include the third detection results, which are obtained by the third management platform sending detection information to the second database.

[0029] In one possible implementation, the first storage system also stores indicator data of the first database and indicator data of the third database, and the switching module is further configured to: determine whether to switch the first database to the main database based on the indicator data of the first database and the indicator data of the third database stored in the first storage system.

[0030] In one possible implementation, the first storage system also stores the value of a failover key, which includes a first value and a second value, wherein the first value indicates that the failover state is not in progress and the second value indicates that the failover state is in progress.

[0031] Before switching the first database to the primary database, the switching module is also used to: lock the value of the fault switching key stored in the first storage system if the value of the fault switching key is the first value and the value of the fault switching key is not locked, and modify the value of the fault switching key stored in the first storage system to the second value.

[0032] After switching the first database to the primary database, the switching module is also used to: modify the value of the failover key stored in the first storage system to the first value, and release the lock on the value of the failover key.

[0033] In one possible implementation, the first storage system also stores the value of a normal switching key, which includes a third value, a fourth value, and a fifth value. The third value indicates that the system is not in a normal switching state, the fourth value indicates that the system is in a normal switching state during the downgrade phase, and the fifth value indicates that the system is in a normal switching state during the master-slave phase.

[0034] The switching module is also used to: change the value of the normal switching key stored in the first storage system to a fourth value when the value of the normal switching key stored in the first storage system is a third value; the second management platform is used to switch the second database to the backup database when the value of the normal switching key stored in the first storage system is a fourth value, and to change the value of the normal switching key stored in the first storage system to a fifth value.

[0035] The switching module is also used to: switch the first database to the primary database when the value of the normal switching key stored in the first storage system is the fifth value, and modify the value of the normal switching key stored in the first storage system to the third value.

[0036] In one possible implementation, the infrastructure also includes a second storage system for providing the IP address of the main database to the client. The switching module is further configured to: modify the IP address of the main database stored in the second storage system to the IP address of the first database, so that the client can establish a connection with the first database through the IP address of the first database.

[0037] Thirdly, this application also provides a computing device cluster. The cluster may include at least one computing device, each computing device including a processor and a memory. The processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, causing the computing device cluster to perform the method provided by the first aspect or any possible implementation thereof.

[0038] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium includes computer program instructions that, when executed by a cluster of computing devices, cause the cluster of computing devices to perform the method provided by the first aspect or any possible implementation thereof.

[0039] Fifthly, this application also provides a computer program product. The computer program product includes computer program instructions. When the computer program instructions are executed by a cluster of computing devices, the cluster of computing devices performs the method provided by the first aspect or any possible implementation thereof.

[0040] Any of the devices, computing device clusters, computer storage media, or computer program products provided above are used to execute the methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects of the corresponding solutions in the corresponding methods provided above, and will not be repeated here. Attached Figure Description

[0041] Figure 1 This is a schematic diagram of the structure of a database system provided in an embodiment of this application;

[0042] Figure 2 This is a schematic diagram of the public and private keys stored in the first storage system provided in this application embodiment;

[0043] Figure 3 This is a schematic diagram of the structure of a database management platform provided in an embodiment of this application;

[0044] Figure 4 This is a schematic diagram of another database system structure provided in an embodiment of this application;

[0045] Figure 5 This is a flowchart of a database management method provided in an embodiment of this application;

[0046] Figure 6 The embodiments provided in this application are based on Figure 5 The diagram illustrates the database management method shown.

[0047] Figure 7a and Figure 7bThis is a schematic diagram illustrating a method for switching a backup database to a primary database, as provided in an embodiment of this application.

[0048] Figure 7c This is a schematic diagram illustrating the locking and modification of the value of the fault switching key of the first storage system, as provided in an embodiment of this application.

[0049] Figure 7d This is a schematic diagram illustrating the determination of database-based indicator data to upgrade to the main database, provided in an embodiment of this application.

[0050] Figure 8 This is a flowchart of another database management method provided in an embodiment of this application;

[0051] Figure 9 This is a schematic diagram of the structure of a database management device provided in an embodiment of this application;

[0052] Figure 10 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application;

[0053] Figure 11 and Figure 12 This is a schematic diagram of the structure of a computing device cluster provided in an embodiment of this application. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be described below with reference to the accompanying drawings.

[0055] In the description of the embodiments of this application, the words "exemplary," "for example," or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplary," "for example," or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the words "exemplary," "for example," or "for instance" is intended to present the relevant concepts in a specific manner.

[0056] In the description of the embodiments in this application, the term "and / or" is merely a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, B existing alone, and A and B existing simultaneously. Furthermore, unless otherwise stated, the term "multiple" means two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals.

[0057] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and their variations all mean "including but not limited to," unless otherwise specifically emphasized.

[0058] Before introducing the embodiments of this application, the technical terms mentioned in the embodiments of this application will be explained below.

[0059] A region refers to the deployment location of a data center cluster. Data center clusters within a region share a range of public services, such as elastic computing, block storage, object storage, virtual private cloud (VPC) networks, elastic public networks, and image services.

[0060] An availability zone (AZ) is a collection of one or more data centers within a region. Each AZ has its own independent power, cooling, and water supply. Logically, within an AZ, computing, network, and storage resources are further divided into multiple sub-clusters. Multiple AZs within a region are connected via high-speed fiber optic cables to meet users' needs for building high-availability systems across AZs.

[0061] The primary database is the main database that handles all user read and write operations. It is updated in real-time and contains the latest data. Under normal circumstances, all applications and users connect to the primary database to read and write data. The primary database is responsible for handling all transactions, ensuring data consistency and integrity.

[0062] A standby database is a copy of the primary database and typically does not directly handle application read / write requests. The standby database receives the latest data from the primary database through a data synchronization mechanism to maintain data synchronization with the primary database. In the event of a failure of the primary database, the standby database can quickly switch over to become the primary database and take over the service, thereby ensuring business continuity. Primary and standby databases are key components in database high availability and disaster recovery strategies. They are typically used to ensure that data is not lost and that services can be restored as quickly as possible in the event of a database system failure.

[0063] In database management scenarios, to ensure that database failures do not affect normal user operation, a primary and a backup database are typically configured. This configuration allows for rapid failover to the backup database in the event of a primary database outage or unexpected interruption, thus ensuring uninterrupted business operations. Currently, centralized management is commonly used for primary and backup databases. However, this centralized management has a potential risk: if the management platform itself fails, the high availability capabilities of the primary and backup databases will be lost, making failover impossible and impacting normal business operations.

[0064] Therefore, this application provides a database management method that can solve the above problems.

[0065] The database management method provided in this application embodiment is applied to a first management platform, which runs on an infrastructure and manages a first database, which serves as a backup database for a second database. A second management platform is also deployed on the infrastructure and manages the second database. The method includes: sending probe information to the second database to obtain a first probe result; sending the first probe result to a first storage system, wherein the first storage system is deployed on the infrastructure and stores all probe results related to the second database, including the first probe result and a second probe result, the second probe result being obtained by the second management platform sending probe information to the second database; determining whether the second database is faulty based on all probe results; and switching the first database to the primary database if the second database fails. In this scheme, the database cluster is configured with a distributed management platform, which includes multiple management platforms, i.e., each database in the primary database and at least one backup database is configured with a management platform. This avoids the problem of the database cluster being unable to perform primary / backup switching if a single management platform fails. In a distributed management system, a first storage system is added as a communication medium between various management platforms. The management platform decides whether the main database is faulty based on all the probe results related to the main database stored in the first storage system, which can prevent a management platform from mistakenly believing that the main database is faulty.

[0066] Figure 1 This is a schematic diagram of the structure of a database system provided in an embodiment of this application. For example... Figure 1 As shown, a database system may include a cloud management platform and infrastructure. The cloud management platform is used to manage the infrastructure. The infrastructure may include one or more regions, and each region may include one or more availability zones (AZs). Each AZ may include one or more compute nodes and / or one or more storage nodes.

[0067] The infrastructure may specifically include one or more database clusters, a primary storage system, and a secondary storage system. The database clusters, primary storage system, and secondary storage system are described below.

[0068] exist Figure 1 In a database cluster, the database cluster is used to provide database (DB) services to users. To improve the high availability of the database cluster, it can be deployed in a single Availability Zone (AZ), across AZs, or across regions.

[0069] Specifically, a database cluster can include multiple databases forming a primary / standby protection network, and a management platform for each database. These databases and their management platforms can be deployed in a single Availability Zone (AZ), multiple AZs within a single region, or multiple regions. Figure 1 For example, a database cluster may include a first database (DB1), a second database (DB2), and a third database (DB3). DB1 to DB3 form a primary-backup protection system, with DB2 as the primary database and DB1 and DB3 as backup databases for DB2. The first database is configured with a first management platform (management platform 1), the second database is configured with a second management platform (management platform 2), and the third database is configured with a third management platform (management platform 3). DB1 and its management platform 1 are deployed in AZ1, DB2 and its management platform 2 are deployed in AZ2, and DB3 and its management platform 3 are deployed in AZ3. AZ1, AZ2, and AZ3 may belong to the same region or different regions. Of course, in some embodiments, DB1 to DB3 and their management platforms may also be deployed in a single AZ.

[0070] After introducing the structure and deployment of the database cluster, the first storage system will be introduced below.

[0071] exist Figure 1 In this system, the first storage system provides management information to the management platforms of each database in the database cluster. Specifically, the management information stored in the first storage system can include key-value pairs of the database cluster. The management platforms of each database in the cluster can perform master-slave failover based on these keys, thereby improving the high availability of the database cluster. The database cluster keys can include public keys and private keys for multiple databases. Public keys are those that allow read and write operations on the management platforms of each database in the cluster. Private keys are those that allow read and write operations on the management platform of that specific database, but only read operations on other databases.

[0072] The following is combined with Figure 2 The format of the common key is introduced.

[0073] The format of a public key can include, for example: Figure 2 The diagram shows the identifier of the database cluster, its common identifier, the IP addresses of the multiple databases included in the cluster, and the identifier of the common management key. Figure 1 Taking DB2 to DB3 as an example, the IP addresses of multiple databases in the common key are the IP addresses of DB2 to DB3. DB2 to DB3 can perform read and write operations on the common key. The common management key can include the IP address of the primary database, the detection results of the primary database by the management platform of each database in the database cluster, the failover status, the normal failover status, the normal failover status message, the primary database repair status, and the forced rebuild identifier.

[0074] The management platforms for each database can use F or S to represent the detection results of the main database. F indicates that the main database is abnormal, and S indicates that the main database is normal.

[0075] The failover key can contain a first value and a second value. The first value indicates that failover is not in progress, and the second value indicates that failover is in progress. For example, the first and second values ​​can be 0 and 1, respectively. Furthermore, the failover state is locked before the standby database is promoted to the primary database, and the lock is released after the promotion. This locking mechanism prevents a split-brain event from occurring among multiple standby databases of the primary database. Specifically, it avoids the situation where multiple standby databases are promoted to the primary database when the primary database fails and needs to be demoted to a standby database.

[0076] The value of the normal handover key can include a third, fourth, and fifth value. The third value indicates that the system is not in a normal handover state, the fourth value indicates that the system is in a normal handover state during the downgrade phase, and the fifth value indicates that the system is in a normal handover state during the master-slave phase. For example, the third, fourth, and fifth values ​​are 00, 01, and 10, respectively.

[0077] The forced rebuild flag indicates a backup delay between the standby and primary databases. After a primary database failure, during recovery, data can be backed up from the standby database based on the forced rebuild flag.

[0078] The following is combined with Figure 2 This section introduces the format of private keys in a database.

[0079] Private keys can include, for example Figure 2 The private keys 1 through 3 are shown. The format of each private key may include... Figure 2 The database cluster identifier, private identifier, database IP address, and private management key identifier are shown.

[0080] The IP address of the database in private key 1 is the IP address of DB1. DB1 can perform read and write operations on private key 1, while DB2 and DB3 can perform read operations on private key 1.

[0081] The IP address of the database in private key 2 is the same as the IP address of DB2. DB2 can perform read and write operations on private key 2, while DB1 and DB3 can perform read operations on private key 2.

[0082] The IP address of the database in private key 3 is the IP address of DB3. DB3 can perform read and write operations on private key 3, while DB2 and DB1 can perform read operations on private key 3.

[0083] A private management key is used to record the metrics data of the database identified by the database IP address included in the private key. These metrics data may include, but are not limited to: availability status, database status, data synchronization time, database restart time, database heartbeat detection time, and the network status of the current database to other databases.

[0084] The primary storage system can adopt a distributed architecture. Specifically, the primary storage system may include multiple primary storage nodes, which can be deployed within a single Availability Zone (AZ), across AZs, or across regions. In some embodiments, a disaster recovery scheme can be deployed among the multiple primary storage nodes to improve the reliability of the primary storage system. For example, reliability can be improved by configuring primary and backup storage nodes among the multiple primary storage nodes. Figure 1 The AZ1 shown is an example. For instance, multiple first storage nodes can be deployed in... Figure 1 As shown in AZ1 to AZ3. For example, multiple first storage nodes can be deployed in addition to... Figure 1 In the other AZs besides AZ1 to AZ3 shown.

[0085] After introducing the first storage system, the second storage system will now be introduced.

[0086] exist Figure 1 In this system, the second storage system is used to provide the IP address of the primary database to the clients of the database cluster.

[0087] Specifically, the client can obtain the IP address of the primary database from the secondary storage system, and then establish a connection with the primary database based on that IP address, thereby accessing the primary database through this connection. After switching the primary / standby relationship between the two databases, the management platform of the standby database needs to modify the IP address of the primary database stored in the secondary storage system.

[0088] The secondary storage system can adopt a distributed architecture. Specifically, the secondary storage system may include multiple secondary storage nodes, which can be deployed in a single Availability Zone (AZ), across AZs, or across regions. In some embodiments, disaster recovery schemes can be deployed among the multiple secondary storage nodes to improve the reliability of the secondary storage system. For example, reliability can be improved by setting up primary and backup storage nodes among the multiple secondary storage nodes. For example, the multiple secondary storage nodes can be deployed in... Figure 1 The AZ1 shown is an example. For instance, multiple second storage nodes can be deployed in... Figure 1 As shown in AZ1 to AZ3. For example, multiple second storage nodes can be deployed in addition to... Figure 1 In the other AZs besides AZ1 to AZ3 shown.

[0089] exist Figure 1 In this system, the database cluster manages the respective databases through their respective management platforms, thereby achieving high availability. These database management platforms can have the same functionality and / or structure. The following section will combine... Figure 3 The database management platform provided in the embodiments of this application will be introduced.

[0090] Figure 3 This application provides a schematic diagram of the database management platform structure. (See attached diagram.) Figure 3 As shown, the management platform can specifically include a tool layer, an access layer, and a service layer.

[0091] The tool layer provides tools that can perform various management functions. Specifically, the tool layer may include... Figure 3 The illustration shows a first type of tool for managing the database, a second type of tool for managing the service layer, and a third type of tool for managing the first and second storage systems. It should be noted that this application does not impose specific limitations on the names or number of tools included in the first, second, and third types of tools. In practical applications, different tools can be configured according to the required functionalities.

[0092] The first type of tool is used to implement database operation and maintenance functions. These functions may include, but are not limited to, starting the database, stopping the database, and querying the database's identity information. The database's identity information includes whether it is a primary database or a backup database.

[0093] The second type of tool is used to implement the operation and maintenance functions of the service layer. The operation and maintenance functions of the service layer may include, but are not limited to, starting the service layer and stopping the service layer.

[0094] The third type of tool is used to implement operation and maintenance functions for the first and second storage systems. These functions may include, but are not limited to, starting and stopping the first and second storage systems, and providing operation and maintenance personnel with information stored in the first and second storage systems.

[0095] The access layer provides the operational interface for atomic capabilities, assisting the tool layer and service layer in implementing corresponding functions. Specifically, the access layer may include... Figure 4 The diagram shows a first type of operation interface for the database and a second type of operation interface for the first and second storage systems. The first type of operation interface supports calls to the first type of tools and service layer, while the second type of operation interface supports calls to the service layer. It should be noted that this application does not impose specific limitations on the first type of operation interface or the names and number of interfaces contained within it. In practical applications, different interfaces can be configured according to the supported operations.

[0096] The first type of tools and service layer can call the first type of operation interface to perform database operations. The first type of tools' database operations include: starting the database, stopping the database, and querying identity information. The service layer's database operations include: promoting a backup database to a primary database and demoting a primary database to a backup database.

[0097] The service layer can call the second type of operation interface to perform operations on the first and second storage systems. Operations performed by the service layer on the first storage system include: reading, writing, and locking operations on public and private keys, and reading and writing operations on the IP address of the main database stored in the second storage system.

[0098] The service layer is used to provide high availability services for the database cluster. Specifically, the service layer can include collection services, failover services, and normal failover services.

[0099] The collection service is used to collect metric data from this database. The collection service is also used to perform write operations on the first storage system and the second storage system. Write operations performed by the collection service on the first storage system may include, but are not limited to, writing metric data from this database into the first storage system.

[0100] The failover service is used to probe the primary database and determine the probe results from the management platform. It also switches the primary / standby relationship between the primary and standby databases in the event of a primary database failure. Specific functions of the failover service are detailed later in the documentation. Figure 5 The details of each step are omitted here.

[0101] The normal failover service is used to switch the primary / standby relationship between the primary and standby databases when the primary database is functioning normally. Specific functions of the normal failover service can be found in the following sections. Figure 8 The details of each step are omitted here.

[0102] The service layer may also include synchronization services and monitoring services.

[0103] The synchronization service is used to synchronize the public and private keys stored in the first storage system to this database. This synchronization service allows for failover in case the first storage system fails and becomes inoperable. In its implementation, the synchronization service can use asynchronous, semi-synchronous, or fully synchronous methods to synchronize data.

[0104] The monitoring service is used to monitor the status of this database and repair database anomalies. Database anomalies can include one or more of the following: database replication anomalies, database process anomalies, primary database read / write status anomalies, primary database VIP loss, and database identity (primary or backup) anomalies.

[0105] The monitoring service can also be used to monitor anomalies in the first storage system and / or the second storage system. Anomalies in the first storage system may include abnormal key values ​​(including null or incorrect values) and key value cleanup. The monitoring service can also be used to monitor one or more of the following: the status of the first storage system, the status of the second storage system, the status of nodes in the first storage system, and the status of nodes in the second storage system. Furthermore, the monitoring service can send an anomaly notification to the cloud management platform when one or more of the following indicate an anomaly: the status of the first storage system, the status of the second storage system, the status of nodes in the first storage system, and the status of nodes in the second storage system. This anomaly notification indicates that one or more of the following: the first storage system, the second storage system, the nodes in the first storage system, and the nodes in the second storage system are experiencing an anomaly.

[0106] In some embodiments, the database cluster can also include, for example: Figure 4 The structure shown. In Figure 4 In the illustrated architecture, the database cluster does not include a second storage system, and the database cluster can operate in Virtual Internet Protocol (VIP) mode. In VIP mode, clients can access the primary database through the VIP address of the database cluster. When switching the primary / standby relationship between the two databases, the standby database's management platform can configure the standby database based on the VIP address, thereby attaching the VIP address to the standby database. In this way, clients can still access the new primary database through the VIP address.

[0107] Next, based on the above description, a database management method provided by an embodiment of this application will be described. This method is used to determine whether the primary database in a database cluster has failed, and to switch a backup database to the primary database in the event of a failure.

[0108] Figure 5 This is a flowchart of a database management method provided in an embodiment of this application. Figure 5 The method shown can be applied to Figure 1 or Figure 4 The management platform for any database in the database cluster shown. For example... Figure 5 As shown, the method may include S501 to S504. The following example uses a management platform 1 applied to DB1, combined with... Figure 6 The steps performed by the DB1 management platform 1 shown are as follows: Figure 5 The steps shown will be explained. It should be noted that... Figure 6 Therefore Figure 1 and Figure 4 The example shown uses DB2 as the primary database, and DB1 and DB3 as backup databases for DB2.

[0109] S501, send probe information to the second database and obtain the first probe result of the second database.

[0110] With DB2 as the primary database and DB1 and DB3 as backup databases for DB2, Figure 6 Taking T1 as an example, management platform 1 performs status probing on DB2 and obtains the first probing result. In specific implementation, management platform 1 can periodically send probing information to DB2 using heartbeat detection technology. The probing information can include heartbeat requests. After receiving a heartbeat request, DB2 replies with a heartbeat response to management platform 1. If no heartbeat response is received within a certain period of time, management platform 1 can determine that DB2 is abnormal and record the first probing result indicating an abnormality. If a heartbeat response is received within a certain period of time, management platform 1 can determine that DB2 is normal and record the first probing result indicating normal operation.

[0111] S502, send the first detection result to the first storage system.

[0112] After receiving the first detection result, management platform 1, as follows Figure 6 As shown in T2, the first probe results can be written to the first storage system so that other management platforms can obtain them. Thus, when the primary database includes multiple backup databases, each database's management platform can obtain the probe results from the management platforms of other databases through the first storage system. In other words, the first storage system stores all probe results related to DB2.

[0113] When DB2's backup database includes DB1 and DB3, the total probe results include the first probe result, the second probe result, and the third probe result. The second probe result is obtained by management platform 2 sending probe information to DB2. The third probe result is obtained by management platform 3 sending probe information to DB2. The process by which management platform 2 obtains the second probe result, and the process by which management platform 3 obtains the third probe result, can be referred to in S501 above, and will not be repeated here. It is understood that when DB2's backup database includes DB1, the total probe results include the first probe result and the second probe result.

[0114] S503: Obtain all detection results from the first storage system and determine whether the second database has failed based on all detection results.

[0115] like Figure 6 As shown in T3, management platform 1 can obtain all probe results from the first storage system to determine whether DB2 has failed. If the backup databases for DB2 include DB1 and DB3, and all probe results from management platforms 1 through 3 indicate anomalies, management platform 1 can determine that DB2 has failed. If any probe result indicates normal operation, management platform 1 can also determine that DB2 has failed. If any probe result indicates anomalies, management platform 1 can determine that DB2 has not failed.

[0116] Furthermore, if the detection results obtained by management platform 1 are missing from the results of any one or two of management platforms 1 to 3, management platform 1 may wait for a preset time. If, within this preset time, management platform 1 obtains complete detection results from the first storage system, management platform 1 will continue to perform arbitration. If, within this preset time, complete detection results are still not obtained, management platform 1 may consider the management platform with missing detection results as having abandoned arbitration. In the case of being deemed to have abandoned arbitration, management platform 1 can determine whether DB2 has failed based on the existing detection results.

[0117] For example, if all probe results lack the third probe result from management platform 3, and management platform 3 fails to write the third probe result to the first storage system within the preset time, management platform 1 will be unable to obtain the third probe result within that preset time. Management platform 1 can then consider management platform 3 as having waived arbitration. Subsequently, management platform 1 determines whether DB2 has failed based on the probe results from management platform 1 and management platform 2. Specifically, if the probe results from management platform 1 and management platform 2 indicate anomalies, management platform 1 can determine that DB2 has failed; otherwise, management platform 1 determines that DB2 has not failed.

[0118] For example, if all detection results lack those from management platforms 2 and 3, and if management platforms 2 and 3 fail to write their detection results to the first storage system within the preset time, management platform 1 will be unable to obtain their detection results within that time. Management platform 1 can then consider management platforms 2 and 3 as having waived arbitration. Subsequently, management platform 1 determines whether DB2 has failed based on the first detection result. Specifically, if the first detection result indicates an anomaly, management platform 1 determines that DB2 has failed; otherwise, management platform 1 determines that DB2 has not failed.

[0119] S504: If it is determined that the second database has failed, switch the first database to become the primary database.

[0120] If it is determined that DB2 has failed, such as Figure 6 As shown in T4, management platform 1 can switch DB1 to the primary database.

[0121] The management platform 1 can switch DB1 to the primary database by changing the IP address of the primary database stored in the first storage system to the IP address of DB1.

[0122] For example, in Figure 1 In the scenario shown, the second storage system stores the IP address of the primary database, which is the IP address of DB2, such as... Figure 7a As shown, management platform 1 can first change the IP address of the primary database stored in the first storage system to the IP address of DB1, and then change the IP address of the primary database stored in the second storage system to the IP address of DB1. For example... Figure 7a As shown, the IP address of the main database obtained by the client from the second storage system is the IP address of DB1. Subsequently, the client can establish a connection with DB1 based on the IP address of DB1 and access DB1.

[0123] For example, in Figure 4 In the scenario shown, such as Figure 7b As shown, management platform 1 can first modify the IP address of the primary database stored in the first storage system to the IP address of DB1, and then configure DB1 according to the VIP address, mounting the VIP address onto DB1. In this way, clients can access DB1 using the VIP address. When the primary and backup databases are deployed across availability zones or regions, if different availability zones or regions are typically deployed in independent physical network environments, each availability zone or region uses an independent subnet, and IP address conflicts may exist in different subnets, the database cluster cannot use the Virtual Internet Protocol (VIP) mode. In this case, the database cluster can configure a second storage system and switch the primary database as described in the previous paragraph.

[0124] Before switching DB1 to the primary database, management platform 1 can obtain the failover key value from the first storage system. If the failover key value is the first value and is not locked, management platform 1 locks the failover key value and modifies the first value to the second value. After switching DB1 to the primary database, management platform 1 can release the lock on the failover key value in the first storage system and then modify it back to the first value. By locking the failover key value and modifying it to the second value, management platform 1 prevents other management platforms from performing the primary database upgrade operation, thus preventing a split-brain problem between multiple backup database management platforms. For example... Figure 7c As shown, before switching DB1 to the primary database, if the failover key stored in the first storage system has a first value and is not locked, then the failover key value stored in the first storage system is locked, and the first value is changed to the second value. Then, after switching DB1 to the primary database, the lock on the failover key value is released, and the failover key value is changed back to the first value.

[0125] Furthermore, before switching DB1 to the primary database, management platform 1 can obtain indicator data for each database from the first storage system and determine the target DB to be switched to the primary database based on this data. After determining the target DB, if the target DB is DB1, management platform 1 executes S504 to switch DB1 to the primary database; otherwise, it terminates the primary database upgrade operation. If the target DB is DB1, management platform 3 for DB3 can terminate the primary database upgrade operation. For example, as... Figure 7d As shown, management platform 1 can obtain indicator data for DB1 and DB3 from the first storage system, and determine the target DB based on the indicator data for DB1 and DB3. Taking data synchronization time as an example, management platform 1 can determine the target DB based on the data synchronization time of DB1 and DB3. For example, if DB1 has the most recent data synchronization time, DB1 can be used as the target DB. If DB3 has the most recent data synchronization time, DB3 can be used as the target DB. For example, as... Figure 7d As shown, if DB1 is the target DB, the management platform 1 executes step S504 to switch DB1 to the primary database.

[0126] It should be noted that the above refers to... Figure 5 The steps described are based on management platform 1 as an example. For management platform 2 of DB2 and management platform 3 of DB3, management platform 1 and management platform 3 can perform the above steps. Figure 5The steps are shown below. Management platform 1 and management platform 3 execute the above steps. Figure 5 The specific process of the steps shown can be referred to the above description of the execution of management platform 1, and will not be repeated in this embodiment.

[0127] In combination Figure 5 After introducing the failover process for database clusters, the following section will combine... Figure 8 This section describes the normal process of switching between master and standby relationships in a database cluster.

[0128] Figure 8 This is a flowchart of another database management method provided in the embodiments of this application. Figure 8 The method shown can be applied to Figure 1 The database cluster shown or Figure 4 The database cluster shown. (As shown) Figure 8 As shown, the method may include S801 to S804. The following uses... Figure 1 and Figure 4 Taking DB2 as the primary database and DB1 as the backup database for DB2 as an example, Figure 8 The steps shown are explained below.

[0129] S801, Management Platform 1 obtains the value of the normal switching key from the first storage system and modifies the value of the normal switching key stored in the first storage system from the third value to the fourth value.

[0130] In this embodiment, management platform 1 can execute S801 upon receiving a normal switchover instruction from the cloud management platform. Specifically, the cloud management platform can send a normal switchover instruction to management platform 1 as needed to adjust the primary / standby relationship between DB1 and DB2. In other words, the normal switchover instruction indicates that DB1 will be promoted to primary database and DB2 will be demoted to standby database.

[0131] In this embodiment, the value of the normal switch key is the third value when there is no need to switch the primary / standby relationship, indicating that it is not in a normal switch state. When the management platform 1 determines that a primary / standby relationship switch is required, it can modify the third value to a fourth value. Thus, the management platform 2 of DB2 can detect the change in the value of the normal switch key and switch DB2 to the standby database. The third and fourth values ​​are, for example, 00 and 01 as described above.

[0132] S802, Management Platform 2 obtains the value of the normal switchover key from the first storage system. If the value of the normal switchover key is the fourth value, it switches DB2 as the backup database, and then modifies the value of the normal switchover key to the fifth value. The fifth value can be, for example, 10 as mentioned above.

[0133] When management platform 2 detects that the normal switch key value is the fourth value, it switches DB2 to the backup database. Then, management platform 2 changes the normal switch key value to the fifth value to notify management platform 1 that DB1 can be promoted to the primary database.

[0134] S803, Management Platform 1 obtains the value of the normal switching key from the first storage system. If the value of the normal switching key is the fifth value, DB1 is switched to become the primary database.

[0135] The process of switching DB1 to the primary database in management platform 1 can be referred to the above. Figure 5 The description of S504 in the text will not be repeated here.

[0136] S804, Management Platform 1 modifies the value of the normal switching key to the third value.

[0137] After the switch is completed, management platform 1 can change the value of the normal switch key to the third value, thereby restoring it to the default value.

[0138] It should be noted that the above refers to... Figure 8 The descriptions of each step are based on the DB1 management platform 1 as an example. For the DB3 management platform 3, the execution process of the management platform 3 can refer to the execution process of DB1 described above, and will not be repeated in this embodiment.

[0139] based on Figure 5 and Figure 8 In addition to the method shown, this application embodiment also provides a database management platform.

[0140] Figure 9 This is a schematic diagram of the structure of a database management device 900 based on distributed technology provided in an embodiment of this application. The database management device 900 can be used to execute... Figure 5 and / or Figure 8 This step in the method shown is to achieve the following: Figure 1 or Figure 4 The database cluster management shown. Figure 9 As shown, the database management device 900 may include a detection module 901 and a switching module 902.

[0141] The detection module 901 is used to send detection information to the second database, obtain the first detection result of the second database, and send the first detection result to the first storage system. The first storage system is deployed on the infrastructure and stores all detection results related to the second database. All detection results include the first detection result and the second detection result. The second detection result is obtained by the second management platform sending detection information to the second database.

[0142] The switching module 902 is used to determine whether the second database is faulty based on all detection results, and, if the second database is faulty, to switch the first database to the main database.

[0143] In one possible implementation, the first storage system includes multiple storage nodes, each of which stores all the detection results. A first database is located in the first node, and a second database is located in the second node. The infrastructure includes multiple nodes, including the first node and the second node.

[0144] In one possible implementation, a first management platform is located in a first node, a second management platform is located in a second node, the first node and the second node are located in different availability zones or different regions of the infrastructure, and multiple storage nodes are located in at least two availability zones or at least two regions of the infrastructure.

[0145] In one possible implementation, a third management platform is also deployed on the infrastructure. The third management platform is used to manage a third database, which is a backup database for the second database. All detection results also include the third detection results, which are obtained by the third management platform sending detection information to the second database.

[0146] In one possible implementation, the first storage system also stores indicator data of the first database and indicator data of the third database. The switching module 902 is further configured to: determine whether to switch the first database to the main database based on the indicator data of the first database and the indicator data of the third database stored in the first storage system.

[0147] In one possible implementation, the first storage system also stores the value of a failover key, which includes a first value and a second value, wherein the first value indicates that the failover state is not in progress and the second value indicates that the failover state is in progress.

[0148] Before switching the first database to the primary database, the switching module 902 is also used to: lock the value of the fault switching key stored in the first storage system when the value of the fault switching key is a first value and the value of the fault switching key is not locked, and modify the value of the fault switching key stored in the first storage system to a second value.

[0149] After switching the first database to the primary database, the switching module 902 is also used to: modify the value of the fault switching key stored in the first storage system to the first value, and release the lock on the value of the fault switching key.

[0150] In one possible implementation, the first storage system also stores the value of a normal switching key, which includes a third value, a fourth value, and a fifth value. The third value indicates that the system is not in a normal switching state, the fourth value indicates that the system is in a normal switching state during the downgrade phase, and the fifth value indicates that the system is in a normal switching state during the master-slave phase.

[0151] The switching module 902 is also used to: modify the value of the normal switching key stored in the first storage system to a fourth value when the value of the normal switching key stored in the first storage system is a third value; the second management platform is used to switch the second database to a backup database when the value of the normal switching key stored in the first storage system is a fourth value, and to modify the value of the normal switching key stored in the first storage system to a fifth value.

[0152] The switching module 902 is also used to: switch the first database to the main database when the value of the normal switching key stored in the first storage system is the fifth value, and modify the value of the normal switching key stored in the first storage system to the third value.

[0153] In one possible implementation, the infrastructure also includes a second storage system for providing the IP address of the main database to the client. The switching module 902 is further configured to: modify the IP address of the main database stored in the second storage system to the IP address of the first database, so that the client can establish a connection with the first database through the IP address of the first database.

[0154] Both the detection module 901 and the switching module 902 can be implemented in software or in hardware. For example, the implementation of the detection module 901 will be described below. Similarly, the implementation of the switching module 902 can be referenced from the implementation of the detection module 901.

[0155] As an example of a software functional unit, the detection module 901 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, the detection module 901 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same Availability Zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.

[0156] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.

[0157] As an example of a hardware functional unit, the detection module 901 may include at least one computing device, such as a server. Alternatively, the detection module 901 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.

[0158] The detection module 901 includes multiple computing devices that can be distributed in the same region or in different regions. Similarly, the detection module 901 includes multiple computing devices that can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the detection module 901 includes multiple computing devices that can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0159] It should be noted that, in other embodiments, the detection module 901 can be used to perform... Figure 5 and / or Figure 8 In any step of the method shown, the switching module 902 can be used to execute Figure 5 and / or Figure 8 Any step in the method shown. The steps implemented by the detection module 901 and the switching module 902 can be specified as needed and implemented by the detection module 901 and the switching module 902 respectively. Figure 5 and / or Figure 8 The different steps in the method shown implement all the functions of the database management device 900.

[0160] This application also provides a computing device 1000. For example... Figure 10 As shown, the computing device 1000 includes a bus 1001, a processor 1002, a memory 1003, and a communication interface 1004. The processor 1002, the memory 1003, and the communication interface 1004 communicate with each other via the bus 1001. The computing device 1000 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 100.

[0161] Bus 1001 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 10 The bus 104 may be represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 104 may include a path for transmitting information between various components of the computing device 100 (e.g., memory 1003, processor 1002, communication interface 1004).

[0162] The processor 1002 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0163] The memory 1003 may include volatile memory, such as random access memory (RAM). The processor 1002 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0164] The memory 1003 stores executable program code, and the processor 1002 executes the executable program code to implement the functions of the aforementioned detection module 901 and switching module 902, thereby achieving... Figure 5 and / or Figure 8 Method. That is, the memory 1003 stores the method for execution. Figure 5and / or Figure 8 The instructions for the method.

[0165] Alternatively, the memory 1003 stores executable code, and the processor 1002 executes the executable code to implement the functions of the aforementioned database management device 900, thereby achieving... Figure 5 and / or Figure 8 Method. That is, the memory 1003 stores the method for execution. Figure 5 and / or Figure 8 The instructions for the method.

[0166] The communication interface 1004 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 1000 and other devices or communication networks.

[0167] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0168] like Figure 11 As shown, the computing device cluster includes at least one computing device 1000. The memory 1003 of one or more computing devices 1000 in the computing device cluster may store the same memory for executing... Figure 5 and / or Figure 8 The instructions for the method.

[0169] In some possible implementations, the memory 1003 of one or more computing devices 1000 in the computing device cluster may also store memory for execution. Figure 5 and / or Figure 8 Part of the instructions of the method. In other words, a combination of one or more computing devices 1000 can jointly execute instructions for performing... Figure 5 and / or Figure 8 The instructions for the method.

[0170] It should be noted that the memory 1003 in different computing devices 1000 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the database management device 900. That is, the instructions stored in the memory 1003 of different computing devices 1000 can implement the functions of one or more modules in the detection module 901 and the switching module 902.

[0171] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 12One possible implementation is shown. For example... Figure 12 As shown, two computing devices 1000A and 1000B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this possible implementation, the memory 1003 in computing device 1000A stores instructions for executing the function of the detection module 901. Simultaneously, the memory 1003 in computing device 1000B stores instructions for executing the function of the switching module 902.

[0172] Figure 12 The connection method between the computing device clusters shown can be based on the database management method provided in this application, therefore, the function implemented by the switching module is handed over to the computing device 1000B for execution.

[0173] It should be understood that Figure 12 The functions of computing device 1000A shown can also be performed by multiple computing devices 1000. Similarly, the functions of computing device 1000B can also be performed by multiple computing devices 1000.

[0174] This application also provides another computing device cluster. The connection relationships between the computing devices in this computing device cluster can be similarly referred to... Figure 11 and Figure 12 The connection method of the computing device cluster. The difference is that the memory 1003 of one or more computing devices 1000 in this computing device cluster can store the same data for execution. Figure 5 and / or Figure 8 The instructions for the method.

[0175] In some possible implementations, the memory 1003 of one or more computing devices 1000 in the computing device cluster may also store memory for execution. Figure 5 and / or Figure 8 Part of the instructions of the method. In other words, a combination of one or more computing devices 1000 can jointly execute instructions for performing... Figure 5 and / or Figure 8 The instructions for the method.

[0176] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform... Figure 5 and / or Figure 8 method.

[0177] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute... Figure 5 and / or Figure 8 method.

[0178] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A database management method based on distributed technology, characterized in that, The method is applied to a first management platform, which runs on an infrastructure. The first management platform manages a first database, which serves as a backup database for a second database. A second management platform is also deployed on the infrastructure, and this second platform manages the second database. The method includes: Send probe information to the second database to obtain the first probe result of the second database; The first detection result is sent to the first storage system, wherein the first storage system is deployed on the infrastructure and stores all detection results related to the second database. The all detection results include the first detection result and the second detection result, and the second detection result is obtained by the second management platform sending detection information to the second database. Determine whether the second database is faulty based on all the detection results; In the event of a failure in the second database, the first database will be switched to become the primary database.

2. The method according to claim 1, characterized in that, The first storage system includes multiple storage nodes, each of which stores all the detection results. The first database is located in the first node, and the second database is located in the second node. The infrastructure includes multiple nodes, including the first node and the second node.

3. The method according to claim 2, characterized in that, The first management platform is located in the first node, the second management platform is located in the second node, the first node and the second node are located in different availability zones or different regions of the infrastructure, and the plurality of storage nodes are located in at least two availability zones or at least two regions of the infrastructure.

4. The method according to claim 2 or 3, characterized in that, A third management platform is also deployed on the infrastructure. The third management platform is used to manage a third database, which is a backup database for the second database. All detection results also include third detection results, which are obtained by the third management platform sending detection information to the second database.

5. The method according to claim 4, characterized in that, The first storage system also stores indicator data from the first database and indicator data from the third database. The method further includes: Based on the indicator data of the first database stored in the first storage system and the indicator data of the third database, it is determined to switch the first database to the primary database.

6. The method according to claim 4 or 5, characterized in that, The first storage system also stores the value of a failover key, which includes a first value and a second value. The first value indicates that the system is not in a failover state, and the second value indicates that the system is in a failover state. Before switching the first database to the primary database, the method further includes: If the value of the fault switching key stored in the first storage system is the first value and the value of the fault switching key is not locked, then the value of the fault switching key is locked, and the value of the fault switching key stored in the first storage system is modified to the second value. After switching the first database to the primary database, the method further includes: Modify the value of the failover key stored in the first storage system to the first value, and release the lock on the value of the failover key.

7. The method according to any one of claims 1-6, characterized in that, The first storage system also stores the value of a normal switchover key, which includes a third value, a fourth value, and a fifth value. The third value indicates that the system is not in a normal switchover state, the fourth value indicates that the system is in a normal switchover state during the degradation phase, and the fifth value indicates that the system is in a normal switchover state during the promotion phase. The method further includes: If the value of the normal switching key stored in the first storage system is the third value, the value of the normal switching key stored in the first storage system is modified to the fourth value; the second management platform is used to switch the second database to a backup database and modify the value of the normal switching key stored in the first storage system to the fifth value when the value of the normal switching key stored in the first storage system is the fourth value. If the value of the normal switching key stored in the first storage system is the fifth value, the first database is switched to the primary database, and the value of the normal switching key stored in the first storage system is modified to the third value.

8. The method according to any one of claims 1-7, characterized in that, The infrastructure also includes a second storage system, which provides the IP address of the primary database to the client. Switching the first database to the primary database further includes: The IP address of the main database stored in the second storage system is modified to the IP address of the first database, so that the client can establish a connection with the first database through the IP address of the first database.

9. A database management device based on distributed technology, characterized in that, The database management device is applied to a first management platform, which runs on an infrastructure. The first management platform manages a first database, which serves as a backup database for a second database. A second management platform is also deployed on the infrastructure, and this second platform manages the second database. The database management device includes: The detection module is used to send detection information to the second database, obtain a first detection result of the second database, and send the first detection result to the first storage system. The first storage system is deployed on the infrastructure and stores all detection results related to the second database. The all detection results include the first detection result and the second detection result. The second detection result is obtained by the second management platform sending detection information to the second database. The switching module is used to determine whether the second database is faulty based on all the detection results, and, in the event that the second database is faulty, to switch the first database to the primary database.

10. The database management device according to claim 9, characterized in that, The first storage system includes multiple storage nodes, each of which stores all the detection results. The first database is located in the first node, and the second database is located in the second node. The infrastructure includes multiple nodes, including the first node and the second node.

11. The database management device according to claim 10, characterized in that, The first management platform is located in the first node, the second management platform is located in the second node, the first node and the second node are located in different availability zones or different regions of the infrastructure, and the plurality of storage nodes are located in at least two availability zones or at least two regions of the infrastructure.

12. The database management device according to claim 10 or 11, characterized in that, A third management platform is also deployed on the infrastructure. The third management platform is used to manage a third database, which is a backup database for the second database. All detection results also include third detection results, which are obtained by the third management platform sending detection information to the second database.

13. The database management device according to claim 12, characterized in that, The first storage system also stores indicator data from the first database and indicator data from the third database. The switching module is further configured to: Based on the indicator data of the first database stored in the first storage system and the indicator data of the third database, it is determined to switch the first database to the primary database.

14. The database management device according to claim 12 or 13, characterized in that, The first storage system also stores the value of a failover key, which includes a first value and a second value. The first value indicates that the system is not in a failover state, and the second value indicates that the system is in a failover state. Before switching the first database to the primary database, the switching module is further configured to: lock the value of the fault switching key when the value of the fault switching key stored in the first storage system is the first value and the value of the fault switching key is not locked, and modify the value of the fault switching key stored in the first storage system to the second value. After switching the first database to the primary database, the switching module is further configured to: modify the value of the fault switching key stored in the first storage system to the first value, and release the lock on the value of the fault switching key.

15. The database management device according to any one of claims 9-14, characterized in that, The first storage system also stores the value of a normal switchover key, which includes a third value, a fourth value, and a fifth value. The third value indicates that the system is not in a normal switchover state, the fourth value indicates that the system is in a normal switchover state during the degradation phase, and the fifth value indicates that the system is in a normal switchover state during the promotion phase. The switching module is further configured to: modify the value of the normal switching key stored in the first storage system to the fourth value when the value of the normal switching key stored in the first storage system is the third value; The second management platform is used to switch the second database to a backup database when the value of the normal switching key stored in the first storage system is the fourth value, and to modify the value of the normal switching key stored in the first storage system to the fifth value. The switching module is further configured to: switch the first database to the primary database when the value of the normal switching key stored in the first storage system is the fifth value, and modify the value of the normal switching key stored in the first storage system to the third value.

16. The database management device according to any one of claims 9-15, characterized in that, The infrastructure also includes a second storage system, which provides the IP address of the main database to the client. The switching module is further configured to: modify the IP address of the main database stored in the second storage system to the IP address of the first database, so that the client can establish a connection with the first database through the IP address of the first database.

17. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1-8.

18. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster causes the computing device cluster to perform the method as described in any one of claims 1-8.

19. A computer-readable storage medium, characterized in that, Includes computer program instructions, which, when executed by a cluster of computing devices, perform the method as described in any one of claims 1-8.