Cloud service system and disaster recovery management method thereof
By establishing network connections between execution components of different cloud management platforms, automated disaster recovery management is achieved in cross-cloud or same-cloud environments. This solves the problems of poor communication and low efficiency of manual operation in existing technologies, and improves the efficiency and reliability of disaster recovery management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-29
- Publication Date
- 2026-03-31
AI Technical Summary
When data clusters and disaster recovery clusters use different cloud management platforms, existing technologies struggle to achieve efficient and automated disaster recovery management, leading to poor communication and low efficiency of manual operations, which affects the reliability and efficiency of disaster recovery management.
By establishing a network connection between the first execution component and the second execution component, instructions can be transmitted and automated management can be achieved, avoiding direct communication dependence between the first management system and the second management system. Disaster recovery management instructions and responses can be transmitted using the network connection between the first execution component and the second execution component, thereby achieving automated disaster recovery management across clouds or within the same cloud.
It enables automated disaster recovery management across cloud or same-cloud environments, improving the efficiency and reliability of disaster recovery management, reducing reliance on manual operations, and ensuring the efficient execution of disaster recovery management.
Smart Images

Figure CN121764993A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud service technology, and in particular to a cloud service system and its disaster recovery management method. Background Technology
[0002] Database systems are used to store business data. To ensure high availability, not only are data clusters created for the database system, but also disaster recovery clusters are created for these data clusters. The entire database system, including the data clusters and the disaster recovery clusters, can be called a database system with disaster recovery (DR) capabilities. The data clusters respond to read and write requests from business clients and synchronize the business data to the disaster recovery cluster. This allows the disaster recovery cluster to take over and continue providing services when the data cluster becomes unavailable due to failure or other reasons, thus ensuring the normal operation of the business.
[0003] Currently, the establishment and termination of disaster recovery relationships between data clusters and disaster recovery clusters, as well as the management during disaster recovery, all require communication between the cloud management system that manages the data clusters and the cloud management system that manages the disaster recovery clusters to ensure the coordination and management of tasks.
[0004] However, in some scenarios, such as when data clusters and disaster recovery clusters use different cloud management platforms for disaster recovery, there may be issues such as network connectivity problems or different account systems between the cloud management system that manages the data cluster and the cloud management system that manages the disaster recovery cluster. This can prevent the two cloud management systems from communicating directly, thus affecting the execution of disaster recovery management operations. Summary of the Invention
[0005] This application provides a cloud service system and its disaster recovery management method. This application can conveniently achieve disaster recovery management between data clusters and disaster recovery clusters in a database system without relying on communication between a first management system and a second management system. The technical solution provided by this application is as follows:
[0006] In a first aspect, this application provides a cloud service system, comprising: a first cloud system and a second cloud system. The first cloud system includes: a first management system and a first data system. The first management system manages the first data system. The first data system provides cloud services. A first execution component is deployed in the first data system. The second cloud system includes: a second management system and a second data system. The second management system manages the second data system. The second data system provides cloud services. A second execution component is deployed in the second data system. A network connection is established between the first execution component and the second execution component. The first management system is used to send a first disaster recovery management instruction to the first execution component, which instructs the first execution component to perform disaster recovery management operations for the data cluster in the first data system for the second cloud system. The first execution component, based on the first disaster recovery management instruction, sends a first disaster recovery management request to the second execution component via a network connection, which requests the execution of disaster recovery management operations for the data cluster in the second cloud system. The second execution component forwards the first disaster recovery management request to the second management system. The second management system, based on the first disaster recovery management request, sends a second disaster recovery management instruction to the second execution component. The second execution component is also used to perform disaster recovery management operations for the data cluster in the second cloud system based on the second disaster recovery management instruction. Furthermore, the first execution component is also used to perform disaster recovery management operations for the data cluster in the first cloud system based on the first disaster recovery management instruction.
[0007] As can be seen from the above, in the cloud service system provided in this application, since the first execution component can send a first disaster recovery management request to the second execution component via a network connection, and the first management system and the first execution component can communicate with each other, as can the second management system and the second execution component, instructions related to disaster recovery management between the first cloud system and the second cloud system can be transmitted through the network connection between the first execution component and the second execution component. Thus, without relying on communication between the first management system and the second management system, instructions related to disaster recovery management can be transmitted between the first cloud system and the second cloud system, enabling convenient disaster recovery management between the data cluster and the disaster recovery cluster in the database system. Furthermore, the first cloud system and the second cloud system can automatically implement disaster recovery management based on the transmitted instructions, without relying on manual execution steps, ensuring the efficiency and reliability of disaster recovery management.
[0008] In one possible implementation, the second execution component may require additional information to perform disaster recovery management operations on the data cluster in the second cloud system. The first disaster recovery management request carries the first information required by the second execution component to perform the disaster recovery management operations on the data cluster. Specifically, the second execution component performs the disaster recovery management operations on the data cluster in the second cloud system based on the second disaster recovery management instructions and the first information.
[0009] Similarly, the second execution component is also used to respond to the second disaster recovery management instruction by sending a first disaster recovery management response to the first execution component via a network connection. The first disaster recovery management response carries second information required by the first execution component to perform disaster recovery management operations for the data cluster. Accordingly, the first execution component is specifically used to perform disaster recovery management operations for the data cluster in the first cloud system based on the second information and the first disaster recovery management instruction.
[0010] In one possible implementation scenario, the first execution component performs disaster recovery management operations for the data cluster in the first cloud system, including executing one or more tasks. This is based on a first disaster recovery management instruction, and includes: the first execution component executing one or more tasks included in the disaster recovery management operation for the data cluster in the first cloud system. When the first execution component performs disaster recovery management operations involving the execution of multiple tasks in a sequential order, this is based on the first disaster recovery management instruction, and includes: the first execution component executing a first task included in the disaster recovery management operation for the data cluster in the first cloud system; the first execution component receiving a third task execution instruction sent by the first management system after the first task has been completed; and the first execution component executing a third task based on the third task execution instruction. The first task and the third task are two of the multiple tasks included in the disaster recovery management operation.
[0011] In some implementation scenarios, the second execution component performs disaster recovery management operations for the data cluster in the second cloud system, including executing one or more tasks. The first execution component performs disaster recovery management operations for the data cluster in the first cloud system, including executing multiple tasks. At least some of the multiple tasks executed by the first execution component depend on the execution results of at least some of the one or more tasks executed by the second execution component. Therefore, the first execution component is specifically used to execute a first task, including the disaster recovery management operation for the data cluster, in the first cloud system based on the first disaster recovery management instruction; the second execution component is specifically used to execute a second task, including the disaster recovery management operation for the data cluster, in the second cloud system based on the second disaster recovery management instruction; the first management system is specifically used to instruct the first execution component to confirm whether the second task has been completed before the first execution component executes the third task, including the disaster recovery management operation for the data cluster; the first execution component is specifically used to confirm whether the second task has been completed to the second execution component via a network connection, and to receive a confirmation result sent by the second execution component to the first execution component via a network connection indicating whether the second task has been completed; the first execution component is specifically used to send the confirmation result to the first management system; the first management system is specifically used to send a third task execution instruction to the first execution component when the confirmation result indicates that the second task has been completed; the first execution component is specifically used to execute the third task in the first cloud system based on the third task execution instruction.
[0012] In one possible implementation, task completion can be indicated by a flag. For example, the second execution component is further configured to record a flag indicating that the second task has been completed after completing the second task. Then, if a flag is recorded, the second execution component sends an acknowledgment to the first execution component via a network connection, indicating that the second task has been completed. This method of indicating task completion by flag is simple, reduces the implementation's dependence on external systems, and minimizes the communication link required to determine task completion.
[0013] In one possible implementation, disaster recovery management operations include one or more of the following: establishing a disaster recovery cluster for the data cluster in the second data system; performing primary / standby failover operations on the data cluster and the disaster recovery cluster; and terminating the disaster recovery relationship between the data cluster and the disaster recovery cluster. This solution supports disaster recovery setup, disaster recovery recovery addition, primary / standby cluster switching, disaster recovery promotion to primary, and disaster recovery termination processes, filling product gaps and reducing customers' cross-cloud costs.
[0014] In one possible implementation, the first cloud system and the second cloud system belong to different cloud resource deployment regions managed by the same cloud management platform. The disaster recovery management solution implemented based on this is a same-cloud disaster recovery solution between different regions under the same cloud management platform. Alternatively, the first cloud system and the second cloud system are managed by different cloud management platforms. The disaster recovery solution implemented based on this is a cross-cloud disaster recovery solution under different cloud management platforms. In this way, this application is equivalent to providing both cross-cloud management platform disaster recovery solutions and same-cloud management platform disaster recovery solutions in cross-regional disaster recovery scenarios, achieving a unified cross-cloud and same-cloud disaster recovery architecture in cross-regional disaster recovery scenarios.
[0015] In one possible implementation, a first management system is used to manage the infrastructure of a first cloud system, and a data cluster in a first data system is a first database instance created based on the infrastructure, which is used to provide database cloud services; a second management system is used to manage the infrastructure of a second cloud system, and a disaster recovery cluster in a second data system is a second database instance created based on the infrastructure, which is used to provide database cloud services.
[0016] Secondly, this application provides a disaster recovery management method for a cloud service system, which includes a first cloud system and a second cloud system. The first cloud system includes a first management system and a first data system. The first management system manages the first data system. The first data system provides cloud services. A first execution component is deployed in the first data system. The second cloud system includes a second management system and a second data system. The second management system manages the second data system. The second data system provides cloud services. A second execution component is deployed in the second data system. A network connection is established between the first execution component and the second execution component. The method includes: a first management system sending a first disaster recovery management instruction to a first execution component, the first disaster recovery management instruction instructing the first execution component to perform a disaster recovery management operation for a data cluster in a first data system targeting a second cloud system; the first execution component, based on the first disaster recovery management instruction, sending a first disaster recovery management request to a second execution component via a network connection, the first disaster recovery management request requesting the execution of a disaster recovery management operation for the data cluster in the second cloud system; the second execution component forwarding the first disaster recovery management request to a second management system; the second management system, based on the first disaster recovery management request, sending a second disaster recovery management instruction to the second execution component; the second execution component, based on the second disaster recovery management instruction, performing a disaster recovery management operation for the data cluster in the second cloud system; and the first execution component, based on the first disaster recovery management instruction, performing a disaster recovery management operation for the data cluster in the first cloud system.
[0017] In one possible implementation, the first disaster recovery management request carries first information required by the second execution component to perform disaster recovery management operations for the data cluster. The second execution component performs disaster recovery management operations for the data cluster in the second cloud system based on the second disaster recovery management instruction, including: the second execution component performs disaster recovery management operations for the data cluster in the second cloud system based on the second disaster recovery management instruction and the first information.
[0018] Optionally, the method further includes: a second execution component responding to a second disaster recovery management instruction by sending a first disaster recovery management response to a first execution component via a network connection, the first disaster recovery management response carrying second information required by the first execution component to perform disaster recovery management operations for the data cluster. Accordingly, the first execution component, based on the first disaster recovery management instruction, performs disaster recovery management operations for the data cluster in the first cloud system, including: the first execution component performing disaster recovery management operations for the data cluster in the first cloud system based on the second information and the first disaster recovery management instruction.
[0019] In one possible implementation, the second execution component performs disaster recovery management operations for the data cluster in the second cloud system based on the second disaster recovery management instructions, including: the second execution component performs a second task in the second cloud system for the disaster recovery management operations for the data cluster based on the second disaster recovery management instructions.
[0020] The first execution component, based on the first disaster recovery management instruction, performs disaster recovery management operations for the data cluster in the first cloud system, including: the first execution component, based on the first disaster recovery management instruction, performs a first task included in the disaster recovery management operation for the data cluster in the first cloud system; before the first execution component performs the third task included in the disaster recovery management operation for the data cluster, the first management system instructs the first execution component to confirm whether the second task has been completed; the first execution component confirms whether the second task has been completed to the second execution component via a network connection, and receives a confirmation result sent by the second execution component to the first execution component via a network connection indicating whether the second task has been completed, and sends the confirmation result to the first management system; if the confirmation result indicates that the second task has been completed, the first management system sends a third task execution instruction to the first execution component; the first execution component performs the third task in the first cloud system based on the third task execution instruction.
[0021] In one possible implementation, the method further includes: after completing the second task, the second execution component records a flag indicating that the second task has been completed, and if the flag is recorded, sends a confirmation result indicating that the second task has been completed to the first execution component via a network connection.
[0022] In one possible implementation, disaster recovery management operations include one or more of the following: establishing a disaster recovery cluster for the data cluster in the second data system; performing primary / standby failover operations on the data cluster and the disaster recovery cluster; and terminating the disaster recovery relationship between the data cluster and the disaster recovery cluster.
[0023] In one possible implementation, the first cloud system and the second cloud system belong to different cloud resource deployment areas managed by the same cloud management platform, or the first cloud system and the second cloud system are managed by different cloud management platforms.
[0024] In one possible implementation, a first management system is used to manage the infrastructure of a first cloud system, and a data cluster in a first data system is a first database instance created based on the infrastructure, which is used to provide database cloud services; a second management system is used to manage the infrastructure of a second cloud system, and a disaster recovery cluster in a second data system is a second database instance created based on the infrastructure, which is used to provide database cloud services.
[0025] Thirdly, this application provides a computing device, including a memory and a processor, wherein the memory stores program instructions and the processor executes the program instructions to implement the functions of the first management system, the first data system, the second management system, or the second data system in the first aspect of this application and any possible implementation thereof.
[0026] Fourthly, this application provides a computing device cluster, including multiple computing devices, each computing device including multiple processors and multiple memories, the multiple memories storing program instructions, and the multiple processors executing the program instructions, causing the computing device cluster to perform the methods provided in the second aspect of this application and any possible implementation thereof.
[0027] Fifthly, this application provides a computer-readable storage medium that is a non-volatile computer-readable storage medium, the computer-readable storage medium including program instructions that, when executed on a computing device, cause the computing device to perform the methods provided in the second aspect of this application and any of its possible implementations.
[0028] Sixthly, this application provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the methods provided in the second aspect of this application and any possible implementation thereof. (See attached drawings.)
[0029] Figure 1 This is a structural diagram of an implementation scenario involving a cloud service system and its disaster recovery management method provided in the embodiments of this application;
[0030] Figure 2This is a schematic diagram illustrating the deployment of basic resources in a data center, provided in an embodiment of this application.
[0031] Figure 3 This is a schematic diagram of the structure of a cloud service system provided in an embodiment of this application;
[0032] Figure 4 This is a flowchart of a disaster recovery management method for a cloud service system provided in an embodiment of this application;
[0033] Figure 5 This is a schematic diagram of another cloud service system provided in an embodiment of this application;
[0034] Figure 6 This is a flowchart illustrating how a first execution component performs disaster recovery management operations for a data cluster in a first cloud system, as provided in an embodiment of this application.
[0035] Figure 7 This is a schematic diagram illustrating whether a task has been completed by a marker, as provided in an embodiment of this application.
[0036] Figure 8 This is a schematic diagram illustrating the implementation process of disaster recovery setup provided in an embodiment of this application;
[0037] Figure 9 This is a schematic diagram illustrating the implementation process of a primary / standby cluster switchover provided in an embodiment of this application;
[0038] Figure 10 This is a schematic diagram illustrating the implementation process of disaster recovery and mitigation provided in an embodiment of this application;
[0039] Figure 11 This is a schematic diagram illustrating the implementation process of disaster recovery setup provided in an embodiment of this application;
[0040] Figure 12 This is a schematic diagram illustrating an implementation process with dependencies on task execution provided in an embodiment of this application;
[0041] Figure 13 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application;
[0042] Figure 14 This is a schematic diagram of the structure of a computing device cluster provided in an embodiment of this application;
[0043] Figure 15 This is a schematic diagram of another computing device cluster structure provided in an embodiment of this application. Detailed Implementation
[0044] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0045] To facilitate understanding, the technologies and background involved in the embodiments of this application will be introduced below.
[0046] Cloud computing is a type of distributed computing that refers to a network that centrally manages and schedules a large number of computing and storage resources to provide on-demand services to users. These computing and storage resources are provided through clusters of computing devices located in data centers. Furthermore, cloud computing can provide users with various types of services, such as Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and Software as a Service (SaaS). Infrastructure as a Service provides virtual machines or other resources as a service to tenants. Platform as a Service provides a development platform as a service to tenants. Software as a Service provides applications (Apps) as a service to customers.
[0047] In hybrid cloud scenarios, the financial and banking industries have increasingly higher requirements for data security and reliability, and correspondingly, higher requirements for database disaster recovery. Current database disaster recovery solutions generally include intra-city disaster recovery solutions and cross-regional disaster recovery solutions. Cross-regional disaster recovery solutions include: intra-cloud resource deployment area (Region) cross-availability zone (AZ) disaster recovery solutions, intra-cloud cross-Region disaster recovery solutions, and cross-cloud disaster recovery solutions. Intra-city disaster recovery solutions can guarantee the security of business data in most scenarios, but in extreme situations, such as fires, earthquakes, wars, and other extreme disasters, it is difficult to guarantee data security. Therefore, cross-regional disaster recovery solutions are needed. Cross-regional disaster recovery solutions typically refer to situations where the distance between the primary and backup data centers is more than 200 kilometers (KM). Even if the primary data center cannot continue to provide services due to the aforementioned extreme disasters in its region, the backup data center still has the ability to continue providing services.
[0048] For building cross-region disaster recovery solutions, most customers use different cloud management platforms to construct data centers across different regions, while some customers use the cross-region capabilities of the same cloud management platform. Currently, there is no unified solution that can be compatible with both cross-cloud and intra-cloud scenarios while ensuring a good recovery point objective (RPO). The following are some existing cross-region disaster recovery solutions:
[0049] 1. Under the same cloud management platform, a single-cluster, multi-replica model is used for cross-regional data center deployment. This approach deploys one cluster across two regions, managing the data centers in both regions as multiple instance nodes (AZs) within a single region. Shards are evenly distributed across different data centers, ensuring that each write request is simultaneously persisted to disk in both data centers, mitigating single-data center failures. However, because this solution uses a single-cluster deployment, it cannot achieve fault isolation. Failures in cluster management components or other regional faults will render the entire cluster service unavailable. Furthermore, when using traditional network-based log synchronization between database primary and standby nodes, increased geographical distance between the nodes will significantly increase transmission latency, directly impacting production service performance. Additionally, this solution is unsuitable for multi-cloud platform customer scenarios.
[0050] 2. Under the same cloud management platform, a dual-cluster solution is adopted to ensure data cluster performance and fault domain isolation. This solution manages two geographically separated clusters as two regions, and data synchronization between the primary and backup clusters is achieved through cascading backup. However, because this solution uses a single cloud management platform, which lacks the capability to withstand extreme disasters, manual operation is required when the cloud management platform is damaged or in emergency situations, leading to inefficiency and a high risk of errors. Furthermore, this solution is incompatible with multi-cloud platform customer scenarios.
[0051] 3. Under different cloud management platforms, a dual-cluster solution is adopted to ensure the performance of the primary cluster and fault domain isolation. Cluster data synchronization relies on the database's own synchronization and replication capabilities, while scheduling between management planes depends on additional services. The two geographically separated clusters are managed as two different cloud environments, and data synchronization between the primary and backup clusters is achieved through cascading backup. Furthermore, disaster recovery operations require establishing management plane networks between the two operation and management platforms, which relies on external cloud services calling relevant interfaces of the two cloud platforms. This solution depends on the development of external tools, leading to compatibility issues. The isolation of management plane networks between different cloud management platforms, coupled with differing authentication services across platforms, makes cross-regional disaster recovery operations impossible.
[0052] 4. Under different cloud management platforms, data migration tools are used to synchronize data from the primary cluster to the disaster recovery cluster using logical replication capabilities. The two clusters, located in different regions, are managed as two distinct cloud environments. Data migration tools are used between the two clusters to read data from the primary cluster and write it to the disaster recovery cluster. This solution relies on data migration tools, and due to the use of logical replication, the migration performance is somewhat lower than physical replication. During peak business periods, the RPO between the primary and backup clusters is relatively high.
[0053] Therefore, for disaster recovery solutions using different cloud management platforms for data clusters and disaster recovery clusters, the establishment and termination of the disaster recovery relationship between the data clusters and disaster recovery clusters, as well as management during the disaster recovery period, all require communication between the two cloud management platforms to ensure task coordination and management. However, in cross-cloud disaster recovery scenarios, there may be situations where the cloud management platforms of the two clouds are not interconnected or have different account systems, making it virtually impossible for the two cloud management platforms to communicate directly, thus affecting the execution of disaster recovery management operations. Similarly, in other scenarios where disaster recovery requires communication between the cloud management plane that manages the data cluster and the cloud management plane that manages the disaster recovery cluster, the two cloud management planes may also be unable to communicate directly, which will also affect the execution of disaster recovery management operations. Although it is currently possible to manually log in to the two clusters separately and manually perform disaster recovery setup and other disaster recovery management operations, relying on manual execution of multiple process steps is significantly less efficient than automating disaster recovery setup and other disaster recovery management operations. Meanwhile, relying on manual execution of multiple process steps inevitably leads to human error, and increases recovery time after a failure, resulting in reduced disaster recovery reliability of the database system. Therefore, neither cloud service management plane communication nor manual setup can effectively solve the current problems in database system disaster recovery management.
[0054] In view of this, embodiments of this application provide a cloud service system and its disaster recovery management method. The cloud service system includes: a first cloud system and a second cloud system. The first cloud system includes: a first management system and a first data system. The first management system is used to manage the first data system. The first data system is used to provide cloud services. A first execution component is deployed in the first data system. The second cloud system includes: a second management system and a second data system. The second management system is used to manage the second data system. The second data system is used to provide cloud services. A second execution component is deployed in the second data system. A network connection is established between the first execution component and the second execution component.
[0055] In this cloud service system, the first management system sends a first disaster recovery management instruction to the first execution component, which instructs the first execution component to perform disaster recovery management operations for the data cluster in the first data system targeting the second cloud system. The first execution component, based on the first disaster recovery management instruction, sends a first disaster recovery management request to the second execution component via a network connection, requesting the execution of disaster recovery management operations for the data cluster in the second cloud system. The second execution component forwards the first disaster recovery management request to the second management system. The second management system, based on the first disaster recovery management request, sends a second disaster recovery management instruction to the second execution component. The second execution component also performs disaster recovery management operations for the data cluster in the second cloud system based on the second disaster recovery management instruction. Furthermore, the first execution component performs disaster recovery management operations for the data cluster in the first cloud system based on the first disaster recovery management instruction.
[0056] In this way, since the first execution component can send a first disaster recovery management request to the second execution component via a network connection, and the first management system and the first execution component can communicate with each other, as can the second management system and the second execution component, disaster recovery management-related instructions between the first cloud system and the second cloud system can be transmitted through the network connection between the first execution component and the second execution component. This allows for the transmission of disaster recovery management-related instructions between the first cloud system and the second cloud system without relying on communication between the first management system and the second management system, facilitating convenient disaster recovery management between the data cluster and the disaster recovery cluster in the database system. Furthermore, the first cloud system and the second cloud system can automatically implement disaster recovery management based on the transmitted instructions, without relying on manual execution steps, ensuring the efficiency and reliability of disaster recovery management.
[0057] This article provides a detailed introduction to the technical solution of this application from multiple perspectives, including implementation scenarios, methods and processes, hardware devices, and software devices.
[0058] The following are examples illustrating the implementation scenarios of the embodiments of this application.
[0059] Figure 1 This is a structural diagram illustrating an implementation scenario of a cloud service system and its disaster recovery management method provided in this application embodiment. For example... Figure 1 As shown, the implementation scenario includes: data center 1 and client 2. Data center 1 and client 2 can establish a communication connection via a network. Optionally, this network can be the Internet, or other networks; this embodiment is not limited to any particular network. Tenants can interact with data center 1 through client 2. For example, a tenant can send cloud service requests and other information to data center 1 through client 2. Data center 1 responds based on the information sent by client 2.
[0060] Data Center 1 houses a large amount of infrastructure owned by the cloud service provider, such as computing resources, storage resources, and network resources. For example, computing resources can be computing devices (such as servers) capable of providing computing power. Figure 1 As shown, data center 1 includes a cloud management platform and infrastructure ( Figure 1 (Not shown in the image). The cloud management platform and the infrastructure are connected via an internal data center network. The cloud management platform is used to manage the infrastructure. The infrastructure is used to provide public cloud services. The infrastructure includes multiple servers. Cloud services are optionally deployed on the servers. Cloud services are implemented by running virtual instances, and are therefore also referred to as virtual instances deployed on servers to implement tenant services. Tenants can send cloud service requests and related information to the server through their client 2. The server can process the cloud service requests and related information and provide cloud services to the tenant based on the processed cloud service requests and related information. For example, the server can perform disaster recovery management using the disaster recovery management method of the cloud service system provided in this application embodiment.
[0061] The cloud management platform can be logically divided into: tenant console, compute management service, network management service, storage management service, authentication service, and image management service. The tenant console provides a user interface or application programming interface (API) for interaction with tenants. The compute management service manages servers running virtual instances and bare metal servers. The network management service manages network services (such as gateways and firewalls). The storage management service manages storage services (such as data bucket services). The authentication service manages tenant accounts and passwords. The image management service manages virtual instance images.
[0062] exist Figure 1 In the illustrated implementation scenario, a data center contains multiple servers. The servers consist of a hardware layer and a software layer. The hardware layer comprises the standard server configuration, including processors, memory, network interface cards (NICs), disks, and buses. The software layer includes the operating system installed and running on the server. This operating system, relative to the virtual machine, can be called the host operating system. The host operating system runs a virtual machine manager (also known as a hypervisor). The hypervisor's role is to implement compute virtualization, network virtualization, and storage virtualization, and to manage the virtual machines.
[0063] The virtual machine manager runs a cloud management platform client. This client receives control plane commands from the cloud management platform, creates virtual instances on the server based on these commands, and manages the virtual instances throughout their lifecycle. For example, the client can monitor the hardware resource usage of the server in real time and report it to the cloud management platform. When the cloud management platform confirms that a virtual instance needs to be created on a specific server, it sends a virtual instance creation command to the client on that server. Upon receiving the command, the client creates the virtual instance on that server. In this way, tenants can create, manage, log in to, and operate virtual instances within the data center through the cloud management platform.
[0064] Servers can run virtual machines of different specifications. Virtual machine specifications are categorized as: general-purpose computing, memory-optimized, ultra-large memory, etc., with specific specifications under each type. After a tenant selects a virtual machine specification, the cloud management platform selects a server in the data center that supports that specification and ensures sufficient idle hardware resources on that server. Then, it creates and configures the virtual machine with that specification on that server. Configuring servers through the cloud management platform allows for the analysis and planning of server hardware resources. Based on the server's hardware performance, it plans the corresponding computing products for the physical hardware, such as planning virtual machines of different specifications, to meet the diverse needs of different tenants. Furthermore, differentiated pricing strategies can be implemented based on the performance differences of virtual machines of different specifications. For example, high-performance virtual instances can be sold at a higher price, while ordinary performance virtual instances can be sold at a lower price, allowing tenants to purchase virtual instances as needed.
[0065] In one implementation, such as Figure 2As shown, the location of basic resources in a data center can be described using cloud resource deployment regions (regions) and availability zones (AZs). Tenants can choose to deploy cloud services based on resources within a specific region or AZ. A region is defined by geographical location and network latency. Using the same resource pool within the same region can be understood as sharing common services such as elastic computing, block storage, object storage, virtual private cloud (VPC) networks, elastic internet protocol (EIP) addresses, and images. Regions are divided into general-purpose regions and dedicated regions. A general-purpose region refers to a region that provides general cloud services to public tenants. A dedicated region refers to a region that hosts the same type of business or provides business services to specific tenants. A region typically includes multiple AZs. Multiple AZs within a region are connected via high-speed fiber optic cables to meet the needs of tenants building high-availability systems across AZs. An AZ is one or more... Figure 2 The data center shown is a collection of data centers. Within an Availability Zone (AZ), computing, networking, and storage resources are logically divided into multiple clusters.
[0066] Tenants can send instructions to the cloud management platform through their client 2 to create, manage, log in to, and operate virtual instances on the server, and use the cloud services provided by these virtual instances. For example, the cloud management platform can provide an access interface. This interface can be provided either as a user interface or an API. Tenants can operate their client to remotely access the access interface to register a cloud account and password on the cloud management platform, and then log in using these accounts and passwords. The cloud management platform can also authenticate the cloud account and password. After successful authentication, the tenant can further select and purchase a virtual instance with specific specifications (processor, memory, disk) on the cloud management platform. After the tenant successfully purchases the virtual instance, the cloud management platform provides the tenant with a remote login account and password for the purchased virtual instance. The tenant can use the remote login account and password to remotely log in to the virtual instance on their client, install and run their application within the virtual instance, and use the application to implement their business operations.
[0067] Client 2 can be selected from computers, personal computers, laptops, mobile phones, smartphones, tablets, cloud servers, portable mobile terminals, multimedia players, e-book readers, wearable devices, smart home appliances, artificial intelligence devices, smart wearable devices, smart in-vehicle devices, or Internet of Things devices, etc.
[0068] In one implementation, the disaster recovery management method of the cloud service system provided in this application embodiment can be implemented by running an executable program on a computing device in data center 1. Optionally, the disaster recovery management method of the cloud service system provided in this application embodiment can be applied to a cloud service management system. The cloud service management system is deployed on a server managed by a cloud management platform. The cloud service management system can implement the disaster recovery management method of the cloud service system provided in this application embodiment by running the executable program of the disaster recovery management method of the cloud service system provided in this application embodiment. Furthermore, the executable program implementing the disaster recovery management method of the cloud service system can be presented in the form of an application installation package. After the server installs the application installation package, it can implement the disaster recovery management method of the cloud service system provided in this application embodiment by running the executable program therein.
[0069] It should be understood that the above content is an exemplary description of the implementation scenarios of the disaster recovery management method for the cloud service system provided in the embodiments of this application, and does not constitute a limitation on the implementation scenarios of the disaster recovery management method for the cloud service system. Those skilled in the art will know that as business needs change, the implementation scenarios can be adjusted according to application requirements, and the embodiments of this application do not specifically limit them. Furthermore, when the disaster recovery management method for the cloud service system provided in the embodiments of this application is applied to other scenarios, the executable program of the method can also be presented in the form of an application installation package or in other ways, and the embodiments of this application do not list them all.
[0070] The following example illustrates the implementation process of the disaster recovery management method for a cloud service system provided in this application. This method is applied to a cloud service system. Figure 3 This is a schematic diagram of the structure of a cloud service system provided in an embodiment of this application. For example... Figure 3 As shown, the cloud service system includes: a first cloud system and a second cloud system. The first cloud system includes: a first management system and a first data system. The first management system manages the first data system. The first data system provides cloud services. A first execution component is deployed in the first data system. The second cloud system includes: a second management system and a second data system. The second management system manages the second data system. The second data system provides cloud services. A second execution component is deployed in the second data system. A network connection is established between the first execution component and the second execution component. In this application, the first data system is also used to deploy a data cluster. The second data system is also used to deploy a disaster recovery cluster.
[0071] In one possible implementation, the first management system may be deployed on a first management plane of a cloud platform. The first management plane is a network plane used to run the first management system. The first data system may be deployed on a first data plane. The first data plane is a network plane used for processing and transmitting data. For example, the first management system is a cloud management platform or a component therein, used to manage the infrastructure of the first cloud system. The data cluster in the first data system is a first database instance created based on this infrastructure. The first database instance is used to provide database cloud services. The first database instance may be an instance cluster comprising multiple instances. The second management system is a cloud management platform or a component therein, used to manage the infrastructure of the second cloud system. The disaster recovery cluster in the second data system is a second database instance created based on this infrastructure. The second database instance is used to provide database cloud services. The second database instance may be an instance cluster comprising multiple instances.
[0072] Optionally, the first cloud system and the second cloud system may belong to different cloud resource deployment regions managed by the same cloud management platform. In this case, the cloud management platform can manage both the first and second data systems. The first and second management systems can be subsystems within the same cloud management platform. Management operations of the first data system by the cloud management platform are performed by the first management system. Management operations of the second data system by the cloud management platform are performed by the second management system. The disaster recovery management solution implemented based on this is a co-cloud disaster recovery solution between different regions under the same cloud management platform.
[0073] Alternatively, the first and second cloud systems can be managed by different cloud management platforms. For example, the first management system can be managed by the first cloud management platform, in which case the first management system can be considered a subsystem of the first cloud management platform, used to manage the first data system. The second management system can be managed by the second cloud management platform, in which case the second management system can be considered a subsystem of the second cloud management platform, used to manage the second data system. The disaster recovery solution implemented based on this is a cross-cloud disaster recovery solution across different cloud management platforms.
[0074] In this way, this application provides both cross-cloud management platform disaster recovery solutions and solutions under the same cloud management platform in cross-regional disaster recovery scenarios, achieving a unified cross-cloud and same-cloud disaster recovery architecture in cross-regional disaster recovery scenarios. Furthermore, the cloud management platform of this application can be a public cloud, private cloud, or hybrid cloud, etc., and this application embodiment does not specifically limit it.
[0075] In one possible implementation, both the first and second execution components can be implemented in software. For example, both the first and second execution components can be agent components. Both the first and second management systems can be implemented in software or hardware. For example, when the first and second cloud systems are managed by different cloud management platforms, the first and second management systems are implemented through their respective cloud management platforms. In this way, the first and second execution components are equivalent to agent components built into the disaster recovery instance across cloud management platforms. Communication between the two agent components transmits instructions between the different cloud management platforms, enabling communication between them and solving the problem of communication barriers between different management systems in cross-cloud management platform scenarios. For example, it solves the problem of communication barriers between cloud management platforms due to different account systems and authentication / authorization requirements in cross-cloud scenarios. Another example is when the first and second cloud systems belong to different cloud resource deployment regions managed by the same cloud management platform, the first and second management systems are implemented through management components deployed on the same cloud management platform in different cloud resource deployment regions. In this way, the first execution component and the second execution component are equivalent to agent components built into instances in different cloud resource deployment areas of the same cloud. Through communication between the two agent components, instructions to the management components in different cloud resource deployment areas of the same cloud are transmitted, thereby realizing communication between the management components in different cloud resource deployment areas of the same cloud.
[0076] Figure 4 This is a flowchart illustrating a disaster recovery management method for a cloud service system provided in an embodiment of this application. Figure 4 As shown, the disaster recovery management method of this cloud service system includes the following steps:
[0077] Step 401: The first management system sends a first disaster recovery management instruction to the first execution component. The first disaster recovery management instruction is used to instruct the first execution component to perform disaster recovery management operations for the second cloud system for the data cluster in the first data system.
[0078] In the first cloud system, the first management system acts as the manager, instructing the first execution component to perform specified operations. The first execution component acts as the executor, performing specified operations based on the instructions of the first management system. A communication connection is established between the first management system and the first execution component. When the first management system needs the first execution component to perform disaster recovery management operations for the second cloud system on the data cluster in the first data system, it sends a first disaster recovery management instruction to the first execution component through this communication connection, instructing the first execution component to perform the corresponding disaster recovery management operation. For example, when an application used to implement tenant services has disaster recovery management requirements, it sends a disaster recovery management instruction to the first management system. After receiving the disaster recovery management instruction, the first management system sends a first disaster recovery management instruction to the first execution component through the communication connection between the first management system and the first execution component, instructing the first execution component to perform the corresponding disaster recovery management operation. The communication connection between the first management system and the first execution component can be a network connection, or other types of connections; this embodiment does not specifically limit the specific type of connection.
[0079] When disaster recovery management operations also involve performing operations on other cloud systems, the first disaster recovery management instruction needs to carry indication information for those other cloud systems so that the first execution component can interact with those other cloud systems based on this indication information during the execution of the disaster recovery management operation. For example, when the disaster recovery management operation also involves performing related operations in the disaster recovery cluster of the second data system, the first disaster recovery management instruction carries the Internet Protocol (IP) address of the first execution component in the second cloud system.
[0080] In this application, disaster recovery management operations include one or more of the following: disaster recovery setup, primary / standby cluster switchover, disaster recovery promotion to primary, or disaster recovery termination. Disaster recovery setup refers to establishing a disaster recovery cluster for the data cluster in the first cloud system within the second cloud system, and establishing a disaster recovery relationship between the data cluster and the disaster recovery cluster. Disaster recovery setup includes initial setup of the disaster recovery relationship and re-establishment of the disaster recovery relationship. Initial setup of the disaster recovery relationship means that a disaster recovery relationship had never been established between the data cluster and the disaster recovery cluster before this establishment of the disaster recovery relationship. Re-establishment of the disaster recovery relationship means that a disaster recovery relationship had been established between the data cluster and the disaster recovery cluster before this establishment of the disaster recovery relationship, but that relationship was terminated before this establishment of the disaster recovery relationship. One example of re-establishing the disaster recovery relationship is adding the disaster recovery cluster back. Adding the disaster recovery cluster back means that, with the disaster recovery cluster in the second data system serving as the primary cluster and the data cluster in the first data system meeting the working conditions, the disaster recovery relationship between the data cluster and the disaster recovery cluster is re-established. The disaster recovery relationship between the data cluster and the disaster recovery cluster is terminated because a reliable disaster recovery relationship cannot be formed between them. For example, after a primary / standby cluster switchover is performed on the data cluster and the disaster recovery cluster, the disaster recovery relationship between them is terminated. For ease of description, the initial establishment of the disaster recovery relationship will be referred to as disaster recovery setup, and the rebuilding of the disaster recovery relationship will be explained using the example of adding back the disaster recovery relationship. Primary / standby cluster switchover refers to performing a primary / standby switchover operation on the data cluster in the first data system and the disaster recovery cluster in the second data system. The primary cluster provides services externally, and the standby cluster is used to replace the primary cluster in providing services externally when the primary cluster fails. Disaster recovery promotion to primary means that when the data cluster in the first data system is the primary cluster, and it cannot continue to provide services externally due to failure or other reasons, the disaster recovery cluster in the second data system is promoted to primary cluster to provide services externally. Disaster recovery termination means terminating the disaster recovery relationship between the data cluster in the first data system and the disaster recovery cluster in the second data system.
[0081] Step 402: Based on the first disaster recovery management instruction, the first execution component sends a first disaster recovery management request to the second execution component via a network connection. The first disaster recovery management request is used to request the execution of disaster recovery management operations for the data cluster in the second cloud system.
[0082] The first execution component performs disaster recovery management operations for the data cluster in the first data system targeting the second cloud system. This requires the cooperation of the second execution component to perform related operations within the second cloud system. Furthermore, the first and second execution components are used to exchange information between the first and second management systems. Therefore, after receiving the first disaster recovery management instruction from the first management system, the first execution component needs to send a first disaster recovery management request to the second execution component via a network connection, so that the second execution component can perform disaster recovery management operations for the data cluster in the second cloud system. As described above, the first disaster recovery management instruction carries the IP address of the first execution component in the second cloud system; therefore, the first execution component can send the first disaster recovery management request to the second execution component via a network connection based on this IP address.
[0083] Optionally, the first disaster recovery management request carries first information required by the second execution component to perform disaster recovery management operations on the data cluster. For example, when the disaster recovery management operation also involves performing related operations in the second data system, and these operations require configuration and authentication information of the data cluster in the first cloud system, the first disaster recovery management request also carries first information including the data cluster's configuration and authentication information. The configuration information may optionally include the Internet Protocol (IP) addresses of the data nodes in the data cluster and the data cluster's topology information, etc. The data cluster's topology information is used to indicate the components deployed in the data cluster and their connections. The authentication information is used for authentication and authorization operations in the first cloud system. For example, the authentication information may be the data cluster's management account and password.
[0084] Step 403: The second execution component forwards the first disaster recovery management request to the second management system.
[0085] In the second cloud system, the second management system acts as the manager, instructing the second execution component to perform specified operations. The second execution component acts as the executor, performing specified operations based on the instructions of the second management system. A communication connection is established between the second management system and the second execution component. After receiving a first disaster recovery management request from the first execution component, the second execution component forwards the request to the second management system through this communication connection, allowing the second management system to instruct the second execution component on how to respond to the request. When the first disaster recovery management request carries first information required for the second execution component to perform disaster recovery management operations on the data cluster, the second execution component also needs to carry this first information when forwarding the request to the second management system, so that the second management system can decide how to respond to the request based on this information. Optionally, after receiving the first disaster recovery management request, the second execution component can also send an acknowledgment message (such as an ACK message) to the first execution component to indicate that it has received the request.
[0086] Communication between the second management system and the second execution component may employ a one-way communication mechanism. This one-way mechanism requires that after the second execution component sends information to the second management system, the information must be consumed by the second management system before the second execution component can receive it. Optionally, this one-way communication mechanism also requires that information sent by the second management system to the second execution component can be directly received by the second execution component. In one possible implementation, such as... Figure 5 As shown, the communication between the second management system and the second execution component (agent component) adopts the Kafka communication mechanism. The Kafka communication mechanism ensures that information sent by the second management system to the second execution component can be directly received by the second execution component, and that information sent by the second execution component to the second management system needs to be consumed by the second management system before the second management system can obtain the information. In this Kafka communication mechanism, the second execution component needs to report information to the second management system to the corresponding Kafka topic, and the second management system needs to consume the information from the corresponding Kafka topic before it can obtain the information. Similarly, the communication between the first management system and the first execution component can also optionally adopt the Kafka communication mechanism. By using a one-way communication mechanism between the management system and the execution component, direct access of the execution component to the management system can be avoided, achieving isolation between the execution component and the management system and ensuring the security of the management system. It should be noted that other communication mechanisms can also be used between the management system and the execution component, such as message queues, but this embodiment does not specifically limit them.
[0087] Step 404: The second management system sends a second disaster recovery management instruction to the second execution component based on the first disaster recovery management request.
[0088] After receiving the first disaster recovery management request forwarded by the second execution component, the second management system needs to decide whether to execute the disaster recovery management operation requested in the first disaster recovery management request. If the second management system decides to execute the disaster recovery management operation requested in the first disaster recovery management request, it sends a second disaster recovery management instruction to the second execution component to instruct it to execute the requested operation. In one possible implementation, when deciding whether to execute the disaster recovery management operation requested in the first disaster recovery management request, the second management system may optionally perform authentication and authorization operations based on the authentication information carried in the first disaster recovery management request. When the result of the authentication and authorization operations indicates that the second cloud system has the authority to execute the requested disaster recovery management operation, it determines to execute the requested operation. If the second management system decides to execute the requested operation, it also stores data related to the disaster recovery management operation in a management metadata table to facilitate management of the execution process based on this data.
[0089] Step 405: The second execution component performs disaster recovery management operations for the data cluster in the second cloud system based on the second disaster recovery management instructions.
[0090] After receiving the second disaster recovery management instruction, the second execution component can perform disaster recovery management operations on the data cluster in the second cloud system according to the instructions of the second disaster recovery management instruction. Optionally, in response to the second disaster recovery management instruction, the second execution component also sends a first disaster recovery management response to the first execution component via a network connection. The first disaster recovery management response may optionally carry second information required by the first execution component to perform disaster recovery management operations on the data cluster. For example, when performing disaster recovery management operations in the first cloud system requires configuration information and authentication information of the disaster recovery cluster in the second cloud system, the first disaster recovery management response also carries second information including the configuration information and authentication information of the disaster recovery cluster. The configuration information of the disaster recovery cluster includes the IP addresses of the data nodes in the disaster recovery cluster and the topology information of the disaster recovery cluster, etc. The topology information of the disaster recovery cluster is used to indicate the components deployed in the disaster recovery cluster and their connection relationships, etc. The authentication information is used for authentication and authorization operations in the first cloud system. For example, the authentication information is the management account and password of the disaster recovery cluster.
[0091] As described above, the first disaster recovery management request may optionally carry first information required by the second execution component to perform disaster recovery management operations for the data cluster. In this case, the second execution component performs disaster recovery management operations for the data cluster in the second cloud system based on the second disaster recovery management instruction, including: the second execution component performs disaster recovery management operations for the data cluster in the second cloud system based on the second disaster recovery management instruction and the first information.
[0092] In one possible implementation scenario, the second execution component performs disaster recovery management operations for the data cluster in the second cloud system, including executing one or more tasks. This involves the second execution component performing disaster recovery management operations for the data cluster in the second cloud system based on the second disaster recovery management instructions. Specifically, the tasks required for disaster recovery management operations such as disaster recovery setup, primary / standby cluster switching, disaster recovery recovery addition, and disaster recovery recovery discontinuation are not detailed in this application; please refer to relevant technologies for more information.
[0093] Step 406: The first execution component performs disaster recovery management operations for the data cluster in the first cloud system based on the first disaster recovery management instruction.
[0094] After receiving the first disaster recovery management instruction, the first execution component can perform disaster recovery management operations for the data cluster in the first cloud system according to the instructions of the first disaster recovery management instruction. When the second execution component sends a first disaster recovery management response to the first execution component through a network connection, and the first disaster recovery management response carries second information, the first execution component performs disaster recovery management operations for the data cluster in the first cloud system based on the first disaster recovery management instruction, including: the first execution component performs disaster recovery management operations for the data cluster in the first cloud system based on the first disaster recovery management instruction and the second information.
[0095] It should be noted that when the first disaster recovery management response carries the second information, after receiving the first disaster recovery management response, the first execution component needs to forward the first disaster recovery management response to the first management system. After receiving the first disaster recovery management response, the first management system needs to decide whether to execute the disaster recovery management operation requested by the first disaster recovery management request based on the first disaster recovery management response. If it decides to execute the disaster recovery management operation requested by the first disaster recovery management request, it sends a third disaster recovery management instruction to the first execution component to instruct the first execution component to execute the disaster recovery management operation requested by the first disaster recovery management request. In one possible implementation, when deciding whether to execute the disaster recovery management operation requested by the first disaster recovery management request, the first management system performs authentication and authorization operations based on the authentication information carried in the first disaster recovery management response. When the authentication and authorization operation results indicate that the first cloud system has the authority to execute the disaster recovery management operation requested by the first disaster recovery management request, it determines to execute the disaster recovery management operation requested by the first disaster recovery management request. When the first management system determines the disaster recovery management operation requested by the first disaster recovery management request, it will also store the data related to the execution of the disaster recovery management operation in the management metadata table so as to manage the process of executing the disaster recovery management operation based on the data.
[0096] In one possible implementation scenario, the first execution component performs disaster recovery management operations for the data cluster in the first cloud system, including executing one or more tasks. This is achieved by the first execution component executing one or more tasks included in the disaster recovery management operation based on a first disaster recovery management instruction. When the first execution component performs disaster recovery management operations involving multiple tasks with a sequential order, it performs the same operations in the first cloud system based on the first disaster recovery management instruction. This includes: the first execution component executing a first task included in the disaster recovery management operation based on the first disaster recovery management instruction; the first execution component receiving a third task execution instruction sent by the first management system after the first task is completed; and the first execution component executing a third task based on the third task execution instruction. The first and third tasks are two of the multiple tasks included in the disaster recovery management operation. The specific implementation process of the tasks required for disaster recovery management operations such as disaster recovery setup, primary / standby cluster switching, disaster recovery addition / removal, etc., is not detailed in this application; please refer to relevant technologies.
[0097] In some implementation scenarios, the second execution component execution step 405 includes executing one or more tasks, and the first execution component execution step 406 includes executing multiple tasks. Furthermore, at least some of the multiple tasks included in the first execution component execution step 406 depend on the execution results of at least some of the tasks included in the one or more tasks included in the second execution component execution step 405. Therefore, as... Figure 6 As shown, the first execution component performs disaster recovery management operations for the data cluster in the first cloud system based on the first disaster recovery management instruction, including:
[0098] Step 4061: The first execution component performs the first task, including the disaster recovery management operation for the data cluster, in the first cloud system based on the first disaster recovery management instruction.
[0099] Step 4062: Before the first execution component executes the third task, which is part of the disaster recovery management operation for the data cluster, the first management system instructs the first execution component to confirm whether the second task has been completed.
[0100] After completing the first task, the first execution component sends a notification to the first management system indicating that the first task has been completed. Upon receiving this notification, the first management system, since executing the third task depends on the execution result of the second task by the second execution component, instructs the first execution component to confirm whether the second task has been completed before executing the third task (which is part of the disaster recovery management operation for the data cluster) in the first cloud system. The communication method between the first management system and the first execution component is similar to that between the second management system and the second execution component; it will not be elaborated upon here.
[0101] Step 4063: The first execution component confirms with the second execution component via network connection whether the second task has been completed.
[0102] After receiving an instruction from the first management system to confirm whether the second task has been completed, the first execution component, based on this instruction, sends a confirmation message to the second execution component via a network connection to confirm whether the second task has been completed. In one possible implementation, the first execution component sends a query request to the second execution component via a network connection, the query request being used to check whether the second task has been completed.
[0103] Step 4064: The first execution component receives the confirmation result sent by the second execution component to the first execution component via the network connection, which indicates whether the second task has been completed, and sends the confirmation result to the first management system.
[0104] The first execution component confirms with the second execution component via a network connection whether the second task has been completed. The second execution component then queries whether the second task has been completed and, based on the query result, sends a confirmation result to the first execution component via the network connection, indicating whether the second task has been completed. After receiving the confirmation result, the first execution component sends a confirmation result to the first management system, so that the first management system can send an instruction to the first execution component to perform the next operation based on the confirmation result.
[0105] In one possible implementation, such as Figure 7 As shown, task completion can be indicated by a flag. For example, after completing the second task, the second execution component records a flag indicating that the second task has been completed. When the second execution component queries whether the second task has been completed, it can check if a flag indicating completion has been recorded. If a flag indicating completion has been recorded, the second execution component determines that the second task has been completed. If no flag indicating completion has been recorded, the second execution component determines that the second task has not been completed. The flag indicating completion can be recorded in a local file of the second execution component or another specified location. Using flags to indicate task completion is simple, reduces external dependencies, and minimizes the communication links required to determine task completion.
[0106] Step 4065: If the confirmation result indicates that the second task has been completed, the first management system sends a third task execution instruction to the first execution component.
[0107] If the first management system confirms that the second task has been completed, it sends a third task execution instruction to the first execution component to instruct the first execution component to execute the third task. If the first management system confirms that the second task has not been completed, it continues to instruct the first execution component to confirm whether the second task has been completed, until it receives a confirmation that the second task has been completed, or instructs the component to wait until a timeout occurs and terminate the disaster recovery operation.
[0108] Step 4066: The first execution component executes the third task in the first cloud system based on the third task execution instructions.
[0109] When the first execution component performs disaster recovery management operations on the data cluster in the first cloud system based on the first disaster recovery management instruction, including more tasks, it can execute the remaining tasks according to the above process until all disaster recovery management operations are completed. Simultaneously, when at least some of the tasks in step 405 performed by the second execution component depend on the execution results of at least some of the tasks in step 406 performed by the first execution component, the second management system and the second execution component can also refer to the above process to check whether the tasks they need to execute depend on have been completed. The implementation process will not be elaborated here.
[0110] Steps 401 to 407 above mainly describe the interaction process between the first management system, the first execution component, the second execution component, and the second management system during disaster recovery management. The following example, using the execution component implemented through the agent component, with the first cloud system belonging to Cloud 1 and the second cloud system belonging to Cloud 2, illustrates the implementation process of disaster recovery management, including disaster recovery setup, primary / standby cluster switching, disaster recovery recovery addition, and disaster recovery recovery termination.
[0111] Disaster recovery setup refers to establishing a disaster recovery cluster for the data cluster in the first cloud system within the second cloud system, and establishing a disaster recovery relationship between the data cluster and the disaster recovery cluster. In this context, the disaster recovery management-related content mentioned earlier refers to disaster recovery setup-related content. For example, the first disaster recovery management instruction is the first disaster recovery setup instruction, the disaster recovery management operation is the disaster recovery setup operation, the first disaster recovery management request is the first disaster recovery setup request, the second disaster recovery management instruction is the second disaster recovery setup instruction, the first disaster recovery management response is the first disaster recovery setup response, and the third disaster recovery management instruction is the third disaster recovery setup instruction. Figure 8 As shown, the disaster recovery setup process includes the following steps:
[0112] 1) After receiving the tenant's application instruction to establish disaster recovery management for the data cluster in Cloud 1, the management system of Cloud 1 (i.e., the first management system) issues the first disaster recovery setup instruction to the agent component of Cloud 1 (i.e., the first execution component). The first disaster recovery setup instruction carries the IP of the agent of Cloud 2.
[0113] 2) The agent component of Cloud 1 sends a first disaster recovery setup request to the agent component of Cloud 2 (i.e., the second execution component) through a network connection. The first disaster recovery setup request carries the first information required for the disaster recovery cluster of Cloud 2 to establish a disaster recovery relationship with the data cluster in Cloud 1, such as the configuration information and authentication information of the data cluster in Cloud 1.
[0114] 3) After receiving the first disaster recovery setup request sent by the agent component of cloud 1, the agent component of cloud 2 reports the first disaster recovery setup request to the corresponding topic in the Kafka of cloud 2, and sends an ACK message to the agent component of cloud 1 to indicate that the agent component of cloud 2 has received the first disaster recovery setup request.
[0115] 4) The management system of Cloud 2 (i.e., the second management system) consumes the first disaster recovery setup request from the corresponding topic in Cloud 2's Kafka. Based on the first disaster recovery setup request and the first information it carries, it decides whether to establish a disaster recovery relationship with Cloud 1's data cluster. If it determines that a disaster recovery relationship has been established with Cloud 1's data cluster, it sends a second disaster recovery setup instruction to Cloud 2's agent component and stores the first information in Cloud 2's management metadata table. Specifically, deciding whether to establish a disaster recovery relationship with Cloud 1's data cluster based on the first disaster recovery setup request and the first information it carries includes: performing authentication and authorization based on the authentication information in the first information; if authentication and authorization are successful, it determines that a disaster recovery relationship has been established with Cloud 1's data cluster.
[0116] 5) After receiving the second disaster recovery setup instruction from the management system of Cloud 2, the agent component of Cloud 2 sends a first disaster recovery setup response to the agent component of Cloud 1 via a network connection. This first response carries the second information required for the agent component of Cloud 1 to establish a disaster recovery relationship with the disaster recovery cluster of Cloud 2, such as the configuration and authentication information of the disaster recovery cluster in Cloud 2. Furthermore, the agent component of Cloud 2 executes the task flow for establishing a disaster recovery relationship with the data cluster in Cloud 1 in parallel based on the first information. The implementation process of the task flow for establishing a disaster recovery relationship with the data cluster in Cloud 1 by the agent component of Cloud 2 includes: creating a second database instance with the same specifications in Cloud 2 as the first database instance implementing the data cluster in Cloud 1, and configuring this second database instance as a backup instance of the first database instance. The second database instance is used to implement the disaster recovery cluster.
[0117] 6) After receiving the first disaster recovery setup response from the agent component of cloud 2, the agent component of cloud 1 reports the first disaster recovery setup response to the corresponding topic of Kafka in cloud 1, and sends an ACK message to the agent component of cloud 2 to indicate that the agent component of cloud 1 has received the first disaster recovery setup response.
[0118] 7) The management system of Cloud 1 consumes the first disaster recovery setup response from the corresponding topic in Cloud 1's Kafka. Based on the first disaster recovery setup response and the second information it carries, it decides whether to establish a disaster recovery relationship with the disaster recovery cluster of Cloud 2. If it determines that a disaster recovery relationship has been established with the disaster recovery cluster of Cloud 2, it sends a third disaster recovery setup instruction to the agent component of Cloud 1 and stores the second information in the management metadata table of Cloud 1. The decision on whether to establish a disaster recovery relationship with the disaster recovery cluster of Cloud 2 based on the first disaster recovery setup response and the second information it carries includes: performing authentication and authorization based on the authentication information in the second information; if authentication and authorization are successful, it determines that a disaster recovery relationship has been established with the disaster recovery cluster of Cloud 2.
[0119] 8) After receiving the third disaster recovery setup instruction from the management system of Cloud 1, the agent component of Cloud 1 executes a task flow to establish a disaster recovery relationship with the disaster recovery cluster in Cloud 2 based on the second information. This task flow includes: synchronizing data from the data cluster to the disaster recovery cluster, and configuring the second database instance as a backup instance for the first database instance. In one possible implementation, when both the data cluster in Cloud 1 and the disaster recovery cluster in Cloud 2 are distributed relational databases (such as Gauss DB), streaming replication is used to synchronize data from the distributed relational database in Cloud 1 to the distributed relational database in Cloud 2.
[0120] Once the agent components of both Cloud 1 and Cloud 2 have completed the task flow for establishing the disaster recovery relationship, the disaster recovery relationship can be established. At this point, the data cluster in Cloud 1 and the disaster recovery cluster in Cloud 2 have established a disaster recovery relationship.
[0121] Disaster recovery reinstatement refers to re-establishing the disaster recovery relationship between the data cluster and the disaster recovery cluster when the disaster recovery cluster in the second data system acts as the primary cluster and the data cluster in the first data system is operational. The implementation process is detailed in the section on disaster recovery setup and will not be elaborated upon here.
[0122] Primary / standby cluster failover refers to performing a primary / standby switchover operation on the data cluster in the first data system and the disaster recovery cluster in the second data system. The original data cluster providing services is switched to a standby cluster, and the original disaster recovery cluster is switched to the primary cluster providing services. In this case, the disaster recovery management content mentioned earlier pertains to the primary / standby cluster failover. For example, the first disaster recovery management instruction is the first primary / standby failover instruction, the disaster recovery management operation is the primary / standby failover operation, the first disaster recovery management request is the first primary / standby failover request, and the second disaster recovery management instruction is the second primary / standby failover instruction. During primary / standby cluster failover, since a disaster recovery relationship has already been established between the data cluster and the disaster recovery cluster, the first disaster recovery management request may not carry the first information. After receiving the second primary / standby failover instruction from the management system of Cloud 2, the agent component of Cloud 2 may optionally not send the first disaster recovery management response to the agent component of Cloud 1. The management system of Cloud 1 may optionally not send the third disaster recovery management instruction to the agent component of Cloud 1. For example... Figure 9 As shown, the process of switching between primary and standby clusters includes the following steps:
[0123] 1) After receiving the tenant's instruction to perform a primary / standby switch, the management system of Cloud 1 sends the first primary / standby switch instruction to the agent component of Cloud 1, triggering the primary / standby switch workflow of Cloud 1.
[0124] 2) After receiving the first primary / standby switchover instruction, the agent component of Cloud 1 sends the first primary / standby switchover request to the agent component of Cloud 2 through the network connection, and executes the workflow of switching the data cluster of Cloud 1 to the standby cluster in parallel based on the first primary / standby switchover instruction.
[0125] 3) After receiving the first primary / standby switch request sent by the agent component of cloud 1, the agent component of cloud 2 reports the first primary / standby switch request to the corresponding topic of Kafka in cloud 2, and sends ACK information to the agent component of cloud 1 to indicate that the agent component of cloud 2 has received the first primary / standby switch request.
[0126] 4) The management system of Cloud 2 consumes the first primary / standby switch request in the corresponding topic of Cloud 2's Kafka, sends the second primary / standby switch instruction to the agent component of Cloud 2, and triggers the primary / standby switch workflow of Cloud 2.
[0127] 5) After receiving the second primary / standby switch command, the agent component of Cloud 2 executes the workflow to switch the disaster recovery cluster to the primary cluster based on the second primary / standby switch command.
[0128] Once the agent components of both Cloud 1 and Cloud 2 have completed their respective primary / standby switchover workflows, the primary / standby switchover will be completed.
[0129] It should be noted that the primary / standby cluster switchover can be initiated from either the data cluster or the disaster recovery cluster. The above process is illustrated using the example where the primary / standby cluster switchover is initiated by the management system in Cloud 1, where the data cluster resides. Furthermore, there are multiple scenarios for performing a primary / standby cluster switchover. For example, when both the data cluster and the disaster recovery cluster are functioning normally, a primary / standby cluster switchover can be performed in scenarios requiring cluster management drills.
[0130] Disaster recovery deactivation refers to severing the disaster recovery relationship between the data cluster in the first data system and the disaster recovery cluster in the second data system. In this case, the disaster recovery management-related content mentioned earlier becomes the disaster recovery deactivation-related content. For example, the first disaster recovery management instruction becomes the first disaster recovery deactivation instruction, the disaster recovery management operation becomes the disaster recovery deactivation operation, the first disaster recovery management request becomes the first disaster recovery deactivation request, and the second disaster recovery management instruction becomes the second disaster recovery deactivation instruction. During disaster recovery deactivation, since a disaster recovery relationship has already been established between the data cluster and the disaster recovery cluster, the first disaster recovery management request may not carry the first information. After receiving the second primary / standby switchover instruction sent by the management system of Cloud 2, the agent component of Cloud 2 may optionally not send the first disaster recovery management response to the agent component of Cloud 1. The management system of Cloud 1 may optionally not send the third disaster recovery management instruction to the agent component of Cloud 1. For example... Figure 10 As shown, the disaster recovery and mitigation process includes the following steps:
[0131] 1) After receiving the tenant's instruction to cancel disaster recovery, the management system of Cloud 1 issues the first disaster recovery cancellation command to the agent component of Cloud 1, triggering the workflow to cancel the disaster recovery relationship between the data cluster of Cloud 1 and the disaster recovery cluster of Cloud 2.
[0132] 2) After receiving the first disaster recovery deactivation instruction, the agent component of Cloud 1 sends the first disaster recovery deactivation request to the agent component of Cloud 2 through the network connection, and executes the workflow to deactivate the disaster recovery relationship between the data cluster of Cloud 1 and the disaster recovery cluster of Cloud 2 in parallel based on the first disaster recovery deactivation instruction.
[0133] 3) After receiving the first disaster recovery cancellation request sent by the agent component of cloud 1, the agent component of cloud 2 reports the first disaster recovery cancellation request to the corresponding topic of Kafka in cloud 2, and sends an ACK message to the agent component of cloud 1 to indicate that the agent component of cloud 2 has received the first disaster recovery cancellation request.
[0134] 4) After the management system of Cloud 2 consumes the first disaster recovery cancellation request in the corresponding topic of Cloud 2's Kafka, it sends the second disaster recovery cancellation instruction to the agent component of Cloud 2, triggering the primary and backup switchover workflow of Cloud 2.
[0135] 5) After receiving the second disaster recovery deactivation instruction, the agent component of Cloud 2 executes the workflow to deactivate the disaster recovery relationship between the data cluster of Cloud 1 and the disaster recovery cluster of Cloud 2 based on the second disaster recovery deactivation instruction.
[0136] Once the agent components of Cloud 1 and Cloud 2 have completed their respective workflows for removing the disaster recovery relationship between the data cluster of Cloud 1 and the disaster recovery cluster of Cloud 2, the disaster recovery removal will be completed.
[0137] Disaster recovery promotion refers to the process of promoting a disaster recovery cluster in a second data system to become the primary cluster when the primary cluster in the first data system is unable to continue providing services due to failures or other reasons. This ensures timely business recovery. The disaster recovery promotion process does not involve interaction between the first and second execution components, but this application also applies to this scenario. In this case, the disaster recovery management-related content mentioned earlier is related to disaster recovery promotion. For example, the disaster recovery management operation is the disaster recovery promotion operation, and the second disaster recovery management instruction is the second disaster recovery promotion instruction. The implementation process is briefly explained below. Figure 11 As shown, the disaster recovery setup process includes the following steps:
[0138] 1) After receiving the instruction from the tenant's application to upgrade the disaster recovery system to the primary system, the management system of Cloud 2 sends a second instruction to the agent component of Cloud 2 to trigger the disaster recovery system upgrade workflow of Cloud 2.
[0139] 2) After receiving the second disaster recovery upgrade command, the agent component of Cloud 2 executes the workflow to upgrade the disaster recovery cluster of Cloud 2 to the master cluster based on the second disaster recovery upgrade command.
[0140] Once the agent component of Cloud 2 completes the disaster recovery upgrade workflow, the disaster recovery upgrade can be completed.
[0141] As described above, the first execution component's execution step 406 includes executing multiple tasks, at least some of which depend on the execution results of at least some of the tasks included in the second execution component's execution step 405. The following example of disaster recovery setup illustrates the implementation process where task execution has dependencies. Assume the entire cross-cloud disaster recovery setup task is simplified as follows: Cloud A needs to execute Task1 and Task2, Cloud B needs to execute Task3, and the execution of Task2 on Cloud A depends on the completion of Task3 on Cloud B. At this time, the first disaster recovery setup instruction sent by the management system of Cloud A to the execution component of Cloud A includes: the management system of Cloud A sending a first disaster recovery setup instruction instructing the execution of Task1 to the execution component of Cloud A, and the management system of Cloud A sending a first disaster recovery setup instruction instructing the execution of Task2 to the execution component of Cloud A. The second disaster recovery setup instruction sent by the management system of Cloud B to the execution component of Cloud B includes: the management system of Cloud B sending a second disaster recovery setup instruction instructing the execution of Task3 to the execution component of Cloud B. Figure 12 This is a flowchart of the implementation process. For example... Figure 12 As shown:
[0142] 1) The management system of Cloud A sends an instruction to the execution component of Cloud A to execute the first disaster recovery setup command for Task 1.
[0143] 2) The execution component of Cloud A executes Task1 based on the first disaster recovery setup instructions. After execution, it records a marker in the local file indicating that Task1 has been completed.
[0144] 3) The B Cloud management system sends an instruction to the B Cloud execution component to execute the second disaster recovery setup command for Task 3.
[0145] 4) The execution component of Cloud B executes Task3 based on the second disaster recovery setup instructions. After execution, it records a marker in a local file indicating that Task3 has been completed.
[0146] 5) The management system of Cloud A queries whether Task1 has been completed. After Task1 has been completed and before Task2 needs to be executed, it sends a query instruction to the execution component of Cloud A to query whether Task3 has been completed.
[0147] 6) Based on the query command, the execution component of cloud A sends a query request to the execution component of cloud B via network connection to query whether Task3 has been completed.
[0148] 7) Based on the query request, the execution component of Cloud B checks its local file to see if a marker indicating that Task 3 has been completed is recorded. If the marker indicating that Task 3 has been completed is recorded in the local file, a confirmation result indicating that Task 3 has been completed is sent to the execution component of Cloud A via the network connection. If the marker indicating that Task 3 has been completed is not recorded in the local file, a confirmation result indicating that Task 3 has not been completed is sent to the execution component of Cloud A via the network connection.
[0149] 8) After receiving the confirmation result sent by the execution component of cloud B, the execution component of cloud A reports the confirmation result to the corresponding topic of Kafka in cloud A, and sends an ACK message to the execution component of cloud B to indicate that the confirmation result has been received.
[0150] 9) After the management system of Cloud A consumes the confirmation result from the corresponding topic in Cloud A's Kafka, if the confirmation result indicates that Task 3 has been completed, it sends a first disaster recovery setup instruction to the execution component of Cloud A, instructing it to execute Task 2. If the confirmation result indicates that Task 3 has not been completed, it waits for a specified time and then repeats steps 5) to 8) above until the confirmation result indicates that Task 3 has been completed, or until the timeout expires and the disaster recovery setup process is terminated.
[0151] 10) The execution component of Cloud A executes Task2 based on the first disaster recovery setup instructions. After execution, it records a marker in a local file indicating that Task2 has been completed.
[0152] As can be seen from the above, in the cloud service system and disaster recovery management method provided in this application, since the first execution component can send a first disaster recovery management request to the second execution component via a network connection, and the first management system and the first execution component can communicate with each other, as can the second management system and the second execution component, instructions related to disaster recovery management between the first cloud system and the second cloud system can be transmitted through the network connection between the first execution component and the second execution component. Thus, without relying on communication between the first management system and the second management system, instructions related to disaster recovery management can be transmitted between the first cloud system and the second cloud system, enabling convenient disaster recovery management between the data cluster and the disaster recovery cluster in the database system. Furthermore, the first cloud system and the second cloud system can automatically implement disaster recovery management based on the transmitted instructions, without relying on manual execution steps, ensuring the efficiency and reliability of disaster recovery management.
[0153] Furthermore, this application provides both cross-cloud management platform disaster recovery solutions and solutions within the same cloud management platform for cross-regional disaster recovery scenarios. It achieves a unified cross-cloud and same-cloud disaster recovery architecture in cross-regional disaster recovery scenarios, ensuring the high reliability of the primary and backup dual-cluster disaster recovery capabilities through a unified architecture. Simultaneously, data synchronization between the data cluster and the disaster recovery cluster in this solution does not rely on additional data synchronization tools. The solution supports disaster recovery setup, disaster recovery recovery recovery, primary / backup cluster switching, disaster recovery promotion to primary, and disaster recovery deactivation processes, filling product gaps and reducing cross-cloud costs for customers.
[0154] It should be noted that the order of steps in the disaster recovery management method for the cloud service system provided in this application embodiment can be appropriately adjusted, and steps can also be added or removed as needed. Any variations that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application, and therefore will not be elaborated further.
[0155] The following describes an example of a virtual device in an embodiment of this application.
[0156] The above describes the disaster recovery management method of the cloud service system according to the embodiments of this application. Corresponding to the above method, the embodiments of this application also provide a cloud service system. Figure 3 This is a schematic diagram of the structure of a cloud service system provided in an embodiment of this application. Based on Figure 3 The following are several components shown, which Figure 3 The cloud service system shown is capable of performing the above. Figure 4 All or part of the operations shown. It should be understood that the cloud service system may include more additional components than those shown, or may omit some of the components shown; this application embodiment does not impose limitations in this regard. For example... Figure 3 As shown, the cloud service system includes: a first cloud system and a second cloud system. The first cloud system includes: a first management system and a first data system. The first management system manages the first data system, and the first data system provides cloud services. A first execution component is deployed in the first data system. The second cloud system includes: a second management system and a second data system. The second management system manages the second data system, and the second data system provides cloud services. A second execution component is deployed in the second data system. A network connection is established between the first execution component and the second execution component. The functions of each part are as follows:
[0157] The first management system is used to send a first disaster recovery management instruction to the first execution component. The first disaster recovery management instruction is used to instruct the first execution component to perform disaster recovery management operations for the second cloud system for the data cluster in the first data system.
[0158] The first execution component is used to send a first disaster recovery management request to the second execution component via a network connection based on the first disaster recovery management instruction. The first disaster recovery management request is used to request the execution of disaster recovery management operations for the data cluster in the second cloud system.
[0159] The second execution component is used to forward the first disaster recovery management request to the second management system.
[0160] The second management system is used to send a second disaster recovery management instruction to the second execution component based on the first disaster recovery management request.
[0161] The second execution component is also used to perform disaster recovery management operations for the data cluster in the second cloud system based on the second disaster recovery management instructions.
[0162] The first execution component is also used to perform disaster recovery management operations for the data cluster in the first cloud system based on the first disaster recovery management instructions.
[0163] In one possible implementation, the first disaster recovery management request carries first information required by the second execution component to perform disaster recovery management operations on the data cluster. The second execution component is specifically used to perform disaster recovery management operations on the data cluster in the second cloud system based on the second disaster recovery management instruction and the first information.
[0164] The second execution component is also used to respond to the second disaster recovery management instruction by sending a first disaster recovery management response to the first execution component via a network connection. The first disaster recovery management response carries second information required by the first execution component to perform disaster recovery management operations for the data cluster. Accordingly, the first execution component is specifically used to perform disaster recovery management operations for the data cluster in the first cloud system based on the second information and the first disaster recovery management instruction.
[0165] In one possible implementation, the first execution component is specifically used to execute a first task, including disaster recovery management operations for the data cluster, in the first cloud system based on a first disaster recovery management instruction. The second execution component is specifically used to execute a second task, including disaster recovery management operations for the data cluster, in the second cloud system based on a second disaster recovery management instruction. The first management system is specifically used to instruct the first execution component to confirm whether the second task has been completed before the first execution component executes a third task, including disaster recovery management operations for the data cluster. The first execution component is specifically used to confirm whether the second task has been completed to the second execution component via a network connection, and to receive a confirmation result sent by the second execution component to the first execution component via the network connection indicating whether the second task has been completed. The first execution component is specifically used to send the confirmation result to the first management system. The first management system is specifically used to send a third task execution instruction to the first execution component if the confirmation result indicates that the second task has been completed. The first execution component is specifically used to execute the third task in the first cloud system based on the third task execution instruction.
[0166] In one possible implementation, the second execution component is further configured to, after completing the second task, record a flag indicating that the second task has been completed, and, if the flag is recorded, send a confirmation result indicating that the second task has been completed to the first execution component via a network connection.
[0167] In one possible implementation, disaster recovery management operations include one or more of the following: establishing a disaster recovery cluster for the data cluster in the second data system; performing primary / standby failover operations on the data cluster and the disaster recovery cluster; and terminating the disaster recovery relationship between the data cluster and the disaster recovery cluster.
[0168] In one possible implementation, the first cloud system and the second cloud system belong to different cloud resource deployment areas managed by the same cloud management platform, or the first cloud system and the second cloud system are managed by different cloud management platforms.
[0169] In one possible implementation, a first management system is used to manage the infrastructure of a first cloud system, and a data cluster in a first data system is a first database instance created based on the infrastructure, which is used to provide database cloud services; a second management system is used to manage the infrastructure of a second cloud system, and a disaster recovery cluster in a second data system is a second database instance created based on the infrastructure, which is used to provide database cloud services.
[0170] Here, please refer to the description in the previous method embodiments for the detailed working process of the first management system, the first data system, the second management system, the second data system, the first execution component, and the second execution component. The embodiments of this application will not repeat the description here.
[0171] The following provides examples illustrating the basic hardware structures involved in the embodiments of this application.
[0172] This application also provides a computing device 1300. For example... Figure 13 As shown, the computing device 1300 includes a bus 1302, a processor 1304, a memory 1306, and a communication interface 1308. The processor 1304, the memory 1306, and the communication interface 1308 communicate with each other via the bus 1302. The computing device 1300 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 1300.
[0173] Bus 1302 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 13 The bus 1302 may be represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 1302 may include a path for transmitting information between various components of the computing device 1300 (e.g., memory 1306, processor 1304, communication interface 1308).
[0174] The processor 1304 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0175] The memory 1306 may include volatile memory, such as random access memory (RAM). The processor 1304 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0176] The memory 1306 stores executable program code, and the processor 1304 executes the executable program code to implement the functions of the aforementioned first management system, first data system, second management system, or second data system, thereby realizing the disaster recovery management method of the cloud service system. That is, the memory 1306 stores instructions for executing the disaster recovery management method of the cloud service system.
[0177] The communication interface 1308 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 1300 and other devices or communication networks.
[0178] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0179] like Figure 14 As shown, the computing device cluster includes multiple computing devices 1300. The memory 1306 of each of the multiple computing devices 1300 in the computing device cluster can store instructions for executing disaster recovery management methods for the cloud service system. Different computing devices 1300 can be used to implement the functions of a first management system, a first data system, a second management system, or a second data system, respectively. For example, the memory 1306 of each of the multiple computing devices 1300 in this computing device cluster stores partial instructions for executing disaster recovery management methods for the cloud service system. In other words, the combination of multiple computing devices 1300 can jointly execute the instructions for executing disaster recovery management methods for the cloud service system.
[0180] It should be noted that the memory 1306 in different computing devices 1300 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the cloud service system. That is, the instructions stored in the memory 1306 of different computing devices 1300 can implement the functions of the first management system, the first data system, the second management system, or the second data system.
[0181] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN), a local area network (LAN), or similar. Figure 15 One possible implementation is shown. For example... Figure 15As shown, four computing devices 1300A, 1300B, 1300C, and 1300D are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this possible implementation, the memory 1306 in computing device 1300A stores instructions for implementing the functions of the first management system. Simultaneously, the memory 1306 in computing device 1300B stores instructions for implementing the functions of the first data system. The memory 1306 in computing device 1300C stores instructions for implementing the functions of the second management system. The memory 1306 in computing device 1300D stores instructions for implementing the functions of the second data system.
[0182] It should be understood that Figure 15 The functions of computing device 1300A shown can also be performed by multiple computing devices 1300. Similarly, the functions of computing device 1300B can also be performed by multiple computing devices 1300. The functions of computing device 1300C can also be performed by multiple computing devices 1300.
[0183] This application also provides a computer program product containing instructions. The computer program product may be software or program products containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product runs on at least one computing device, it causes the at least one computing device to execute a disaster recovery management method for a cloud service system.
[0184] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform a disaster recovery management method for a cloud service system, or instruct the computing device to perform a disaster recovery management method for a cloud service system.
[0185] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0186] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the raw data and executable code involved in this application were obtained with full authorization.
[0187] In the embodiments of this application, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. The term "at least one" refers to one or more, and the term "multiple" refers to two or more, unless otherwise expressly defined.
[0188] In this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.
[0189] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.
Claims
1. A cloud service system, characterized by The cloud service system comprises: a first cloud system and a second cloud system, the first cloud system comprises: a first management system and a first data system, the first management system is used for managing the first data system, the first data system is used for providing cloud service, the first data system is deployed with a first execution component, the second cloud system comprises: a second management system and a second data system, the second management system is used for managing the second data system, the second data system is used for providing cloud service, the second data system is deployed with a second execution component, and a network connection is established between the first execution component and the second execution component; The first management system is configured to send a first disaster recovery management instruction to the first execution component, the first disaster recovery management instruction being used to instruct the first execution component to perform a disaster recovery management operation for a data cluster in the first data system on the second cloud system; The first execution component is configured to send a first disaster recovery management request to the second execution component through the network connection based on the first disaster recovery management instruction, the first disaster recovery management request being used to request to perform a disaster recovery management operation for the data cluster in the second cloud system; The second execution component is configured to forward the first disaster recovery management request to the second management system; The second management system is configured to send a second disaster recovery management instruction to the second execution component based on the first disaster recovery management request; The second execution component is further configured to perform a disaster recovery management operation for the data cluster in the second cloud system based on the second disaster recovery management instruction; The first execution component is further configured to perform a disaster recovery management operation for the data cluster in the first cloud system based on the first disaster recovery management instruction.
2. The system of claim 1, wherein The first disaster recovery management request carries first information required by the second execution component to perform a disaster recovery management operation for the data cluster, and the second execution component is specifically configured to perform a disaster recovery management operation for the data cluster in the second cloud system based on the second disaster recovery management instruction and the first information; The second execution component is further configured to send a first disaster recovery management response to the first execution component through the network connection in response to the second disaster recovery management instruction, the first disaster recovery management response carrying second information required by the first execution component to perform a disaster recovery management operation for the data cluster; The first execution component is specifically configured to perform a disaster recovery management operation for the data cluster in the first cloud system based on the second information and the first disaster recovery management instruction.
3. The system of claim 1 or 2, wherein The first execution component is specifically configured to perform a first task included in a disaster recovery management operation for the data cluster in the first cloud system based on the first disaster recovery management instruction; The second execution component is specifically configured to perform a second task included in a disaster recovery management operation for the data cluster in the second cloud system based on the second disaster recovery management instruction; The first management system is specifically configured to instruct the first execution component to confirm whether the second task is executed completely before the first execution component executes a third task included in a disaster recovery management operation for the data cluster; The first execution component is specifically configured to confirm, through the network connection, whether the second task is executed completely to the second execution component, and receive the confirmation result sent by the second execution component to the first execution component through the network connection, the confirmation result being used to indicate whether the second task is executed completely; The first execution component is specifically configured to send the confirmation result to the first management system; The first management system is specifically configured to send a third task execution instruction to the first execution component in a case where the confirmation result indicates that the second task is executed completely; The first execution component is specifically configured to execute the third task in the first cloud system based on the third task execution instruction.
4. The system of claim 3, wherein The second execution component is further configured to record a mark indicating that the second task is completed after the second task is completed, and send, to the first execution component through the network connection, a confirmation result indicating that the second task is executed completely in a case where the mark is recorded.
5. The system of any one of claims 1 to 4, wherein, The disaster recovery management operation includes one or more of the following: establishing a disaster recovery cluster of the data cluster in the second data system; performing a master-slave switchover operation on the data cluster and the disaster recovery cluster; releasing the disaster recovery relationship between the data cluster and the disaster recovery cluster.
6. The system of any one of claims 1 to 5, wherein, The first cloud system and the second cloud system belong to different cloud resource deployment regions managed by a same cloud management platform, or the first cloud system and the second cloud system are respectively managed by different cloud management platforms.
7. The system of any one of claims 1 to 6, wherein The first management system is configured to manage infrastructure possessed by the first cloud system, the data cluster in the first data system being a first database instance created based on the infrastructure, and the first database instance being configured to provide a database cloud service; The second management system is configured to manage infrastructure possessed by the second cloud system, the disaster recovery cluster in the second data system being a second database instance created based on the infrastructure, and the second database instance being configured to provide a database cloud service.
8. A disaster recovery management method of a cloud service system, characterized by, The cloud service system includes a first cloud system and a second cloud system, the first cloud system includes a first management system and a first data system, the first management system is configured to manage the first data system, the first data system is configured to provide a cloud service, and the first data system is deployed with a first execution component, the second cloud system includes a second management system and a second data system, the second management system is configured to manage the second data system, the second data system is configured to provide a cloud service, and the second data system is deployed with a second execution component, a network connection is established between the first execution component and the second execution component, and the method includes: The first management system sends a first disaster recovery management instruction to the first execution component, where the first disaster recovery management instruction is used to instruct the first execution component to perform a disaster recovery management operation for a data cluster in the first data system to the second cloud system; The first execution component sends a first disaster recovery management request to the second execution component based on the first disaster recovery management instruction through the network connection, where the first disaster recovery management request is used to request to perform a disaster recovery management operation for the data cluster in the second cloud system; The second execution component forwards the first disaster recovery management request to the second management system; The second management system sends a second disaster recovery management instruction to the second execution component based on the first disaster recovery management request; The second execution component performs a disaster recovery management operation for the data cluster in the second cloud system based on the second disaster recovery management instruction; The first execution component performs a disaster recovery management operation for the data cluster in the first cloud system based on the first disaster recovery management instruction.
9. The method of claim 8, wherein The first disaster recovery management request carries first information required by the second execution component to perform a disaster recovery management operation for the data cluster, and the second execution component performs a disaster recovery management operation for the data cluster in the second cloud system based on the second disaster recovery management instruction includes: The second execution component performs a disaster recovery management operation for the data cluster in the second cloud system based on the second disaster recovery management instruction and the first information; The method further includes: The second execution component sends a first disaster recovery management response to the first execution component through the network connection in response to the second disaster recovery management instruction, where the first disaster recovery management response carries second information required by the first execution component to perform a disaster recovery management operation for the data cluster; The first execution component performs a disaster recovery management operation for the data cluster in the first cloud system based on the first disaster recovery management instruction includes: The first execution component performs a disaster recovery management operation for the data cluster in the first cloud system based on the second information and the first disaster recovery management instruction.
10. The method of claim 8 or 9, wherein The second execution component performs a disaster recovery management operation for the data cluster in the second cloud system based on the second disaster recovery management instruction includes: The second execution component performs a second task included in the disaster recovery management operation for the data cluster in the second cloud system based on the second disaster recovery management instruction; The first execution component performs a disaster recovery management operation for the data cluster in the first cloud system based on the first disaster recovery management instruction includes: The first execution component performs a first task included in the disaster recovery management operation for the data cluster in the first cloud system based on the first disaster recovery management instruction; The first management system instructs the first execution component to confirm whether the second task is executed before the first execution component executes a third task included in a disaster recovery management operation for the data cluster; The first execution component confirms whether the second task is executed by the second execution component through the network connection, and receives the confirmation result sent by the second execution component to the first execution component through the network connection, and sends the confirmation result to the first management system; The first management system sends a third task execution instruction to the first execution component if the confirmation result indicates that the second task is executed; The first execution component executes the third task in the first cloud system based on the third task execution instruction.
11. The method of claim 10, wherein, The method further comprises: The second execution component records a mark indicating that the second task is completed after completing the second task, and sends a confirmation result indicating that the second task is executed to the first execution component through the network connection if the mark is recorded.
12. The method of any one of claims 8 to 11, wherein, The disaster recovery management operation includes one or more of the following: Establishing a disaster recovery cluster of the data cluster in the second data system; Performing a master-slave switchover operation on the data cluster and the disaster recovery cluster; Releasing the disaster recovery relationship between the data cluster and the disaster recovery cluster.
13. The method of any one of claims 8 to 12, wherein, The first cloud system and the second cloud system belong to different cloud resource deployment areas managed by the same cloud management platform, or the first cloud system and the second cloud system are managed by different cloud management platforms.
14. The method of any one of claims 8 to 13, wherein: The first management system is configured to manage infrastructure of the first cloud system, and the data cluster in the first data system is a first database instance created based on the infrastructure, and the first database instance is configured to provide a database cloud service; The second management system is configured to manage infrastructure of the second cloud system, and the disaster recovery cluster in the second data system is a second database instance created based on the infrastructure, and the second database instance is configured to provide a database cloud service.
15. A cluster of computing devices, characterized in that, The computing device cluster comprises a plurality of computing devices, the plurality of computing devices comprising a plurality of processors and a plurality of memories, the plurality of memories storing program instructions, and the plurality of processors executing the program instructions to cause the computing device cluster to perform the method of any one of claims 8 to 14.
16. A computer readable storage medium characterized by: The program instructions, when executed on a computing device, cause the computing device to perform the method of any one of claims 8 to 14.
17. A computer program product comprising instructions, characterized in that, The instructions, when executed on a computing device cluster, cause the computing device cluster to perform the method of any one of claims 8 to 14.