Resource specification change fault processing method and computing device
By introducing change processes and processing processes into the computing service, we automatically monitor resource specification change operations and execute fault processing strategies, and solve the long recovery time problems caused by resource specification change failures, improving fault processing efficiency and business continuity.
Patent Information
- Application Number
- CN202410008202.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-03
- Publication Date
- 2025-07-04
AI Technical Summary
When a computing service fails during resource specification changes, the failure recovery time is long, affecting business continuity and user experience.
By introducing change processes and processing processes into the computing service, we automatically monitor the execution results of resource specification changes, and execute corresponding fault handling strategies when a failure occurs, including releasing resources, data snapshots, address switching and other operations.
It realizes automated perception and rapid processing of resource specification changes failures, reduces fault recovery time, and improves business continuity and user experience.
Smart Images

Figure CN120256172A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and in particular, to a method for handling resource specification change failures and a computing device. Background Art
[0002] During the life cycle of a computing service, if the business volume of the computing service changes, it is usually necessary to adjust the resource specification of the computing service. For example, when the business volume increases, the resource specification can be upgraded to support the business; when the business volume decreases, the resource specification can be downgraded to reduce costs. During the process of changing the resource specification of a computing service, failures may occur in the resource specification change due to some unpredictable reasons.
[0003] In the related art, if a failure occurs during the process of changing the resource specification, the computing service will issue an alarm. After receiving the alarm, the operation and maintenance personnel will perform fault troubleshooting, identify the fault, and then handle the fault to restore the normal operation of the computing service. This method of the computing service issuing an alarm and the operation and maintenance personnel manually handling the fault will prolong the fault recovery time of the computing service. If the fault recovery time is long, the business of the computing service will be damaged due to untimely processing, affecting the user experience. Summary of the Invention
[0004] This application provides a method for handling resource specification change failures and a computing device, which can automatically recover the failures that occur during the resource specification change of a computing service and solve the problem of a long fault recovery time.
[0005] In a first aspect, this application provides a method for handling resource specification change failures. This method can be applied to a computing node on which a computing service is deployed. The method includes: obtaining an execution result of a change operation, where the change operation is used to change the resource specification of the computing service; determining whether the change operation has failed according to the execution result; and in a case where it is determined that the change operation has failed, executing a fault handling strategy corresponding to the change operation.
[0006] In the above solution, the computing service can sense whether a failure occurs during the resource specification change based on the execution result of the change operation, and in the case of a failure, execute the corresponding fault handling strategy, which can enable the computing service to automatically sense and handle the fault, improve the fault handling efficiency, and solve the problem of a long fault recovery time.
[0007] In a possible implementation manner of the first aspect, the computing service includes a first program, and the method includes: executing the first program to execute the method.
[0008] In the above solution, the first program may be the program code of the change process in the computing service. When the change process runs on a computing node, the computing node executes this first program and executes this method. In this way, during the process of the change process of the computing service to implement resource specification changes, it can monitor each operation step in the change process, sense whether each operation fails, and execute corresponding processing strategies.
[0009] In a possible implementation manner of the first aspect, the computing service includes a first program and a second program, and the method further includes: executing the first program to execute the change operation, obtain the execution result of the change operation, and determine whether the change operation fails according to the execution result; in the case where it is determined that the change operation fails, execute the second program to execute the fault handling strategy corresponding to the change operation.
[0010] In the above solution, the first program may be the program code of the change process in the computing service, and the second program may be the program code of the processing process in the computing service. When the change process runs on a computing node, the computing node executes this first program. In the case where it is determined that the change operation fails, a processing process runs on the computing node, and the computing node executes the second program to execute the fault handling strategy corresponding to the change operation.
[0011] In a possible implementation manner of the first aspect, the first program includes a data interface. In the case where it is determined that the change operation fails, the method further includes: sending the execution result to the second program through the data interface; executing the second program to determine the fault handling strategy corresponding to the execution of the change operation according to the execution result.
[0012] In the above solution, during the process of the computing node implementing the fault handling method by running the change process and the processing process, the two processes can communicate through the data interface, so that information indicating a failure can be transmitted.
[0013] In a possible implementation manner of the first aspect, the change operation includes a resource allocation operation, a data snapshot mounting operation, a data consistency check operation, an address switching operation, or a management layer data update operation.
[0014] Among them, the fault handling strategy corresponding to the resource allocation operation includes: releasing the resources of the target specification applied for by the resource allocation operation, and releasing the conflict lock added by the resource allocation operation to the computing service.
[0015] Among them, the fault handling strategy corresponding to the data snapshot mounting operation includes: releasing the data snapshot created by the data snapshot mounting operation, releasing the resources of the target specification, and releasing the conflict lock.
[0016] Among them, the fault handling strategy corresponding to the data consistency check operation includes: releasing the backup data created by the data consistency check operation, releasing the data snapshot, releasing the resources of the target specification, and releasing the conflict lock.
[0017] Among them, the fault handling strategy corresponding to the address switching operation includes: switching back the address of the computing service, releasing the backup data, releasing the data snapshot, releasing the resources of the target specification, and releasing the conflict lock.
[0018] Among them, the fault handling strategy corresponding to the management layer data update operation includes: updating the resource specification of the computing service stored in the management node corresponding to the computing service again to the resource specification of the computing service after the resource specification change.
[0019] In a possible implementation manner of the first aspect, the computing service includes a cloud database.
[0020] In a second aspect, the present application provides a resource specification change fault handling device. The device is applied to a computing node on which a computing service is deployed. The device includes: a sensing module and a processing module.
[0021] Among them, the sensing module is used to obtain the execution result of a change operation, where the change operation is used to change the resource specification of the computing service, and determine whether the change operation fails according to the execution result.
[0022] Among them, the processing module is used to execute the fault handling strategy corresponding to the change operation in the case where it is determined that the change operation fails.
[0023] In a possible implementation manner of the second aspect, the computing service includes a first program.
[0024] Among them, the sensing module is used to execute the first program to obtain the execution result of the change operation and determine whether the change operation fails according to the execution result. The processing module is used to execute the first program to execute the fault handling strategy corresponding to the change operation in the case where it is determined that the change operation fails.
[0025] In a possible implementation manner of the second aspect, the computing service includes a first program and a second program.
[0026] Among them, the sensing module is used to execute the first program to obtain the execution result of the change operation and determine whether the change operation fails according to the execution result. The processing module is used to execute the second program to execute the fault handling strategy corresponding to the change operation in the case where it is determined that the change operation fails.
[0027] In a possible implementation manner of the second aspect, the first program includes a data interface. In the case where it is determined that the change operation fails, the sensing module is further used to: send the execution result to the second program through the data interface. The processing module is further used to execute the second program to determine the fault handling strategy corresponding to the execution of the change operation according to the execution result.
[0028] In a possible implementation manner of the second aspect, the change operation includes a resource allocation operation, a data snapshot mounting operation, a data consistency check operation, an address switching operation, or a management layer data update operation.
[0029] Among them, the fault handling strategy corresponding to the resource allocation operation includes: releasing the resources of the target specification applied for by the resource allocation operation, and releasing the conflict lock added to the computing service by the resource allocation operation.
[0030] Among them, the fault handling strategy corresponding to the data snapshot mounting operation includes: releasing the data snapshot created by the data snapshot mounting operation, releasing the resources of the target specification, and releasing the conflict lock.
[0031] Among them, the fault handling strategy corresponding to the data consistency check operation includes: releasing the backup data created by the data consistency check operation, releasing the data snapshot, releasing the resources of the target specification, and releasing the conflict lock.
[0032] Among them, the fault handling strategy corresponding to the address switching operation includes: switching back the address of the computing service, releasing the backup data, releasing the data snapshot, releasing the resources of the target specification, and releasing the conflict lock.
[0033] Among them, the fault handling strategy corresponding to the management layer data update operation includes: re-updating the resource specification of the computing service stored in the management node corresponding to the computing service to the resource specification of the computing service after the resource specification change.
[0034] In a possible implementation manner of the second aspect, the computing service includes a database service.
[0035] In a third aspect, the present application provides a processor. The processor is applied to a computing device in which a computing service is deployed. The processor is used to execute the resource specification change fault handling method provided in the first aspect or any possible implementation manner of the first aspect.
[0036] In a fourth aspect, the present application provides a computing device. The computing device includes a processor and a memory. The processor is used to execute a computer program stored in the memory to implement the resource specification change fault handling method provided in the foregoing first aspect or any possible implementation manner of the first aspect.
[0037] In a fifth aspect, the present application provides a computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the resource specification change fault handling method provided in the first aspect or any possible implementation manner of the first aspect.
[0038] In a sixth aspect, the present application provides a computer program product containing instructions. When the computer program product runs on a computer, the computer is caused to execute the resource specification change fault handling method provided in the first aspect or any possible implementation manner of the first aspect.
[0039] Any of the foregoing provided devices, computer storage media, or computer program products is used to execute the method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding solutions in the corresponding methods provided above, which will not be elaborated here. Description of the Drawings
[0040] Figure 1 is a schematic diagram of resource specification change for a cloud database provided by an embodiment of the present application;
[0041] Figure 2 is a flowchart of a resource specification change fault handling method provided by an embodiment of the present application;
[0042] Figure 3 is a schematic diagram of a change operation included in resource specification change for a cloud database provided by an embodiment of the present application;
[0043] Figure 4a is a schematic diagram of a resource allocation operation for a cloud database provided by an embodiment of the present application;
[0044] Figure 4b is a schematic diagram of a fault handling strategy corresponding to a resource allocation operation provided by an embodiment of the present application;
[0045] Figure 5a is a schematic diagram of a data snapshot mounting operation for a cloud database provided by an embodiment of the present application;
[0046] Figure 5b It is a schematic diagram of a fault handling strategy corresponding to a data snapshot mounting operation provided by an embodiment of the present application;
[0047] Figure 6a It is a schematic diagram of an operation for performing data consistency verification on a cloud database provided by an embodiment of the present application;
[0048] Figure 6b It is a schematic diagram of a fault handling strategy corresponding to a data consistency verification operation provided by an embodiment of the present application;
[0049] Figure 7a It is a schematic diagram of an operation for performing VIP switching on a cloud database provided by an embodiment of the present application;
[0050] Figure 7b It is a schematic diagram of a fault handling strategy corresponding to a VIP switching operation provided by an embodiment of the present application;
[0051] Figure 8a It is a schematic diagram of an operation for performing management layer data update on a cloud database provided by an embodiment of the present application;
[0052] Figure 8b It is a schematic diagram of a fault handling strategy corresponding to a management layer data update operation provided by an embodiment of the present application;
[0053] Figure 9 It is a flowchart of another method for handling resource specification change faults provided by an embodiment of the present application;
[0054] Figure 10 It is a schematic structural diagram of a device for handling resource specification change faults provided by an embodiment of the present application;
[0055] Figure 11 It is a schematic structural diagram of a computing device provided by an embodiment of the present application. Detailed implementation manners
[0056] In order to make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.
[0057] In the description of the embodiments of the present application, words such as "exemplary", "for example", or "for illustration" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary", "for example", or "for illustration" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example", or "for illustration" is intended to present relevant concepts in a specific manner.
[0058] In the description of the embodiments of the present application, the term "and / or" is merely a relationship describing associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, B exists alone, and both A and B exist simultaneously. Additionally, unless otherwise specified, the meaning of the term "plural" refers to two or more. For example, a plurality of systems refers to two or more systems, and a plurality of screen terminals refers to two or more screen terminals.
[0059] Furthermore, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0060] Before introducing the embodiments of the present application, the following first introduces the nouns that appear in the embodiments of the present application.
[0061] A computing node (hereinafter referred to as a node) can be a physical computing device (such as a server), or can be a virtual computer system (such as a virtual machine) determined through virtualization technology on a physical computing device, etc.
[0062] A computing service is a software program deployed on a computing node for providing specific services to users. Computing services can include but are not limited to database services, cloud services, or cloud databases.
[0063] A cloud database is a database service deployed in a cloud environment and provided to users in the form of a cloud service. A cloud database can consist of multiple database instances, and the database instances can be deployed in one or more computing nodes. To improve data security in a multi-user environment and protect user privacy, the cloud database can use virtualization technology to divide these computing nodes into multiple virtual networks, and each virtual network can run at least one independent database instance. Different database instances provide database services to different users, and the data between the database instances deployed in different virtual networks is isolated from each other. Moreover, in a cloud database, a load balancer (or a proxy server, and other components) can be used to manage and distribute user requests. Communication can be carried out between the load balancer and the nodes in the virtual network, and between the nodes of each virtual network, based on the virtual internet protocol (VIP). Each load balancer and each node in the virtual network correspond to an independent VIP address. In this way, users can communicate with the load balancer of the cloud database through an application (APP) to perform data storage, query, and management operations in the cloud database. That is to say, the operation requests of users for the cloud database first reach the load balancer of the cloud database. The load balancer will forward the operation requests to the computing nodes in the corresponding virtual network for execution according to the user's identity or other parameters. In this way, users only need to know the domain name of the cloud database to communicate with the cloud database, without needing to be aware of the VIP addresses of the computing nodes.
[0064] A resource specification is the resources on which a computing service provides corresponding services, and specifically can include computing resources, storage resources, network resources, etc. Computing resources can include the number of central processing units (CPUs) or the number of CPU cores occupying the computing nodes. The size of the computing resources determines the amount of business data that the computing service can process. Storage resources can include the size of the memory occupying the computing nodes. The size of the storage resources determines the amount of business data that the computing service can store. Network resources can include the network bandwidth occupying the computing nodes. The size of the network resources determines the data transmission ability between the computing service and the external environment.
[0065] Resource specification change means upgrading or downgrading the resource specification of a computing service according to the specific business volume of the computing service. Taking a cloud database as an example, as Figure 1 shown, the cloud database is deployed on Node 1 with a resource specification of 1U2G (1 CPU, 2G storage space). If the business volume of the cloud database increases, the resource specification of the cloud database is changed. As Figure 1As shown in the figure, after the resource specification is changed, the cloud database is deployed in Node 2, and the resource specification is changed to 2U4G (2 CPUs, 4G storage space). After the resource specification is changed, the cloud database can handle more traffic, thereby providing better services to users.
[0066] The infrastructure service layer node is a node used to manage computing resources, storage resources, network resources, and other basic resources in the cloud environment. When a computing service needs to be created or the resource specification of a computing service needs to be changed, resources can be applied for from the IAAS layer node, and the IAAS layer allocates resources to create a computing service or change the resource specification of a computing service.
[0067] A data snapshot is a copy of a data set at a certain point in time. In the storage field, a data snapshot is a data protection technology. When data is damaged for some reason, the data can be restored to the data at a certain point in time through the snapshot technology.
[0068] The management layer is used to manage the metadata of computing services and storage computing services. The metadata includes the resource specifications of computing services.
[0069] During the process of changing the resource specification of a computing service, the resource specification change may fail due to various abnormal problems. In a related technology, the change process of the computing service monitors the abnormal problems of the computing service during the resource specification change process. If an abnormality occurs, an alarm is issued; after receiving the alarm, the operation and maintenance personnel determine the fault by analyzing the alarm, and then manually process the fault.
[0070] In order to avoid the resource specification change failure from affecting the business of the computing service, it is usually necessary to process the failure in a timely manner. In the solution implemented by the related technology, when the change process of the computing service monitors an abnormality, an alarm is issued, and then the operation and maintenance personnel identify the fault based on the alarm and manually process the fault, which will prolong the fault handling time and the timeliness is poor. In other words, this solution cannot ensure that the resource specification change failure is processed quickly, and in severe cases, it will affect the business of the computing service, thereby reducing the user experience.
[0071] Therefore, the embodiment of the present application provides a method for processing resource specification change failures, which can solve the above problems.
[0072] In the method provided by the embodiments of the present application, after performing a change operation in the resource specification change of the computing service, the computing node deploying the computing service determines whether a failure occurs according to the execution result of the change operation, and in the case where a failure is determined to occur, executes a failure handling strategy corresponding to the change operation. This method can achieve automatic perception of resource specification change failures and automatic handling of failures. Compared with the related art above where operation and maintenance personnel determine failures and handle them manually, the method provided by the embodiments of the present application can improve the speed of failure handling, avoid the business being affected due to a long failure handling time, thereby improving the business security of the computing service and improving the user experience.
[0073] Taking a cloud database as an example, the following combines Figure 1 and Figure 2 to introduce in detail the resource specification change failure handling method provided by the embodiments of the present application.
[0074] Figure 2 FIG. is a flowchart of a resource specification change failure handling method provided by the embodiments of the present application. This method can be applied to a cloud database in a computing node and is used to automatically perceive and handle resource specification change failures of the cloud database. Among them, the cloud database includes a change process and a processing process, Figure 2 The method shown can be executed by the change process and the processing process.
[0075] As Figure 2 shown, this method may include the following S201-S204.
[0076] S201, the change process executes a change operation and determines the execution result of the change operation.
[0077] In this embodiment, in the initial state, the cloud database is deployed in nodes 1 and 3 as shown in Figure 4a Node 1 is the primary node of the cloud database, and node 3 is the standby node of the cloud database. When it is necessary to change the resource specification of the cloud database, the target specification of the cloud database is clarified, and then the change process of the cloud database is started, and the computing node executes the change process. When the computing node executes the change process, it executes the resource specification change process of the cloud database and changes the resource specification of the cloud database to the target specification.
[0078] In this embodiment, the resource specification change process of the cloud database can be divided into five change operations executed in sequence as shown in Figure 3 The five change operations are respectively a resource allocation operation, a data snapshot mounting operation, a data consistency verification operation, a virtual Internet protocol VIP address switching operation, and a management layer data update operation.
[0079] The following introduces the five change operations respectively. It should be noted thatFigure 3 The five change operations shown are only an exemplary division of the resource specification change process in the embodiments of the present application and do not constitute a limitation to the present application. In other embodiments, the resource specification change process can be divided into other types of operations.
[0080] For the resource allocation operation, when the computing node starts to execute the change process, the resource allocation operation is executed.
[0081] Specifically, as Figure 4a shown, the change process adds a conflict lock to the cloud database to prevent the cloud database from responding to other change requests, and then sends a resource request to the IAAS layer node. This resource request is used to apply for resources of the target specification from the IAAS layer node; after receiving the resource request, the IAAS layer node allocates resources of the target specification.
[0082] After the IAAS layer node successfully allocates resources, it can send a first message indicating successful resource allocation to the computing node; and when it determines that the resources of the cloud platform cannot support the resource specification change, it can send a second message indicating failed resource allocation to the computing node. Thus, the change process can determine the execution result of the resource allocation operation according to the message received from the IAAS layer. Specifically, when receiving the first message, it is determined that the execution result of the change process is successful; when receiving the second message, it is determined that the execution result of the change process fails. Among them, the first message may also include the address information of node 2 and node 4. Node 2 is the primary node of the cloud database after the resource specification change, and node 4 is the standby node of the cloud database after the resource specification change. The address information may include the VIP addresses of node 2 and node 4.
[0083] For the data snapshot mounting operation, the change process executes the data snapshot mounting operation in the case where the execution result of the resource allocation operation is determined to be successful.
[0084] Specifically, as Figure 5a shown, the change process takes a snapshot of the data set of the cloud database at the current time to obtain a data snapshot, and then mounts the data snapshot to node 2 and node 4. Among them, the change process can send the data snapshot to node 2 and node 4 based on the VIP addresses of node 2 and node 4. After receiving the data snapshot, node 2 and node 4 store the data snapshot in their corresponding storage spaces.
[0085] When Node 2 and Node 4 successfully store the data snapshot, they can send a third message indicating successful mounting to Node 1, and when it is determined that the data snapshot cannot be successfully stored, they can send a fourth message indicating failed mounting to Node 1. Thus, the change process can determine the execution result of the data snapshot mounting operation based on the messages received from Node 2 and Node 4. If the third message is received from both Node 2 and Node 4, it is determined that the execution result of the data snapshot mounting operation is successful; otherwise, it is failed.
[0086] For the data consistency check operation, the change process performs the data consistency check operation in the case where it is determined that the execution result of the data snapshot mounting operation is successful.
[0087] Specifically, since Node 1 does not stop during snapshot taking and snapshot mounting and still provides the database service to users, this may cause differences between the data snapshot and the data in the cloud database after snapshot mounting. Therefore, as Figure 6a shown, the change process can establish a streaming replication relationship between Node 1 and Node 2, and between Node 1 and Node 4, and use Node 2 and Node 4 as standby nodes of Node 1 to back up the data of Node 1. Moreover, the change process only executes the next change operation when it determines that the data stored in Node 2 and the data stored in Node 4 are both consistent with the data stored in Node 1, which can prevent the loss of data in the cloud database. Among them, after establishing the streaming replication relationship, the change process can back up the data stored in the cloud database from the time of snapshot taking to the current time to Node 2 and Node 4. Node 2 and Node 4 store the backup data in their respective storage spaces. Node 2 and Node 4 can also send messages indicating whether the backup data is successfully stored or failed to Node 1, and Node 1 determines the execution result of the data consistency check operation based on the messages from Node 2 and Node 4. The specific process can refer to the foregoing explanation of the data snapshot mounting operation and will not be elaborated here.
[0088] For the VIP address switching operation, the change process performs the VIP address switching operation in the case where it is determined that the execution result of the data consistency check operation is successful.
[0089] Specifically, after the change process successfully executes the data consistency check operation, all the data of the cloud database deployed on Node 1 is migrated to Node 2. Therefore, as Figure 7a shown, the change process can update the VIP address of the node providing the database service to users to the VIP address of Node 2, and Node 2 replaces Node 1 to provide the database service to users. Among them, before switching the VIP address, the VIP address of the node providing the database service to users is the VIP address of Node 1. After switching the VIP address, the operation requests sent by users through the application APP can reach Node 2 via the load balancer of the cloud database, and Node 2 executes the operation requests.
[0090] For the management layer data update operation, in the case where the change process determines that the execution result of the VIP address switching operation is successful, the management layer data update operation is executed.
[0091] Specifically, as Figure 8a shown, before switching the VIP address, the change process originally ran on Node 1. After switching the VIP address, the change process runs on Node 2. Therefore, Node 2 can update the specification information of the cloud database stored in the management layer database where the management node of the cloud database is deployed to the target specification.
[0092] S202. The change process determines whether the change operation fails according to the execution result.
[0093] In this embodiment, if the execution result of the change operation indicates that the change operation is successful, the change process determines that the change operation has not failed; if the execution result of the change operation indicates that the change operation fails, the change process determines that the change operation has failed.
[0094] In the case where no failure occurs, Node 1 continues to execute the change process to execute the next operation of the change operation. For example, in the case where the execution result of the resource allocation operation indicates that the resource allocation operation is successful, the change process executes the data snapshot mounting operation.
[0095] In the case where a failure occurs, Node 1 executes S203.
[0096] S203. In the case where it is determined that the change operation has failed, the change process sends the execution result of the change operation to the processing process.
[0097] In this embodiment, the change process can communicate with the processing process through the data interface. The change process can send the execution results of each change operation to the processing process through the data interface in the case where the execution result of the change operation indicates that the change operation fails. Then, Node 1 executes the processing process, and the processing process obtains the execution results of each change operation from the data interface.
[0098] S204. The processing process executes the fault handling strategy corresponding to the change operation.
[0099] In this embodiment, after receiving the execution result of the change operation, the processing process executes the fault handling strategy corresponding to the change operation.
[0100] For the resource allocation operation, as Figure 4bAs shown, after receiving the execution result of the change operation, the processing process releases the resources of the target specification and the conflict lock. Specifically, the processing process can send a release request to the IAAS layer node to release the resources of the target specification, that is, release the resources corresponding to node 2 and node 4. After receiving the release request, the IAAS layer node reclaims the resources of the target specification and then sends a message indicating successful release to node 1. After receiving the message indicating successful release, the processing process releases the conflict lock of the cloud database.
[0101] For the data snapshot mounting operation, as Figure 5b shown, after receiving the execution result of the data snapshot mounting operation, the processing process releases the data snapshot, releases the resources of the target specification, and releases the conflict lock. Releasing the data snapshot includes deleting the data snapshot stored on node 1 and notifying node 2 and node 4 to delete the stored data snapshot. After releasing the data snapshot, the resources of the target specification are released and the conflict lock is released. The specific process refers to the introduction of the fault handling strategy corresponding to the above resource allocation operation.
[0102] For the data consistency check operation, as Figure 6b shown, after receiving the execution result of the data snapshot mounting operation, the processing process releases the backup data, releases the data snapshot, releases the resources of the target specification, and releases the conflict lock. Releasing the backup data includes notifying node 2 and node 4 to delete the stored backup data. The specific process of releasing the data snapshot, releasing the resources of the target specification, and releasing the conflict lock can refer to the introduction of the fault handling strategy corresponding to the above data snapshot mounting operation.
[0103] For the VIP address switching operation, as Figure 7b shown, after receiving the execution result of the VIP switching operation, the processing process reverts the VIP address, releases the backup data, releases the data snapshot, releases the resources of the target specification, and releases the conflict lock.
[0104] For the management layer data update operation, as Figure 8b shown, after receiving the execution result of the management layer data update operation, at this time node 2 has started to provide services to users. The failure of the management layer data update operation will not affect the users' services. Therefore, it is only necessary to update the specification information of the cloud database stored in the management layer database again.
[0105] In the above embodiments, by setting a processing process in the computing service to monitor the change process, it is possible to automatically sense whether a failure occurs in the resource specification change process, and in the case of a failure, execute the corresponding fault handling strategy in a timely manner, achieving the purpose of quickly processing the fault. In this embodiment, setting the processing process and the change process to jointly sense and handle faults can avoid making major changes to the change process and reduce the program development cost.
[0106] Based on Figure 2 the method embodiments shown above, the embodiments of the present application further provide a method for handling resource specification change failures. This method can be applied to a cloud database, and the change process automatically senses and handles resource specification change failures in the cloud database.
[0107] Figure 9 is a flowchart of a method for handling resource specification change failures provided by the embodiments of the present application. As Figure 9 shown, this method may include the following S901 - S903.
[0108] S901, perform a change operation and determine the execution result of this change operation.
[0109] S902, determine whether a failure occurs in this change operation according to the execution result.
[0110] S903, in the case where it is determined that a failure occurs in this change operation, execute the failure handling strategy corresponding to this change operation.
[0111] In this embodiment, the specific processes of S901 - S903 may respectively refer to the descriptions of S201, S203, and S204 in the method embodiments shown above, and will not be elaborated here. Figure 2 In the above embodiments, during the process of performing resource specification changes in the cloud database, the change process automatically senses whether a failure occurs in the resource specification change process, and in the case of a failure, promptly executes the corresponding failure handling strategy, achieving the purpose of quickly handling failures. Compared with
[0112] the embodiments shown, having the change process sense and handle failures can reduce the communication overhead between processes, thereby reducing the resource consumption of computing nodes. Figure 2 Based on
[0113] Based on Figure 2 and Figure 9 the method embodiments shown above, the embodiments of the present application further provide a device for handling resource specification change failures.
[0114] Figure 10 is a schematic structural diagram of a failure handling device 1000 provided by the embodiments of the present application. This failure handling device 1000 includes a sensing module 1001 and a processing module 1002.
[0115] Among them, the sensing module 1001 is used to determine whether a failure occurs in the change operation according to the execution result of the change operation.
[0116] Among them, the processing module 1002 is used to execute the failure handling strategy corresponding to this change operation in the case where it is determined that a failure occurs.
[0117] It should be noted that Figure 10 When the resource specification change fault handling device 1000 provided by the illustrated embodiment executes the resource specification change fault handling method, only the division of the above functional modules is used as an example for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the resource specification change fault handling device provided by the above embodiment and Figure 2 or Figure 9 The resource specification change fault handling method embodiment shown belongs to the same concept. For the specific implementation process, please refer to the method embodiment and will not be elaborated here.
[0118] Figure 11 FIG. 11 is a schematic hardware structure diagram of a computing device 1100 provided by an embodiment of the present application.
[0119] The above cloud database is deployed in the computing device 1100. Refer to Figure 11 , the computing device 1100 includes a processor 1101, a memory 1102, a communication interface 1103, and a bus 1104. The processor 1101, the memory 1102, and the communication interface 1103 are connected to each other through the bus 1104. The processor 1101, the memory 1102, and the communication interface 1103 can also be connected in other connection manners except the bus 1104.
[0120] Among them, the processor 1101 can be a general-purpose processor, and the general-purpose processor can be a processor that executes specific steps and / or operations by reading and executing the content stored in a memory (such as the memory 1102). For example, the general-purpose processor can be a central processing unit (CPU). The processor 1101 can include at least one circuit to execute Figure 2 or Figure 9 All or part of the steps of the resource specification change fault handling method provided by the illustrated embodiment.
[0121] Among them, the memory 1102 can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, optical memory, hard disk, etc. The memory 1102 is specifically used to store the computer programs corresponding to the sensing module and the processing module. When the processor 1101 executes the computer programs corresponding to the sensing module and the processing module, the processor 1101 implements the above Figure 2 or Figure 9 steps in the method embodiments shown.
[0122] Among them, the communication interface 1103 includes interfaces such as input / output (I / O) interfaces, physical interfaces, and logical interfaces for implementing the interconnection of components inside the computing device 1100, as well as interfaces for implementing the interconnection of the computing device 1100 with other devices (such as other computing devices or user devices). The physical interface can be an Ethernet interface, a fiber optic interface, an ATM interface, etc.
[0123] Among them, the bus 1104 can be of any type and is a communication bus for implementing the interconnection of the processor 1101, the memory 1102, and the communication interface 1103, such as a system bus.
[0124] The above components can be respectively disposed on independent chips, or at least partially or entirely disposed on the same chip. Whether to independently dispose each component on different chips or integrate them on one or more chips often depends on the needs of product design. The embodiments of the present application do not limit the specific implementation forms of the above components.
[0125] Figure 11 The illustrated computing device 1100 is merely exemplary. During implementation, the computing device 1100 may further include other components, which are not listed one by one herein.
[0126] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0127] It can be understood that the various numerical numbers involved in the embodiments of the present application are only for the convenience of description and are not used to limit the scope of the embodiments of the present application. It should be understood that in the embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0128] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present application. It should be understood that the above is only the specific embodiment of the present invention and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made on the basis of the technical solution of the present application should be included in the protection scope of the present application.
Claims
1. A method for processing resource specification change failures, characterized in that, Applied to a computing node, where the computing node deploys a computing service, the method includes: Obtain the execution result of a change operation, where the change operation is used to change the resource specification of the computing service; Determine whether the change operation fails according to the execution result; In the case where it is determined that the change operation fails, execute the fault handling strategy corresponding to the change operation.
2. The method according to claim 1, wherein The computing service includes a first program, and the method includes: Execute the first program to execute the method.
3. The method according to claim 1, wherein The computing service includes a first program and a second program, and the method further includes: Execute the first program to execute the change operation, obtain the execution result of the change operation, and determine whether the change operation fails according to the execution result; In the case where it is determined that the change operation fails, execute the second program to execute the fault handling strategy corresponding to the change operation.
4. The method according to claim 3, characterized in that The first program includes a data interface. In the case where it is determined that the change operation fails, the method further includes: Send the execution result to the second program through the data interface; Execute the second program to determine to execute the fault handling strategy corresponding to the change operation according to the execution result.
5. The method according to any one of claims 1-4, characterized in that, The change operation includes a resource allocation operation, a data snapshot mounting operation, a data consistency check operation, an address switching operation, or a management layer data update operation.
6. The method according to claim 5, wherein The fault handling strategy corresponding to the resource allocation operation includes: releasing the resources of the target specification applied for by the resource allocation operation, and releasing the conflict lock added by the resource allocation operation to the computing service; The fault handling strategy corresponding to the data snapshot mounting operation includes: releasing the data snapshot created by the data snapshot mounting operation, releasing the resources of the target specification, and releasing the conflict lock; The fault handling strategy corresponding to the data consistency check operation includes: releasing the backup data created by the data consistency check operation, releasing the data snapshot, releasing the resources of the target specification, and releasing the conflict lock; The fault handling strategy corresponding to the address switching operation includes: reverting the address of the computing service, releasing the backup data, releasing the data snapshot, releasing the resources of the target specification, and releasing the conflict lock; The fault handling strategy corresponding to the management layer data update operation includes: re-updating the resource specification of the computing service stored in the management node corresponding to the computing service to the resource specification of the computing service after the resource specification change.
7. The method according to any one of claims 1 to 6, characterized in that, The computing service includes a database service.
8. A processor, characterized in that, Applied to a computing device, where the computing device deploys a computing service, and the processor is used to execute the method according to any one of claims 1-7.
9. A computing device, characterized in that, The computing device includes: a processor and a memory, and the processor is used to execute the computer program stored in the memory to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Includes instructions that, when run on a computer, cause the computer to execute the method according to any one of claims 1 to 7.