A fault handling system, method and apparatus
Patent Information
- Application Number
- CN202310400712.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-14
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2043-04-14
AI Technical Summary
然而,现阶段大数据云平台需要对各个受影响的数据服务进行单独故障处理,而这种方法缺乏普适性,导致处理故障需要消耗较长时间,影响大数据云平台的正常使用
[0028]上述第二方面至第六方面的有益效果,具体请参照上述第一方面中相应设计可以达到的技术效果,这里不再重复赘述。
Smart Images

Figure CN116401019B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data, and in particular to a fault handling system, method and apparatus. Background Technology
[0002] An availability zone refers to a physical area where power and network are independent of each other. Fault isolation can be achieved between different availability zones, meaning that if one availability zone fails, it will not affect the normal operation of other availability zones. When extreme uncontrollable events such as natural disasters or power outages cause failures in the underlying infrastructure, it may cause availability zone failures, thereby making all data within the availability zone unavailable.
[0003] Currently, the big data cloud platform handles some of the bank's core business. When an availability zone fails, the big data cloud platform needs to handle the failure. However, at present, the big data cloud platform needs to handle each affected data service individually, and this method lacks universality, resulting in a long processing time and affecting the normal use of the big data cloud platform.
[0004] In summary, there is a current need for a fault handling system to improve the speed and efficiency of big data cloud platforms in handling faults when availability zones fail. Summary of the Invention
[0005] This invention provides a fault handling system, method, and apparatus to improve the speed and efficiency of big data cloud platform in handling faults when availability zones fail.
[0006] In a first aspect, the present invention provides a fault handling system, including an availability zone management module, an availability zone support module, and a business service module. The availability zone management module is used to acquire the availability status of an availability zone and the availability status of each cluster contained within the availability zone, wherein each cluster contains at least one data resource. Next, the availability zone support module is used to determine the availability status of each data resource contained in a preset data resource set based on the acquired availability status of the availability zone and the availability status of each cluster contained within the availability zone, wherein each data resource is divided into a primary data resource and at least one backup data resource. Furthermore, when the primary data resource is unavailable, the backup data resource that is available is determined as the new primary data resource in the data resource set. Subsequently, the business service module is used to perform data services based on the new primary data resource.
[0007] In the above method, when an availability zone fails, the availability zone management module will learn that the availability zone is unavailable and that all clusters within that availability zone are also unavailable. It can be understood that when a cluster is unavailable, all data resources stored in that cluster are also unavailable. Next, the availability zone support module will determine the availability status of each data resource in each preset data resource set. If a data resource is unavailable and is the primary data resource of a data resource set, the availability zone support module will designate an available backup data resource in that data resource set as the new primary data resource. After this, the business service module will execute data services based on this new primary data resource. Therefore, by setting up an availability zone support module in the system, the availability zone support module can uniformly confirm the availability status of each data resource in all data resource sets, and automatically replace unavailable primary data resources in the data resource sets with available backup data resources. This means that affected data services are handled uniformly, rather than requiring separate fault handling for each affected data service as in existing technologies. This helps improve the speed and efficiency of fault handling after an availability zone fails.
[0008] Optionally, the master data resource is determined in the following way: the availability zone support module is also used to sort all data resources contained in the data resource set according to priority, and the data resource with the highest priority is used as the master data resource.
[0009] By using the above method, the highest priority data resource can be selected as the master data resource of its data resource set, so that the business service module executes data services based on the highest priority data resource, thereby improving the service efficiency and service quality of the business service module.
[0010] Optionally, the availability zone support module is also used to compare the priority of the backup data resource with the priority of the new primary data resource when the backup data resource of any data resource set is restored to an available state; if the priority of the backup data resource is higher than the priority of the new primary data resource, the backup data resource is updated to the new primary data resource.
[0011] In the above method, when the availability zone that has failed returns to normal, the master data resources that were previously unavailable in the data resource set also return to normal. By selecting the data resource with the highest priority as the master data resource of its data resource set, the original standby master data resource is determined as the master data resource, thereby improving the service efficiency and service quality of the business service module.
[0012] Optionally, the availability zone support module is also used to record the set of data resources that are in an unavailable state, wherein the record is used by external components to investigate the affected data resources.
[0013] In the above method, by recording the set of data resources in an unavailable state, external components can be helped to investigate the affected data resources, thereby solving the problem that some data resources are lost or incomplete after the availability zone is restored from failure.
[0014] Optionally, the data resource set includes a composite resource group and a composite data source; the composite resource group includes a primary resource group and at least one backup resource group; the composite data source includes a primary data source and at least one backup data source.
[0015] Secondly, the present invention provides a fault handling method applicable to a fault handling system. The method includes: obtaining the availability status of an availability zone and the availability status of each cluster contained within the availability zone, wherein each cluster contains at least one data resource; determining the availability status of each data resource contained in a preset data resource set based on the availability status of the availability zone and the availability status of each cluster contained within the availability zone, wherein each data resource is divided into a primary data resource and at least one backup data resource; when the primary data resource is unavailable, determining the backup data resource that is available as a new primary data resource in the data resource set; and performing data services based on the new primary data resource.
[0016] Optionally, the master data resource is determined by sorting all data resources contained in the data resource set according to priority, and the data resource with the highest priority is taken as the master data resource.
[0017] Optionally, the method further includes: for any data resource set, if the backup data resource of the data resource set is restored to an available state, comparing the priority of the backup data resource with the priority of the new master data resource; if the priority of the backup data resource is higher than the priority of the new master data resource, updating the backup data resource to the new master data resource.
[0018] Optionally, the method further includes: recording a set of data resources that are in an unavailable state, wherein the record is used by external components to investigate the affected data resources.
[0019] Optionally, the data resource set includes a composite resource group and a composite data source; the composite resource group includes a primary resource group and at least one backup resource group; the composite data source includes a primary data source and at least one backup data source.
[0020] Thirdly, the present invention provides a fault handling apparatus, the data processing apparatus comprising: an acquisition unit, configured to acquire the availability status of an availability zone and the availability status of each cluster contained within the availability zone, wherein each cluster contains at least one data resource; and a processing unit, configured to determine the availability status of each data resource contained in a preset data resource set based on the availability status of the availability zone and the availability status of each cluster contained within the availability zone, wherein each data resource is divided into a primary data resource and at least one backup data resource, and when the primary data resource is unavailable, the backup data resource that is available is determined as the new primary data resource in the data resource set; and to perform data services based on the new primary data resource.
[0021] Optionally, the processing unit is specifically used to sort all the data resources contained in the data resource set according to priority, and the data resource with the highest priority is used as the master data resource.
[0022] Optionally, the processing unit is specifically configured to, for any data resource set, if the backup data resource of the data resource set is restored to an available state, compare the priority of the backup data resource with the priority of the new master data resource; if the priority of the backup data resource is higher than the priority of the new master data resource, update the backup data resource to the new master data resource.
[0023] Optionally, the processing unit is specifically used to record a set of data resources that are in an unavailable state, wherein the record is used by external components to investigate the affected data resources.
[0024] Optionally, the data resource set includes a composite resource group and a composite data source; the composite resource group includes a primary resource group and at least one backup resource group; the composite data source includes a primary data source and at least one backup data source.
[0025] Fourthly, the present invention provides a computing device including at least one processor and at least one memory, wherein the memory stores a computer program that, when executed by the processor, causes the processor to perform any of the fault handling methods described in the second aspect above.
[0026] Fifthly, the present invention also provides a computer-readable storage medium storing a program that, when run on a computer, causes the computer to perform any of the fault handling methods described in the second aspect above.
[0027] In a sixth aspect, the present invention also provides a computer program product including computer-readable instructions that, when executed by a processor, cause the method described in any possible design of the second aspect above to be implemented.
[0028] For details of the beneficial effects of aspects two through six above, please refer to the technical effects that can be achieved by the corresponding design in aspect one above, which will not be repeated here. Attached Figure Description
[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0030] Figure 1 This is a schematic diagram of the architecture of a fault handling system provided in an embodiment of the present invention;
[0031] Figure 2 A schematic diagram of the architecture of an availability zone support module provided in an embodiment of the present invention;
[0032] Figure 3 This is a schematic diagram of the architecture of a business service module provided in an embodiment of the present invention;
[0033] Figure 4 This is a flowchart illustrating a fault handling method provided in an embodiment of the present invention;
[0034] Figure 5 A structural diagram of a fault handling device provided in an embodiment of the present invention;
[0035] Figure 6 This is a schematic diagram of the structure of a computing device provided in an embodiment of the present invention. Detailed Implementation
[0036] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0037] As described in the background section, when an availability zone fails, the big data cloud platform needs to perform targeted fault handling on each service component, which consumes a considerable amount of time and affects the normal use of the big data cloud platform. Therefore, this invention proposes a fault handling system.
[0038] Figure 1 This is a schematic diagram of the architecture of a fault handling system provided in an embodiment of the present invention. Figure 1As shown, the fault handling system includes an availability zone management module 101, an availability zone support module 102, and a service module 103. The fault handling system can communicate with the availability zone, and the communication method can be wired or wireless, without limitation.
[0039] When communicating with the availability zone, the availability zone management module 101 obtains the availability status of the availability zone and the availability status of each cluster contained within the availability zone. Then, the availability zone support module 102 determines the availability status of each data resource contained in a preset data resource set based on the obtained availability status of the availability zone and the availability status of each cluster contained within the availability zone. Each data resource is divided into a primary data resource and at least one backup data resource. For any primary data resource, if the primary data resource is unavailable, the availability zone support module 102 determines the available backup data resource as the new primary data resource in the data resource set. Subsequently, the business service module 103 executes data services based on the new primary data resource.
[0040] In the above description, when an availability zone fails, the availability zone management module learns that the availability zone and all clusters within it are unavailable. This means that when a cluster is unavailable, all data resources stored within that cluster are also unavailable. Next, the availability zone support module determines the availability status of each data resource in each preset data resource set. If a data resource is unavailable and is the primary data resource of a data resource set, the availability zone support module designates an available backup data resource in that data resource set as the new primary data resource. Following this, the business service module executes data services based on this new primary data resource. Therefore, by setting up an availability zone support module in the system, it can uniformly confirm the availability status of each data resource in all data resource sets, and automatically replace unavailable primary data resources with available backup data resources. This provides unified fault handling for affected data services, rather than requiring separate fault handling for each affected data service as in existing technologies. This helps improve the speed and efficiency of fault handling when an availability zone fails.
[0041] To make the solutions of the embodiments of the present invention clearer, the specific functions of each module in the fault handling system are described below.
[0042] Availability Zone Management Module 101:
[0043] The Availability Zone Management Module 101 is used to obtain the availability status of an Availability Zone and the availability status of each cluster within the Availability Zone. The number of Availability Zones (AZs) can be one or more, without limitation. The availability status of an Availability Zone includes whether it is available or unavailable. When extreme unforeseen events such as natural disasters or power outages occur, the Availability Zone may fail, rendering it unavailable. To prevent all Availability Zones from becoming unavailable due to unforeseen events, at least two Availability Zones are typically set up, and these zones are physically isolated. For example, the Availability Zones can be located in different geographical locations, ensuring that each Availability Zone is not affected by failures in other Availability Zones. Furthermore, at least one cluster can be set up within an Availability Zone based on computing, network, and storage resources. Normally, when an Availability Zone is unavailable, all clusters within that zone will also be unavailable; however, when a cluster is unavailable, its Availability Zone may be either available or unavailable.
[0044] Furthermore, the cluster stores at least one data resource used by the big data cloud platform to perform data services. In one possible implementation, information about each availability zone and each cluster within each availability zone can be added to the availability zone management module, and the data resource can be associated with its respective cluster and availability zone. For example, if data resource 1 belongs to cluster 1-1 in availability zone 1, the availability zone management module will associate data resource 1 with cluster 1-1 and availability zone 1. As another example, if data resource 2 belongs to cluster 2-1 in availability zone 2, the availability zone management module will associate data resource 2 with cluster 2-1 and availability zone 2.
[0045] Availability Zone Support Module 102:
[0046] The availability zone support module 102 is used to determine the availability status of each data resource in the preset data resource set based on the availability status of the availability zone and the availability status of each cluster contained in the availability zone. When the primary data resource is unavailable, the standby data resource that is available is determined as the new primary data resource in the data resource set.
[0047] First, the availability zone support module 102 can determine the availability status of each data resource in each availability zone based on the availability status of the availability zone and the availability status of each cluster in the availability zone. In one possible implementation, for a data resource, when both its availability zone and cluster are available, the availability zone support module 102 can determine that the data resource is available; conversely, when at least one of the availability zone and cluster to which the data resource belongs is unavailable, the availability zone support module 102 can determine that the data resource is unavailable. For example, when both the availability zone and cluster to which the data resource belongs are unavailable, the availability zone support module 102 can determine that the data resource is unavailable; another example is when the availability zone to which the data resource belongs is available but the cluster is unavailable, the availability zone support module 102 can determine that the data resource is unavailable. In the aforementioned example, data resource 1 belongs to cluster 1-1 of availability zone 1, and data resource 2 belongs to cluster 2-1 of availability zone 2. If availability zone 1 is available and cluster 1-1 is available, availability zone support module 102 can determine that data resource 1 is available. If availability zone 1 is available and cluster 1-1 is available, availability zone support module 102 can determine that data resource 1 is available. If availability zone 2 is unavailable and cluster 1-1 is unavailable, availability zone support module 102 can determine that data resource 2 is available.
[0048] In this embodiment of the invention, to address availability zone failures, data resources are backed up. By backing up a data resource to clusters in different availability zones, data resource unavailability or loss due to availability zone failures is prevented. All data resources within the same data resource are referred to as a data resource set. A data resource set includes a primary data resource and at least one backup data resource.
[0049] Optionally, the master data resource can be determined as follows: the availability zone support module 102 sorts all data resources contained in the data resource set according to priority, and the data resource with the highest priority is taken as the master data resource.
[0050] The priority of data resources can be manually set by those skilled in the art based on the business type, or it can be automatically set by the fault handling system based on the business type; no specific limitation is made in this regard. Furthermore, the priority can be represented numerically; a smaller priority value can represent a higher priority, or a larger priority value can represent a higher priority; no specific limitation is made in this regard. Table 1 provides an example of a data resource set a. It should be noted that for data resources in data resource set a, a smaller priority value represents a higher priority.
[0051] Table 1 Data Resource Set a
[0052] Data resource a-1 Cluster 1-1, Availability Zone 1 1 Data resource a-2 Cluster 1-2, Availability Zone 1 2 Data resource a-3 Cluster 2-1, Availability Zone 2 3 Data resource a-4 Cluster 3-1, Availability Zone 3 4
[0053] As shown in Table 1, data resource set a includes four data resources: data resource a-1, data resource a-2, data resource a-3, and data resource a-4. Specifically, data resource a-1 is stored in cluster 1-1, availability zone 1, with a priority of 1; data resource a-2 is stored in cluster 1-2, availability zone 1, with a priority of 2; data resource a-3 is stored in cluster 2-1, availability zone 2, with a priority of 3; and data resource a-4 is stored in cluster 3-1, availability zone 3, with a priority of 4. Since a smaller priority value indicates a higher priority for data resources in data resource set a, and data resource a-1 has the lowest priority, data resource a-1 is the primary data resource of data resource set a.
[0054] Furthermore, after determining the availability status of each data resource in each availability zone, the availability zone support module 102 also determines the availability status of each data resource in each data resource set. For any preset data resource set, when its primary data resource is unavailable, the availability zone support module 102 determines the available backup data resource in the data resource set as the new primary data resource.
[0055] Continuing with the aforementioned data resource set a as an example, assume that availability zone 1 is unavailable, availability zone 2 and all clusters within availability zone 2 are available, and availability zone 3 and all clusters within availability zone 3 are available. Table 2 shows the availability status of each data resource in data resource set a.
[0056] Table 2. Availability status of data resource set a
[0057]
[0058] As shown in Table 2, since Availability Zone 1 is unavailable, all clusters within Availability Zone 1 are also unavailable. Therefore, clusters 1-1 and 1-2 are both unavailable. Similarly, since Availability Zone 2 is available, all clusters within Availability Zone 2 are available, therefore cluster 2-1 is available. Furthermore, since Availability Zone 3 is available, all clusters within Availability Zone 3 are available, therefore cluster 3-1 is available. For data resource a-1, since both cluster 1-1 and Availability Zone 1 are unavailable, data resource a-1 is unavailable. For data resource a-2, since both cluster 1-2 and Availability Zone 1 are unavailable, data resource a-2 is unavailable. For data resource a-3, since both cluster 2-1 and Availability Zone 2 are available, data resource a-2 is available. For data resource a-4, since both cluster 3-1 and Availability Zone 3 are available, data resource a-4 is available. However, as described above, the primary data resource of data resource set a is data resource a-1. Currently, data resource a-1 is in an unavailable state. Therefore, the primary data resource of data resource set a is in an unavailable state. The availability zone support module 102 needs to determine the available backup data resource in data resource set a as the new primary data resource.
[0059] In one possible implementation, the availability zone support module 102 can select the highest priority available backup data resource in the data resource set as the new master data resource.
[0060] In another possible implementation, the availability zone support module 102 may also randomly select a standby data resource that is in an available state from the data resource set as the new primary data resource.
[0061] In the example in Table 2, the available backup data resources include data resource a-3 and data resource a-4. Taking the example of the availability zone support module 102 selecting the highest priority available backup data resource as the new primary data resource, since data resource a-3 has the highest priority among the available backup data resources, the availability zone support module 102 determines data resource a-3 as the new primary data resource. At this time, the backup data resources of data resource set a include: data resource a-1 (unavailable), data resource a-2 (unavailable), and data resource a-4 (available).
[0062] Once the fault handling for the availability zone is complete, the availability zone will be restored to an available state, and all clusters within the availability zone will also be restored to an available state, along with the data resources in all clusters within the availability zone.
[0063] Optionally, the availability zone support module 102 is also used to compare the priority of the backup data resource with the priority of the new primary data resource when the backup data resource of any data resource set is restored to an available state; if the priority of the backup data resource is higher than the priority of the new primary data resource, the backup data resource is updated to the new primary data resource.
[0064] Continuing with data resource set a as an example, when the fault handling of Availability Zone 1 is completed and Availability Zone 1 is restored to an available state, all clusters in Availability Zone 1 also restore to an available state, that is, clusters 1-1 and 1-2 are both restored to an available state. Therefore, the data resources in all clusters of Availability Zone 1 are also restored to an available state. For data resource set a, data resources a-1 and a-2 in its spare data resources are restored to an available state. Table 3 shows a schematic table of data resource set a after the fault recovery of Availability Zone 1.
[0065] Table 3 Data resource set after fault recovery in Availability Zone 1 (a)
[0066]
[0067]
[0068] In Table 3, the new primary data resource a-3, standby data resources a-1, a-2, and a-4 are all in an available state. The availability zone support module 102 compares the priority of the standby data resources with the priority of the new primary data resource. Since the priorities of standby data resources a-1 and a-2 are higher than the priority of the new primary data resource, a primary data resource is selected from standby data resources a-1 and a-2. Here, the standby data source with the highest priority is determined as the primary data resource. Therefore, the availability zone support module 102 updates data resource a-1 to the new primary data resource.
[0069] Optionally, data resources may include data sources and resource groups. Typically, data resources may include a resource group and at least one data source. A data resource set may include composite resource groups and composite data sources. A composite data source includes a primary data source and at least one backup data source; a composite resource group includes a primary resource group and at least one backup resource group.
[0070] like Figure 2 The above is a schematic diagram of the architecture of an availability zone support module provided in an embodiment of the present invention. The availability zone support module 102 includes a data source management component 201 and a resource group management component 202.
[0071] In detail, the data source management component 201 is used to determine the availability status of each data source contained in each composite data source within the availability zone based on the availability status of the availability zone and the availability status of each cluster contained within the availability zone. When the primary data source of the composite data source is unavailable, the standby data source in the composite data source that is available is determined as the new primary data source.
[0072] In detail, the resource group management component 202 is used to determine the availability status of each resource group contained in each composite resource group in the availability zone based on the availability status of the availability zone and the availability status of each cluster contained in the availability zone. When the primary resource group of the composite resource group is unavailable, the standby resource group in the composite resource group that is available is determined as the new primary resource group.
[0073] Once the fault handling for the availability zone is complete and the availability zone is restored to an available state, all clusters within the availability zone will also be restored to an available state, and the data sources and resource groups in all clusters within the availability zone will also be restored to an available state.
[0074] Optionally, the data source management component 201 is also configured to, for any composite data source, compare the priority of the backup data source with the priority of the new primary data source when the backup data source of the composite data source is restored to an available state; if the priority of the backup data source is higher than the priority of the new primary data source, update the backup data source to the new primary data source.
[0075] Optionally, the resource group management component 202 is further configured to, for any composite resource group, compare the priority of the standby resource group with the priority of the new primary resource group when the standby resource group of the composite resource group is restored to an available state; if the priority of the standby resource group is higher than the priority of the new primary resource group, update the standby resource group to the new primary resource group.
[0076] Please continue to refer to Figure 1 As shown, the business service module 103 is used to perform data services based on the new master data resources. Specifically, the availability zone support module transmits the master data resources stored in the availability zone cluster to the business service module 103 via routing, and then the business service module 103 performs data services based on the new master data resources.
[0077] In one possible implementation, when the business service module 103 performs data services based on the new master data resource, it can simultaneously write data to the new master data resource and the backup data resource, thereby reducing the impact of availability zone failures on the big data cloud platform.
[0078] Optionally, the business service module 103 includes at least one service component. In one possible implementation, different service components can perform different types of data services; for example, a service component can be a data acquisition component or a data integration component. Each service component corresponds to a data resource set, and the service component performs data tasks based on the master data resources of its data resource set.
[0079] like Figure 3 The above is a schematic diagram of the architecture of a business service module provided in an embodiment of the present invention. The business service module 103 includes a data acquisition component 301, a data integration component 302, a stream computing component 303, and a data service component 304. The data acquisition component 301 is one of the fundamental components of the big data cloud platform, and its main function is to collect data from multiple sources. The data integration component 302 is used to integrate data from different data sources. In the big data cloud platform, the data integration component can help users integrate data scattered across multiple data sources into a single data repository for analysis and processing. The stream computing component 303 is a component for real-time data stream processing. In the big data cloud platform, the stream computing component can help users process, aggregate, filter, compute, and analyze real-time data streams for rapid response. The data service component 304 is a component for providing data services. In the big data cloud platform, the data service component can help users store data in the cloud and provide interfaces so that users can access this data from applications.
[0080] In one possible implementation, for different types of service modules, when an availability zone fails and the primary data resource of the data resource set is unavailable, the rules by which the availability zone support module 102 determines a new primary data resource from the backup data resources can be different. The rules can be priority order or can be set by the operation and maintenance personnel, and there is no specific limitation.
[0081] It should be noted that, Figure 3 This is merely an illustrative example and does not constitute a limitation on the solution. Understandably, the service components in the business service module 103 can be configured by those skilled in the art according to actual needs, without any specific limitations.
[0082] After a failed availability zone is restored, some data resources may be lost or incomplete. To avoid this problem, the availability zone support module 102 is also used to record the set of data resources that are in an unavailable state. The record is used by external components to investigate the affected data resources.
[0083] In detail, the business service module 103 records the data resources executed during service execution. Since the availability zone support module 102 also records the set of data resources in an unavailable state, the business service module 103 and the availability zone support module 102 can be used to determine the data resources affected by an availability zone failure. Afterwards, external components can be used to investigate the affected data resources, i.e., which data resources have been corrupted.
[0084] based on Figure 1 An introduction to the fault handling system. Figure 4 A flowchart of a fault handling method is provided, which includes the following steps:
[0085] Step 401: The availability zone management module 101 obtains the availability status of the availability zone and the availability status of each cluster contained within the availability zone.
[0086] Specifically, the availability zone management module 101 obtains the availability status of each availability zone and the availability status of each cluster within each availability zone.
[0087] Step 402: The availability zone management module 101 sends the availability status of the availability zone and the availability status of each cluster within the availability zone to the availability zone support module 102.
[0088] Specifically, the availability zone support module 102 receives the availability status of the availability zone and the availability status of each cluster within the availability zone.
[0089] Step 403: The availability zone support module 102 determines the availability status of each data resource contained in the preset data resource set.
[0090] Specifically, the availability zone support module 102 determines the availability status of each data resource (primary data resource and backup data resource) in each preset data resource set based on the availability status of the availability zone and the availability status of each cluster contained in the availability zone.
[0091] Step 404: When the master data resource is in an unavailable state, the availability zone support module 102 determines the standby data resource that is in an available state as the new master data resource in the data resource set.
[0092] Specifically, when the master data resource of a certain data resource set is unavailable, the availability zone support module 102 determines the available backup data resource in the data resource set as the new master data resource.
[0093] Step 405: Availability Zone Support Module 102 routes the new master data resource to Business Service Module 103.
[0094] Specifically, the availability zone support module 102 transmits the new master data resources to the business service module 103 via routing.
[0095] Step 406: Business service module 103 performs data services based on the new master data resources.
[0096] Through steps 401 to 406 above, when an availability zone fails, the affected data services can be handled uniformly, thereby improving the speed and efficiency of fault handling.
[0097] Based on the same inventive concept described above, the present invention also provides a fault handling device that can perform the methods described in the embodiments of the invention. The structure of the fault handling device provided by the present invention can be found in [reference needed]. Figure 5 The fault handling device 500 includes an acquisition unit 501 and a processing unit 502. The acquisition unit 501 acquires the availability status of an availability zone and the availability status of each cluster within the availability zone, where each cluster contains at least one data resource. The processing unit 502 determines the availability status of each data resource in a preset data resource set based on the availability status of the availability zone and the availability status of each cluster within the availability zone. Each data resource is divided into a primary data resource and at least one backup data resource. When the primary data resource is unavailable, the available backup data resource is designated as the new primary data resource in the data resource set. Data services are then performed based on the new primary data resource.
[0098] For a more detailed description of the acquisition unit 501 and the processing unit 502 mentioned above, please refer to [reference needed]. Figure 3 The relevant descriptions in the fault handling method embodiments shown are directly obtained and will not be repeated here.
[0099] Based on the same technical concept, the present invention also provides a computing device, such as... Figure 6 As shown, the computing device 600 includes at least one processor 601 and a memory 602 connected to the at least one processor. The specific connection medium between the processor 601 and the memory 602 is not limited in this invention. Figure 6 Taking the connection between the processor 601 and the memory 602 via a bus as an example, the bus can be divided into address bus, data bus, control bus, etc.
[0100] In this invention, memory 602 stores instructions that can be executed by at least one processor 601. At least one processor 601 can execute the steps included in the aforementioned fault handling method by executing the instructions stored in memory 602.
[0101] The processor 601 serves as the control center of the computing device, connecting various parts of the device via various interfaces and lines. It performs fault handling by running or executing instructions stored in the memory 602 and retrieving data stored in the memory 602. Optionally, the processor 601 may include one or more processing units. The processor 601 may integrate an application processor and a modem processor. The application processor primarily handles the operating system, user interface, and applications, while the modem processor primarily handles issuing instructions. It is understood that the modem processor may not be integrated into the processor 601. In some embodiments, the processor 601 and the memory 602 may be implemented on the same chip; in other embodiments, they may be implemented on separate chips.
[0102] Processor 601 can be a general-purpose processor, such as a central processing unit (CPU), digital signal processor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in this invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the fault handling method can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.
[0103] Memory 602, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory 602 may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic storage, magnetic disk, optical disk, etc. Memory 602 can be any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. Memory 602 in this invention can also be a circuit or any other device capable of implementing storage functions for storing program instructions and / or data.
[0104] Based on the same technical concept, the present invention also provides a computer-readable storage medium storing a computer program executable by a computing device, which, when run on the computing device, causes the computing device to perform the steps of the above-described fault handling method.
[0105] Based on the same technical concept, embodiments of the present invention also provide a computer program product, including computer-readable instructions, which, when executed by a processor, enable the above-described fault handling method to be implemented.
[0106] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0107] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0108] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0109] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0110] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A fault handling system, characterized in that, It includes an availability zone management module, an availability zone support module, and a business service module: The availability zone management module is used to obtain the availability status of the availability zone and the availability status of each cluster contained in the availability zone, wherein each cluster contains at least one data resource. The availability zone support module is used to determine the availability status of each data resource in a preset data resource set based on the availability status of the availability zone and the availability status of each cluster contained in the availability zone. This includes: for a data resource, if both the availability zone and the cluster to which the data resource belongs are in an available state, then the data resource is determined to be in an available state; if at least one of the availability zone and the cluster to which the data resource belongs is in an unavailable state, then the data resource is determined to be in an unavailable state. Each data resource is divided into a primary data resource and at least one backup data resource. When the primary data resource is unavailable, the backup data resource that is available is determined as the new primary data resource in the data resource set. The availability zone support module is also used to sort all data resources contained in the data resource set according to priority, with the data resource with the highest priority being the master data resource. The business service module is used to perform data services based on the new master data resources.
2. The fault handling system according to claim 1, characterized in that: The availability zone support module is further configured to, for any data resource set, if the backup data resource of the data resource set is restored to an available state, compare the priority of the backup data resource with the priority of the new primary data resource; if the priority of the backup data resource is higher than the priority of the new primary data resource, update the backup data resource to the new primary data resource.
3. The fault handling system according to claim 1, characterized in that: The availability zone support module is also used to record the set of data resources that are in an unavailable state, wherein the record is used by external components to investigate the affected data resources.
4. The fault handling system according to any one of claims 1-3, characterized in that, The data resource set includes composite resource groups and composite data sources; the composite resource group includes a primary resource group and at least one backup resource group; the composite data source includes a primary data source and at least one backup data source.
5. A fault handling method, characterized in that, Applicable to fault handling systems, the method includes: Obtain the availability status of the availability zone and the availability status of each cluster contained within the availability zone, wherein each cluster contains at least one data resource; Based on the availability status of the availability zone and the availability status of each cluster contained within the availability zone, the availability status of each data resource in the preset data resource set is determined, including: for a data resource, if both the availability zone and the cluster to which the data resource belongs are in an available state, then the data resource is determined to be in an available state; if at least one of the availability zone and the cluster to which the data resource belongs is in an unavailable state, then the data resource is determined to be in an unavailable state. Each data resource is divided into a primary data resource and at least one backup data resource. When the primary data resource is unavailable, the backup data resource that is available is determined as the new primary data resource in the data resource set. All data resources in the data resource set are sorted according to priority, and the data resource with the highest priority is taken as the primary data resource. Data services are performed based on the new master data resource.
6. The method according to claim 5, characterized in that, The method further includes: For any data resource set, if the backup data resource of the data resource set is restored to an available state, the priority of the backup data resource is compared with the priority of the new master data resource; if the priority of the backup data resource is higher than the priority of the new master data resource, the backup data resource is updated to the new master data resource.
7. The method according to claim 5, characterized in that, The method further includes: A set of data resources that are in an unavailable state is recorded, wherein the record is used by external components to investigate the affected data resources.
8. The method according to any one of claims 5-7, characterized in that, The data resource set includes composite resource groups and composite data sources; the composite resource group includes a primary resource group and at least one backup resource group; the composite data source includes a primary data source and at least one backup data source.
9. A computing device, characterized in that, It includes at least one processor and at least one memory, wherein the memory stores a computer program that, when executed by the processor, causes the processor to perform the method according to any one of claims 5 to 8.
10. A computer-readable storage medium, characterized in that, The storage medium stores a program that, when run on a computer, causes the computer to perform the method described in any one of claims 5 to 8.
11. A computer program product, characterized in that, Includes computer-readable instructions that, when executed by a processor, cause the method as described in any one of claims 5 to 8 to be implemented.
Citation Information
Patent Citations
Fault processing method of resources and device
CN105515812A
Alarm system and method for distributed metadata cluster
CN107465560A