Fault isolation method of distributed system and related device

By automatically comparing the target operating parameters and fault parameter ranges in a distributed system, identifying and isolating faults, the problem of fault processing in the existing technology relies on manual operation and maintenance, and the accuracy and efficiency of fault processing are improved.

CN120216231APending Publication Date: 2025-06-27NETSUNION CLEARING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311823616.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, the fault handling of distributed systems relies on manual operation and maintenance, resulting in the fault handling process being unfixed and the efficiency and accuracy are low.

Method used

By obtaining multiple types of target operating parameters at a specified time in a distributed system, comparing these parameters with pre-set fault parameter ranges, automatically identifying and isolating faults, and cutting off the services corresponding to the faults.

Benefits of technology

It realizes automatic identification and isolation of system failures, improves the accuracy and efficiency of fault handling, and reduces the need for manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216231A_ABST
    Figure CN120216231A_ABST
Patent Text Reader

Abstract

The invention provides a fault isolation method of a distributed system and a related device. The method and the device are used for improving system fault processing efficiency and accuracy. The method comprises the steps that multiple types of target operation parameters in the distributed system are obtained every specified duration, and the number of any type of target operation parameters is at least one; for any type of target operation parameter, based on the target operation parameter and a target fault parameter range, whether a target fault corresponding to the type exists in the distributed system is determined, and the target fault parameter range is a preset fault parameter range corresponding to the type; if the fault type exists, determining a target fault isolation strategy corresponding to the type by using a preset corresponding relationship between the fault type and the fault isolation strategy; and performing fault isolation on a target fault corresponding to the type by using the target fault isolation strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] With the rapid development of the network, more and more distributed systems have emerged on the market. For example, distributed storage systems, etc. As the business continues to increase. Among them, if any link in the system's link fails, it may affect the platform business, affect the success rate of the system, and even cause economic losses.

[0003] Therefore, it is necessary to establish a full-platform-level fault emergency handling process that can accurately alarm, quickly respond, and quickly handle. In order to control, mitigate, and eliminate fault risks in a timely and effective manner. Prevent the occurrence of business operation interruption events and ensure the business continuity of the entire link.

[0004] In the prior art, the fault handling method in a distributed system mainly relies on manual operation and maintenance. According to different fault characteristics and the ability levels of operation and maintenance personnel, the adopted fault handling process has certain contingency and differences. And this contingency and difference may directly lead to an unfixed fault handling process, resulting in low efficiency and accuracy of the system's fault handling. Summary of the Invention

[0005] In an exemplary embodiment of the present disclosure, a fault isolation method and related device for a distributed system are provided to improve the processing efficiency and accuracy of system faults.

[0006] A first aspect of the present disclosure provides a method for isolating system faults, the method including:

[0007] At every specified time interval, obtain multiple types of target operating parameters in the distributed system, where the number of target operating parameters of any one type is at least one;

[0008] For any one type of target operating parameter, based on the target operating parameter and the target fault parameter range, determine whether the distributed system has a target fault corresponding to the type, where the target fault parameter range is a pre-set fault parameter range corresponding to the type; and,

[0009] If so, use the pre-set correspondence between the fault type and the fault isolation strategy to determine the target fault isolation strategy corresponding to the type; and,

[0010] Use the target fault isolation strategy to isolate the target fault corresponding to the type to cut off the target service corresponding to the target fault in the distributed system.

[0011] In this embodiment, for any type of target operating parameter, based on the target operating parameter and the target fault parameter range, it is determined whether there is a target fault corresponding to the type in the distributed system; if so, using the pre-set correspondence between the fault type and the fault isolation policy, the target fault isolation policy corresponding to the type is determined, and the target fault corresponding to the type is isolated using the target fault isolation policy to cut off the target service corresponding to the target fault in the distributed system. Thus, in the embodiment of the present application, system faults can be automatically identified and automatically isolated without manual fault detection and handling, so the accuracy and efficiency of fault handling are improved.

[0012] In one embodiment, the determining whether there is a target fault corresponding to the type in the distributed system based on the target operating parameter and the target fault parameter range includes:

[0013] Comparing the target operating parameter with the target fault parameter range;

[0014] If there is a target operating parameter in the target operating parameters that is within the target fault parameter range, it is determined that there is a target fault corresponding to the type in the distributed system, and the service corresponding to the target operating parameter within the target fault parameter range is determined as the target fault;

[0015] If there is no target operating parameter in the target operating parameters that is within the target fault parameter range, it is determined that there is no target fault corresponding to the type in the distributed system.

[0016] In this embodiment, by comparing the target operating parameter with the target fault parameter range to determine whether there is a target fault corresponding to each type in the system, the accuracy of the identified faults is ensured.

[0017] In one embodiment, the type includes server type, database type, network dedicated line type, data center type, and city center type;

[0018] The target fault isolation policy includes at least one of a server fault isolation policy, a database fault isolation policy, a network dedicated line fault isolation policy, a data center fault isolation policy, and a city fault isolation policy;

[0019] The determining the target fault isolation policy corresponding to the type using the pre-set correspondence between the fault type and the fault isolation policy includes:

[0020] If the type is a server type, determine that the target fault isolation policy corresponding to the type is a server fault isolation policy; or,

[0021] If the type is a database type, determine that the target fault isolation policy corresponding to the type is a database fault isolation policy; or,

[0022] If the type is a network dedicated line type, determine that the target fault isolation policy corresponding to the type is a network dedicated line fault isolation policy; or,

[0023] If the type is a data center type, determine that the target fault isolation policy corresponding to the type is a data center fault isolation policy; or,

[0024] If the type is an urban center type, determine that the target fault isolation policy corresponding to the type is an urban fault isolation policy.

[0025] In this embodiment, by using the pre-set correspondence between the fault type and the fault isolation policy, the target fault isolation policy corresponding to the type is determined to ensure that the corresponding fault isolation policy is used to isolate the target fault, improving the accuracy of fault handling.

[0026] In one embodiment, the using the target fault isolation policy to perform fault isolation on the target fault corresponding to the type includes:

[0027] If the number of target faults corresponding to the type is one, use the target fault isolation policy to perform fault isolation on the target fault; or,

[0028] If the number of target faults corresponding to the type is multiple, obtain the priority between the target faults based on the pre-set priority between each service, and use the target fault isolation policy to isolate each target fault in turn according to the priority between the target faults.

[0029] In this embodiment, the fault isolation is performed in a corresponding manner according to the number of target faults to ensure that each target fault can be processed, improving the accuracy of fault handling.

[0030] In one embodiment, before obtaining the target operation parameters of multiple types in the distributed system at every specified duration, the method further includes:

[0031] Respond to the specified duration setting instruction sent by the user, and set the specified duration based on the specified duration setting instruction.

[0032] In this embodiment, the user can set the corresponding specified duration, which can be configured according to the user's needs, meeting the user's needs.

[0033] The second aspect of the present disclosure provides a fault isolation device for a distributed system, and the device includes:

[0034] An acquisition module, configured to acquire multiple types of target operation parameters in the distributed system at intervals of a specified duration, where the number of target operation parameters of any one type is at least one;

[0035] A target fault determination module, configured to, for any one type of target operation parameter, determine whether the distributed system has a target fault corresponding to the type based on the target operation parameter and a target fault parameter range, where the target fault parameter range is a pre-set fault parameter range corresponding to the type; and

[0036] An isolation policy determination module, configured to, if there is a fault, determine a target fault isolation policy corresponding to the type by using a pre-set correspondence between fault types and fault isolation policies; and

[0037] A fault isolation module, configured to perform fault isolation on the target fault corresponding to the type by using the target fault isolation policy to cut off the target service corresponding to the target fault in the distributed system.

[0038] In one embodiment, the target fault determination module is specifically configured to:

[0039] Compare the target operation parameter with the target fault parameter range;

[0040] If there is a target operation parameter among the target operation parameters that is within the target fault parameter range, determine that the distributed system has a target fault corresponding to the type, and determine the service corresponding to the target operation parameter within the target fault parameter range as the target fault;

[0041] If there is no target operation parameter among the target operation parameters that is within the target fault parameter range, determine that the distributed system has no target fault corresponding to the type.

[0042] In one embodiment, the type includes server type, database type, network dedicated line type, data center type, and city center type;

[0043] The target fault isolation policy includes at least one of a server fault isolation policy, a database fault isolation policy, a network dedicated line fault isolation policy, a data center fault isolation policy, and a city fault isolation policy;

[0044] The isolation policy determination module is specifically configured to:

[0045] If the type is a server type, determine that the target fault isolation policy corresponding to the type is a server fault isolation policy; or,

[0046] If the type is a database type, determine that the target fault isolation policy corresponding to the type is a database fault isolation policy; or,

[0047] If the type is a network dedicated line type, determine that the target fault isolation policy corresponding to the type is a network dedicated line fault isolation policy; or,

[0048] If the type is a data center type, determine that the target fault isolation policy corresponding to the type is a data center fault isolation policy; or,

[0049] If the type is an urban center type, determine that the target fault isolation policy corresponding to the type is an urban fault isolation policy.

[0050] In one embodiment, the fault isolation module is specifically configured to:

[0051] If the number of target faults corresponding to the type is one, use the target fault isolation policy to isolate the target fault; or,

[0052] If the number of target faults corresponding to the type is multiple, obtain the priorities among the target faults based on the pre-set priorities among various services, and use the target fault isolation policy to isolate the target faults in sequence according to the priorities among the target faults.

[0053] In one embodiment, the apparatus further includes:

[0054] A specified duration setting module, configured to, before obtaining the target running parameters of multiple types in the distributed system every specified duration, in response to a specified duration setting instruction sent by a user, set the specified duration based on the specified duration setting instruction.

[0055] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including:

[0056] At least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executed by the at least one processor; the instructions are executed by the at least one processor so that the at least one processor can execute the method as described in the first aspect.

[0057] According to a fourth aspect provided by the embodiments of the present disclosure, there is provided a computer storage medium storing a computer program for executing the method as described in the first aspect. Description of the Drawings

[0058] To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for description in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0059] Figure 1 It is a schematic diagram of an applicable scenario according to an embodiment of the present disclosure;

[0060] Figure 2 It is one of the flow schematic diagrams of the fault isolation method for a distributed system according to an embodiment of the present disclosure;

[0061] Figure 3 It is a flow schematic diagram for determining whether there is a target fault corresponding to the type in the system according to an embodiment of the present disclosure;

[0062] Figure 4 It is another flow schematic diagram of the fault isolation method for a distributed system according to an embodiment of the present disclosure;

[0063] Figure 5 It is a fault isolation device for a distributed system according to an embodiment of the present disclosure;

[0064] Figure 6 It is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. Detailed Embodiments

[0065] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the following clearly and completely describes the technical solutions in the embodiments of the present disclosure with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts belong to the scope of protection of the present disclosure.

[0066] The term "and / or" in the embodiments of the present disclosure describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0067] The application scenarios described in the embodiments of the present disclosure are for more clearly illustrating the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those of ordinary skill in the art will know that with the emergence of new application scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems. Among them, in the description of the present disclosure, unless otherwise specified, the meaning of "a plurality of" is two or more.

[0068] In the prior art, mainly relying on manual operation and maintenance, according to different fault characteristics and the ability levels of operation and maintenance personnel, the adopted processing procedures have certain contingency and differences. And this contingency and difference may directly lead to the instability of the fault processing procedure, resulting in low efficiency and accuracy of the system's fault processing.

[0069] Therefore, the present disclosure provides a fault isolation method for a distributed system. For any type of target operating parameter, based on the target operating parameter and the target fault parameter range, it is determined whether the system has a target fault corresponding to the type; if so, using the pre-set correspondence between the fault type and the fault isolation strategy, the target fault isolation strategy corresponding to the type is determined, and the target fault corresponding to the type is isolated using the target fault isolation strategy. Thus, in the embodiments of the present application, the faults of the system can be automatically identified and automatically isolated, without the need for manual fault detection and processing. Therefore, the accuracy and efficiency of fault processing are improved. Next, the solution of the present disclosure will be introduced in detail with reference to the accompanying drawings.

[0070] As Figure 1 shown, an application scenario of a system fault isolation method, which includes a terminal device 110 and a server 120 in this application scenario.

[0071] In a possible application scenario, the server 120 obtains multiple types of target operating parameters in the distributed system at regular intervals. Among them, the number of any type of target operating parameter is at least one; for any type of target operating parameter, the server 120 determines whether the distributed system has a target fault corresponding to the type based on the target operating parameter and the target fault parameter range, where the target fault parameter range is the pre-set fault parameter range corresponding to the type; and if so, the server 120 uses the pre-set correspondence between the fault type and the fault isolation strategy to determine the target fault isolation strategy corresponding to the type; and uses the target fault isolation strategy to isolate the target fault corresponding to the type, so as to cut off the target service corresponding to the target fault in the distributed system and display it through the terminal device 110.

[0072] Among them, Figure 1 information interaction can be carried out between the middle terminal device 110 and the server 120 through a communication network. Among them, the communication method adopted by the communication network can be divided into a wireless communication method or a wired communication method.

[0073] Exemplarily, the server 120 can access the network through cellular mobile communication technology and communicate with the terminal device 110. Among them, the cellular mobile communication technology, for example, includes the fifth-generation mobile communication (5th Generation Mobile Networks, 5G) technology.

[0074] Optionally, the server 120 can access the network through a short-range wireless communication method and communicate with the terminal device 110. Among them, the short-range wireless communication method, for example, includes Wireless Fidelity (Wi-Fi) technology.

[0075] Among them, in the description of the present application, only a single server 120 and a single terminal device 110 are described in detail. However, those skilled in the art should understand that the shown server 120 and terminal device 110 are intended to represent the operations of the server 120 and terminal device 110 involved in the technical solution of the present application. Instead of implying limitations on the number, type, or location of the server 120 and terminal device 110. It should be noted that if additional modules are added to or individual modules are removed from the illustrated environment, the underlying concept of the exemplary embodiments of the present application will not be changed.

[0076] It should be noted that the fault isolation method of the distributed system proposed in the present application is not only applicable to Figure 1 the application scenarios shown, but also applicable to any fault isolation device of a distributed system.

[0077] Next, in combination with the above-described application scenarios, the fault isolation method of the distributed system of the exemplary embodiments of the present application will be described with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown for the convenience of understanding the method and principle of the present application, and the embodiments of the present application are not limited in this regard.

[0078] As Figure 2 shown, it is a schematic flowchart of the fault isolation method of the distributed system of the present disclosure, and may include the following steps:

[0079] Step 201: Obtain multiple types of target operating parameters in the distributed system at regular intervals, where the number of target operating parameters of any one type is at least one;

[0080] The types in the embodiments of the present application include server type, database type, network dedicated line type, data center type, and city center type, etc. The target operating parameters corresponding to each type are pre-set. For example, the target operating parameter corresponding to the server type is the server parameter, and this server parameter is used to represent the operating state of the corresponding server. The target operating parameter corresponding to the database type is the database parameter, and this database parameter is used to represent the state of the database. The target operating parameter corresponding to the network dedicated line type is the network dedicated line operating parameter, and this network dedicated line operating parameter is used to represent the state of the corresponding network cable. The target operating parameter corresponding to the data center type is the data center operating parameter, and this data center operating parameter is used to represent the operating state of a single data center device. The target operating parameter corresponding to the city center type is the city center operating parameter, and this city center operating parameter is used to represent the state of all data center devices.

[0081] It should be noted that: the target operating parameters corresponding to each type in the embodiments of the present application can be set according to the actual situation, and the embodiments of the present application do not limit the target operating parameters corresponding to each type here.

[0082] In one embodiment, the specified duration is determined in the following manner:

[0083] In response to the specified duration setting instruction sent by the user, the specified duration is set based on the specified duration setting instruction.

[0084] It should be noted that: the specified duration in the embodiments of the present application can be 0.1 second, 1 second, 10 seconds, etc. The specified duration in the embodiments of the present application is only for illustrative purposes, but does not limit the specified duration in the embodiments of the present application. The specified duration in the embodiments of the present application can be set according to the actual situation.

[0085] Step 202: For the target operating parameter of any one type, based on the target operating parameter and the target fault parameter range, determine whether there is a target fault corresponding to the type in the distributed system, where the target fault parameter range is a pre-set fault parameter range corresponding to the type;

[0086] The target fault in the embodiments of the present application may be at least one of a server fault, a database fault, a network dedicated line fault, a data center fault, and a city fault. Among them, the server fault in the embodiments of the present application is: when at least one server fails, it is determined that a server fault occurs. The database fault in the embodiments of the present application is: when at least one database fails, it is determined that a database fault occurs. The network dedicated line fault in the embodiments of the present application is: when at least one network cable fails, it is determined that a network dedicated line fault occurs. The data center fault in the embodiments of the present application is: when a single data center device fails, it is determined that a data center fault occurs. The city fault in the embodiments of the present application is: when all data center devices fail, it is determined that a city fault occurs.

[0087] Next, the method for determining whether there is a target fault corresponding to the type in the distributed system in step 202 will be described in detail. As Figure 3 shown, it is a flowchart for determining whether there is a target fault corresponding to the type in the distributed system, which may specifically include the following steps:

[0088] Step 301: Compare the target operating parameter with the target fault parameter range; among them, Table 1 is the correspondence between each type and the fault parameter range:

[0089] Type Fault parameter range Server type A to B Database type M to N Network dedicated line type H to K Data center type C to D City center type E to F … …

[0090] Table 1

[0091] For example, as shown in Table 1, the target fault parameter range corresponding to the target operating parameter of the server type is A to B. The target fault parameter range corresponding to the target operating parameter of the database type is M to N. The target fault parameter range corresponding to the target operating parameter of the network dedicated line type is H to K. The target fault parameter range corresponding to the target operating parameter of the data center type is C to D. The target fault parameter range corresponding to the target operating parameter of the city center type is E to F.

[0092] Step 302: Determine whether there is a target operating parameter within the target fault parameter range in the target operating parameters. If so, execute step 303; if not, execute step 304;

[0093] Step 303: Determine that there is a target fault corresponding to the type in the distributed system, and determine the service corresponding to the target operating parameter within the target fault parameter range as the target fault;

[0094] Step 304: Determine that there is no target fault corresponding to the type in the distributed system.

[0095] For example, if the target operating parameters of the server type include a and b, the target fault parameter range corresponding to the server failure is A to B. If the target operating parameter a is within the target fault parameter range A to B and / or the target operating parameter b is within the target fault parameter range A to B, it is determined that there is a server failure in the system. If the target operating parameter a is not within the target fault parameter range A to B and the target operating parameter b is not within the target fault parameter range A to B, it is determined that there is no server failure in the system.

[0096] Step 203: If it exists, use the pre-set correspondence between the fault type and the fault isolation policy to determine the target fault isolation policy corresponding to the type.

[0097] In one embodiment, step 203 can be specifically implemented in the following five ways:

[0098] Method 1: If the type is the server type, determine that the target fault isolation policy corresponding to the type is the server fault isolation policy.

[0099] The server fault isolation policy in the embodiment of the present application can be to adjust the weights of each server to achieve fault isolation. For example, set the weight of the faulty server to 0, then the server stops working, achieving the isolation of the server.

[0100] Method 2: If the type is the database type, determine that the target fault isolation policy corresponding to the type is the database fault isolation policy.

[0101] The database fault isolation policy in the embodiment of the present application can be to adjust the weights of each database to achieve fault isolation. For example, set the weight of the faulty database to 0, then the database stops working, achieving the isolation of the database.

[0102] Method 3: If the type is the network dedicated line type, determine that the target fault isolation policy corresponding to the type is the network dedicated line fault isolation policy.

[0103] The network dedicated line fault isolation policy in the embodiment of the present application can be to adjust the weights of each network cable to achieve fault isolation. For example, set the weight of the faulty network cable to 0, then the network cable stops working, achieving the isolation of the network cable.

[0104] Method 4: If the type is the data center type, determine that the target fault isolation policy corresponding to the type is the data center fault isolation policy.

[0105] In the embodiments of the present application, the data center fault isolation policy may be to adjust the control identifier corresponding to a single data center device. For example, set the control identifier of the faulty data center device to a specified identifier so that the data center device stops working. The specified identifier in the embodiments of the present application can be set according to the actual situation, and the embodiments of the present application do not limit the specified identifier here.

[0106] Mode Five: If the type is the city center type, determine that the target fault isolation policy corresponding to the type is the city fault isolation policy.

[0107] The city fault isolation policy in the embodiments of the present application may be to adjust the control identifiers corresponding to each data center device. For example, when the control identifiers of all data center devices are set to the specified identifier, so that each device center device stops working.

[0108] It should be noted that: the server fault isolation policy, database fault isolation policy, network dedicated line fault isolation policy, data center fault isolation policy, and city fault isolation policy described above are only for illustrative purposes. The specific isolation methods in the server fault isolation policy, database fault isolation policy, network dedicated line fault isolation policy, data center fault isolation policy, and city fault isolation policy in the embodiments of the present application can be set according to the actual situation, and the embodiments of the present application do not limit the specific methods of each isolation policy here.

[0109] Step 204: Use the target fault isolation policy to isolate the target fault corresponding to the type, so as to cut off the target service corresponding to the target fault in the distributed system.

[0110] In the embodiments of the present application, if the corresponding device stops working, the service corresponding to the device is isolated. For example, if the target isolation policy is the server fault isolation policy, use the server fault isolation policy to isolate the faulty server, that is, stop the working of the faulty server, then the service corresponding to the server in the distributed system will no longer be executed, so as to achieve the isolation of the service, that is, cut off the service in the distributed system.

[0111] In one embodiment, step 204 may be specifically implemented in the following two ways:

[0112] Mode One: If the number of target faults corresponding to the type is one, use the target fault isolation policy to isolate the target fault.

[0113] The embodiments of the present application do not limit the specific methods of the fault isolation policies corresponding to each type, and can be set according to the actual situation.

[0114] Method 2: If the number of target faults corresponding to the type is multiple, then obtain the priorities among the target faults based on the pre-set priorities among the services, and use the target fault isolation strategy to isolate the target faults in sequence according to the priorities among the target faults.

[0115] For example, the priorities among the services are in sequence: Service 1 > Service 2 > Service 3 > Service 4 > Service 5. If the target faults include Service 1, Service 3, and Service 4, then determine the priorities among the target faults as: Service 1 > Service 3 > Service 4.

[0116] To further understand the technical solution of the present disclosure, the following is described in detail in conjunction with Figure 4 and may include the following steps:

[0117] Step 401: In response to a specified duration setting instruction sent by the user, set the specified duration based on the specified duration setting instruction;

[0118] Step 402: Every specified duration, obtain multiple types of target operating parameters in the distributed system, where the number of target operating parameters of any one type is at least one;

[0119] Step 403: For any one type of target operating parameter, compare the target operating parameter with the target fault parameter range;

[0120] Step 404: Determine whether there is a target operating parameter in the target operating parameters that is within the target fault parameter range. If so, execute Step 405; if not, execute Step 406;

[0121] Step 405: Determine that there is a target fault corresponding to the type in the distributed system, and determine the service corresponding to the target operating parameter within the target fault parameter range as the target fault;

[0122] Step 406: Determine that there is no target fault corresponding to the type in the distributed system;

[0123] Step 407: Use the pre-set correspondence between the fault type and the fault isolation strategy to determine the target fault isolation strategy corresponding to the type;

[0124] Step 408: Determine whether the number of target faults corresponding to the type is multiple. If so, execute Step 409; if not, execute Step 410;

[0125] Step 409: Obtain the priorities among the target faults based on the preset priorities among various services, and use the target fault isolation policy to isolate the target faults in sequence according to the priorities among the target faults, so as to cut off the target services corresponding to the target faults in the distributed system;

[0126] Step 410: Use the target fault isolation policy to isolate the target fault, so as to cut off the target service corresponding to the target fault in the distributed system.

[0127] Based on the same inventive concept, the fault isolation method of the distributed system as described above in the present disclosure can also be implemented by a fault isolation device of a distributed system. The effect of the fault isolation device of the distributed system is similar to that of the foregoing method, and will not be elaborated here.

[0128] Figure 5 It is a schematic structural diagram of a fault isolation device of a distributed system according to an embodiment of the present disclosure.

[0129] As Figure 5 shown, the fault isolation device 500 of the distributed system of the present disclosure may include an acquisition module 510, a judgment module 520, an isolation policy determination module 530, and a fault isolation module 540.

[0130] The acquisition module 510 is configured to obtain multiple types of target operation parameters in the distributed system at specified time intervals, where the number of any type of target operation parameter is at least one;

[0131] The judgment module 520 is configured to, for any type of target operation parameter, determine whether there is a target fault corresponding to the type in the distributed system based on the target operation parameter and the target fault parameter range, where the target fault parameter range is a preset fault parameter range corresponding to the type; and,

[0132] The isolation policy determination module 530 is configured to, if any, determine a target fault isolation policy corresponding to the type by using the preset correspondence between the fault type and the fault isolation policy; and,

[0133] The fault isolation module 540 is configured to isolate the target fault corresponding to the type by using the target fault isolation policy.

[0134] In one embodiment, the judgment module 520 is specifically configured to:

[0135] Compare the target operation parameter with the target fault parameter range;

[0136] If there is a target operating parameter within the target fault parameter range among the target operating parameters, it is determined that there is a target fault corresponding to the type in the distributed system, and the service corresponding to the target operating parameter within the target fault parameter range is determined as the target fault;

[0137] If there is no target operating parameter within the target fault parameter range among the target operating parameters, it is determined that there is no target fault corresponding to the type in the distributed system.

[0138] In one embodiment, the type includes server type, database type, network dedicated line type, data center type, and city center type;

[0139] The target fault isolation policy includes at least one of a server fault isolation policy, a database fault isolation policy, a network dedicated line fault isolation policy, a data center fault isolation policy, and a city fault isolation policy;

[0140] The isolation policy determination module 530 is specifically configured to:

[0141] If the type is the server type, it is determined that the target fault isolation policy corresponding to the type is the server fault isolation policy; or,

[0142] If the type is the database type, it is determined that the target fault isolation policy corresponding to the type is the database fault isolation policy; or,

[0143] If the type is the network dedicated line type, it is determined that the target fault isolation policy corresponding to the type is the network dedicated line fault isolation policy; or,

[0144] If the type is the data center type, it is determined that the target fault isolation policy corresponding to the type is the data center fault isolation policy; or,

[0145] If the type is the city center type, it is determined that the target fault isolation policy corresponding to the type is the city fault isolation policy.

[0146] In one embodiment, the fault isolation module 540 is specifically configured to:

[0147] If the number of target faults corresponding to the type is one, the target fault is isolated using the target fault isolation policy; or,

[0148] If the number of target faults corresponding to the type is multiple, then based on the priorities among the pre-set services, obtain the priorities among the target faults, and use the target fault isolation strategy to isolate the target faults in sequence according to the priorities among the target faults.

[0149] In one embodiment, the apparatus further includes:

[0150] A specified duration setting module 550, configured to, before obtaining multiple types of target operating parameters in the distributed system every specified duration, in response to a specified duration setting instruction sent by a user, set the specified duration based on the specified duration setting instruction.

[0151] After introducing a method and apparatus for fault isolation of a distributed system according to an exemplary embodiment of the present disclosure, next, an electronic device according to another exemplary embodiment of the present disclosure will be introduced.

[0152] Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, method, or program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.

[0153] In some possible implementation manners, the electronic device according to the present disclosure may at least include at least one processor and at least one computer storage medium. Among them, the computer storage medium stores program code, and when the program code is executed by the processor, the processor executes the steps in the method for fault isolation of a distributed system according to various exemplary embodiments of the present disclosure described above in this specification. For example, the processor may execute steps 201-204 as shown in Figure 2 shown.

[0154] Next, refer to Figure 6 to describe the electronic device 600 according to this embodiment of the present disclosure. Figure 6 The shown electronic device 600 is only an example, and should not bring any limitation to the functions and usage scopes of the embodiments of the present disclosure.

[0155] As Figure 6 shown, the electronic device 600 is presented in the form of a general electronic device. The components of the electronic device 600 may include but are not limited to: the above at least one processor 601, the above at least one computer storage medium 602, and a bus 603 connecting different system components (including the computer storage medium 602 and the processor 601).

[0156] The bus 603 represents one or more of several types of bus architectures, including a computer storage media bus or a computer storage media controller, a peripheral bus, a processor bus, or a local bus using any of the various bus architectures.

[0157] The computer storage media 602 may include a readable medium in the form of volatile computer storage media, such as random access computer storage media (RAM) 621 and / or cache storage media 622, and may further include read-only computer storage media (ROM) 623.

[0158] The computer storage media 602 may also include a program / utility 625 having a set (at least one) of program modules 624. Such program modules 624 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.

[0159] The electronic device 600 may also communicate with one or more external devices 604 (such as a keyboard, a pointing device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 600, and / or may communicate with any device that enables the electronic device 600 to communicate with one or more other electronic devices (such as a router, a modem, etc.). Such communication may be performed through an input / output (I / O) interface 605. Further, the electronic device 600 may communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 606. As shown in the figure, the network adapter 606 communicates with other modules for the electronic device 600 through the bus 603. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 600, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0160] In some possible embodiments, various aspects of a fault isolation method for a distributed system provided by the present disclosure may also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps in the fault isolation method for a distributed system according to various exemplary embodiments of the present disclosure described above in this specification.

[0161] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access computer storage medium (RAM), a read-only computer storage medium (ROM), an erasable programmable read-only computer storage medium (EPROM or flash memory), an optical fiber, a portable compact disk read-only computer storage medium (CD-ROM), an optical computer storage medium, a magnetic computer storage medium, or any suitable combination of the above.

[0162] The program product for fault isolation of the distributed system according to the embodiments of the present disclosure can adopt a portable compact disk read-only computer storage medium (CD-ROM) and include program code, and can run on an electronic device. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0163] The readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0164] The program code contained on the readable medium can be transmitted by any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.

[0165] The program code for performing the operations of the present disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's electronic device, partially on the user's device, executed as a stand-alone software package, partially on the user's electronic device and partially on a remote electronic device, or entirely on a remote electronic device or server. In cases involving a remote electronic device, the remote electronic device can be connected to the user's electronic device through any type of network including a local area network (LAN) or a wide area network (WAN), or, it can be connected to an external electronic device (e.g., by using an Internet service provider to connect through the Internet).

[0166] It should be noted that although several modules of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-described modules can be embodied in one module. Conversely, the features and functions of one module described above can be further divided and embodied by multiple modules.

[0167] In addition, although the operations of the method of the present disclosure are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the shown operations must be performed to achieve the desired result. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.

[0168] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic computer storage media, CD-ROM, optical computer storage media, etc.) containing computer-usable program code.

[0169] This disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to the disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in a block or multiple blocks.

[0170] These computer program instructions can also be stored in a computer-readable computer storage medium that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable computer storage medium produce a manufactured article including instruction means that implement the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in a block or multiple blocks.

[0171] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in a block or multiple blocks.

[0172] Obviously, those skilled in the art can make various changes and modifications to this disclosure without departing from the spirit and scope of this disclosure. Thus, if these modifications and variations of this disclosure fall within the scope of the claims of this disclosure and their equivalent technologies, this disclosure is also intended to include these changes and modifications.

Claims

1. A fault isolation method for a distributed system, characterized in that The method includes: Obtaining target operation parameters of multiple types in a distributed system at every specified time interval, where the number of target operation parameters of any one type is at least one; For any one type of target operation parameter, determining whether there is a target fault corresponding to the type in the distributed system based on the target operation parameter and a target fault parameter range, where the target fault parameter range is a pre-set fault parameter range corresponding to the type; and If there is, determining a target fault isolation strategy corresponding to the type by using the pre-set correspondence between fault types and fault isolation strategies; and Using the target fault isolation strategy to isolate the target fault corresponding to the type, so as to cut off the target service corresponding to the target fault in the distributed system.

2. The method according to claim 1, wherein The determining whether there is a target fault corresponding to the type in the distributed system based on the target operation parameter and the target fault parameter range includes: Comparing the target operation parameter with the target fault parameter range; If there is a target operation parameter among the target operation parameters that is within the target fault parameter range, determining that there is a target fault corresponding to the type in the distributed system, and determining the service corresponding to the target operation parameter within the target fault parameter range as the target fault; If there is no target operation parameter among the target operation parameters that is within the target fault parameter range, determining that there is no target fault corresponding to the type in the distributed system.

3. The method according to claim 1 or 2, characterized in that, The type includes server type, database type, network dedicated line type, data center type, and city center type; The target fault isolation strategy includes at least one of a server fault isolation strategy, a database fault isolation strategy, a network dedicated line fault isolation strategy, a data center fault isolation strategy, and a city fault isolation strategy; The determining a target fault isolation strategy corresponding to the type by using the pre-set correspondence between fault types and fault isolation strategies includes: If the type is a server type, determining that the target fault isolation strategy corresponding to the type is a server fault isolation strategy; or If the type is a database type, determining that the target fault isolation strategy corresponding to the type is a database fault isolation strategy; or If the type is a network dedicated line type, determining that the target fault isolation strategy corresponding to the type is a network dedicated line fault isolation strategy; or If the type is a data center type, determining that the target fault isolation strategy corresponding to the type is a data center fault isolation strategy; or If the type is a city center type, determining that the target fault isolation strategy corresponding to the type is a city fault isolation strategy.

4. The method according to claim 1, wherein The isolating the target fault corresponding to the type by using the target fault isolation strategy includes: If the number of target faults corresponding to the type is one, isolating the target fault by using the target fault isolation strategy; or If the number of target faults corresponding to the type is multiple, then based on the priorities among the pre-set services, obtain the priorities among the target faults, and use the target fault isolation policy to isolate the target faults in sequence according to the priorities among the target faults.

5. The method according to claim 1, wherein Before obtaining multiple types of target operating parameters in the distributed system at each specified time interval, the method further includes: In response to a specified time interval setting instruction sent by the user, set the specified time interval based on the specified time interval setting instruction.

6. A fault isolation device for a distributed system, characterized in that, The device includes: An acquisition module, configured to obtain multiple types of target operating parameters in the distributed system at each specified time interval, where the number of target operating parameters of any one type is at least one; A target fault determination module, configured to, for any one type of target operating parameter, determine whether there is a target fault corresponding to the type in the distributed system based on the target operating parameter and the target fault parameter range, where the target fault parameter range is a pre-set fault parameter range corresponding to the type; and An isolation policy determination module, configured to, if any, use the pre-set correspondence between the fault type and the fault isolation policy to determine the target fault isolation policy corresponding to the type; and A fault isolation module, configured to use the target fault isolation policy to perform fault isolation on the target fault corresponding to the type, so as to cut off the target service corresponding to the target fault in the distributed system.

7. The device according to claim 6, characterized in that, The target fault determination module is specifically configured to: Compare the target operating parameter with the target fault parameter range; If there is a target operating parameter among the target operating parameters that is within the target fault parameter range, determine that there is a target fault corresponding to the type in the distributed system, and determine the service corresponding to the target operating parameter within the target fault parameter range as the target fault; If there is no target operating parameter among the target operating parameters that is within the target fault parameter range, determine that there is no target fault corresponding to the type in the distributed system.

8. The device according to claim 6 or 7, characterized in that, The type includes server type, database type, network dedicated line type, data center type, and urban center type; The target fault isolation policy includes at least one of a server fault isolation policy, a database fault isolation policy, a network dedicated line fault isolation policy, a data center fault isolation policy, and an urban fault isolation policy; The isolation policy determination module is specifically configured to: If the type is the server type, determine that the target fault isolation policy corresponding to the type is the server fault isolation policy; or If the type is the database type, determine that the target fault isolation policy corresponding to the type is the database fault isolation policy; Or If the type is the network dedicated line type, determine that the target fault isolation policy corresponding to the type is the network dedicated line fault isolation policy; Or If the type is the data center type, determine that the target fault isolation policy corresponding to the type is the data center fault isolation policy; Or If the type is the urban center type, determine that the target fault isolation strategy corresponding to the type is the urban fault isolation strategy.

9. An electronic device, characterized in that, Comprising at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor; the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-5.

10. A computer storage medium, characterized in that, The computer storage medium stores a computer program for executing the method according to any one of claims 1-5.