Device failure restart method, device, electronic device and storage medium

Through the computer's resource utilization rate and group restart strategy, the problem of restart strategy differences caused by system diversity is solved, and rapid and business-free failure recovery is achieved.

CN114328019BActive Publication Date: 2025-09-02CHINA CONSTRUCTION BANK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111627530.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-28
Publication Date
2025-09-02
Estimated Expiration
2041-12-28

AI Technical Summary

Technical Problem

In the prior art, system restart strategies due to system diversity vary greatly, and artificially orchestrating the restart process leads to problems such as long recovery time of production systems and wide business impact.

Method used

By obtaining the total number of machines in the cluster to which the alarm machine belongs and the maximum utilization rate of each target resource for each machine, calculate the total utilization rate of the target resource and obtain the quotient calculation, obtain the number of machine requirements and number of packets, and restart each group of machines in turn.

Benefits of technology

While ensuring resource requirements, the impact of restarting on business services is reduced and the recovery efficiency of the production system is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114328019B_ABST
    Figure CN114328019B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, electronic device and storage medium for restarting a device upon failure, which comprises the following steps: obtaining the total number of machines in a cluster to which an alarm machine belongs and the maximum utilization rate of each target resource of each machine; calculating, for each target resource, the sum of the maximum utilization rates of the target resources of each machine to obtain the total utilization rate corresponding to the target resource; performing a quotient operation on the total utilization rate corresponding to each target resource and a target value to obtain an integer quotient corresponding to each target resource; the target value is 1; calculating the sum of the maximum value of the integer quotients corresponding to each target resource and the target value to obtain the required number of machines; dividing the ratio of the number of machines by the integer quotient of the target value plus the target value to obtain the number of groups; the ratio of the number of machines is equal to the total number of machines divided by the difference between the total number of machines and the required number of machines; grouping the machines according to the number of groups; and restarting the machines in each group in turn.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and more particularly to a method and apparatus for restarting a device after a fault, an electronic device, and a storage medium. Background Art

[0002] With the promotion of banking business, system diversity is increasing, and the system is becoming more and more complex. The failure scenarios that occur in the operation of production systems are also diverse.

[0003] In existing technology, restarting is generally used to resolve most failures that occur during production operations. However, the diversity of systems leads to widely varying restart strategies. To ensure uninterrupted external business operations, various machines in the system must be shut down and started in a rotating manner. This requires manual orchestration of the restart process before a restart can be performed. This waiting period inevitably increases the time it takes for production systems to return to normal, resulting in longer and more widespread business disruptions. Summary of the Invention

[0004] In view of this, the present invention provides a method, apparatus, electronic device and storage medium for restarting a device upon failure, so as to solve the problem that the existing manual restart scheduling has a significant impact on the business.

[0005] A first aspect of the present application provides a method for restarting a device after a fault, the method comprising:

[0006] Obtain the total number of machines in the cluster to which the alarm machine belongs and the maximum usage rate of each target resource of each machine;

[0007] For each target resource, calculate the sum of the highest utilization rates of the target resource of each machine to obtain the total utilization rate corresponding to the target resource;

[0008] Taking the quotient of the total usage rate corresponding to each target resource and the target value to obtain an integer quotient corresponding to each target resource; wherein the target value is 1;

[0009] Calculate the sum of the maximum value of the integer quotients corresponding to each target resource and the target value to obtain the required number of machines;

[0010] The number of groups is obtained by dividing the ratio of the number of machines by the target value and adding the target value to the integer quotient; wherein the ratio of the number of machines is equal to the quotient obtained by dividing the total number of the machines by the difference between the total number of the machines and the required number of machines;

[0011] Grouping each of the machines according to the number of groups;

[0012] Restart each of the machines in each group in turn.

[0013] Optionally, in the device fault restart method provided above, before obtaining the total number of machines in the cluster to which the alarm machine belongs and the maximum usage rate of each target resource of each machine, the method further includes:

[0014] Matching the deployment unit corresponding to the alarm machine and obtaining information about the deployment unit corresponding to the alarm machine;

[0015] Determining whether the deployment unit belongs to a cluster based on the cluster tag in the information of the deployment unit corresponding to the alarm machine;

[0016] If it is determined that the deployment unit belongs to a cluster, the step of obtaining the total number of machines in the cluster to which the alarm machine belongs and the maximum usage rate of each target resource of each machine is executed;

[0017] If it is determined that the deployment unit does not belong to a cluster, target prompt information is fed back; wherein the target prompt information is used to prompt the user to perform manual intervention;

[0018] The step of obtaining the total number of machines in the cluster to which the alarm machine belongs and the maximum usage rate of each target resource of each machine includes:

[0019] Obtain the total number of machines under the deployment unit corresponding to the alarm machine and the maximum usage rate of each target resource of each machine; wherein, each machine in the cluster to which the alarm machine belongs constitutes the deployment unit corresponding to the alarm machine.

[0020] Optionally, in the device fault restart method provided above, before determining whether the deployment unit belongs to a cluster based on the cluster tag in the information of the deployment unit corresponding to the alarm machine, the method further includes:

[0021] Determining whether the deployment unit is related to Netlink according to the Netlink-related tag in the information of the deployment unit;

[0022] If it is determined that the deployment unit is not related to the network connection, executing the cluster tag in the information of the deployment unit corresponding to the alarm machine to determine whether the deployment unit belongs to a cluster;

[0023] If it is determined that the deployment unit is related to the network connection, the target prompt information is fed back.

[0024] Optionally, in the device fault restart method provided above, obtaining the total number of machines in the cluster to which the alarm machine belongs and the maximum usage rate of each target resource of each machine includes:

[0025] Get the total number of machines in the cluster to which the alarm machine belongs, as well as the maximum CPU usage and memory usage of each machine;

[0026] Determine whether there is a special resource indicator identifier in the information of the deployment unit corresponding to the alarm machine;

[0027] If it is determined that a special resource indicator identifier exists in the information of the deployment unit corresponding to the alarm machine, the highest usage rate of the resource corresponding to the special resource indicator identifier of each of the machines is obtained.

[0028] A second aspect of the present application provides a device for restarting a device after a fault, comprising:

[0029] an acquisition unit, configured to acquire the total number of machines in the cluster to which the alarm machine belongs and the maximum usage rate of each target resource of each machine;

[0030] a summation calculation unit, configured to calculate, for each target resource, the sum of the highest utilization rates of the target resource of each of the machines, to obtain a total utilization rate corresponding to the target resource;

[0031] a quotient calculation unit, configured to perform a quotient calculation on the total usage rate corresponding to each target resource and the target value to obtain an integer quotient corresponding to each target resource; wherein the target value is 1;

[0032] a machine quantity determination unit, configured to calculate the sum of the maximum integer quotient corresponding to each target resource and the target value to obtain the required number of machines;

[0033] a grouping quantity determination unit, configured to obtain the grouping quantity by adding the target value to the integer quotient of the machine quantity ratio divided by the target value, wherein the machine quantity ratio is equal to the quotient obtained by dividing the total quantity of each of the machines by the difference between the total quantity of each of the machines and the required quantity of machines;

[0034] A grouping unit, configured to group the machines according to the grouping quantity;

[0035] The restart unit is used to restart each of the machines in each group in turn.

[0036] Optionally, the fault restart device provided above further includes:

[0037] A matching unit, configured to match a deployment unit corresponding to the alarm machine and obtain information about the deployment unit corresponding to the alarm machine;

[0038] A first determining unit is configured to determine, based on a cluster tag in the information of the deployment unit corresponding to the alarm machine, whether the deployment unit belongs to a cluster; wherein, if it is determined that the deployment unit belongs to a cluster, the obtaining unit executes the step of obtaining the total number of machines in the cluster to which the alarm machine belongs and the maximum usage rate of each target resource of each machine;

[0039] A first feedback unit is configured to feedback target prompt information if it is determined that the deployment unit does not belong to a cluster; wherein the target prompt information is used to prompt a user to perform manual intervention;

[0040] Wherein, the acquisition unit includes:

[0041] An acquisition subunit is used to obtain the total number of machines under the deployment unit corresponding to the alarm machine and the maximum utilization rate of each target resource of each machine; wherein, each machine in the cluster to which the alarm machine belongs constitutes the deployment unit corresponding to the alarm machine.

[0042] Optionally, the fault restart device provided above further includes:

[0043] a second judgment unit, configured to judge whether the deployment unit is related to the network connection according to the network connection related tag in the information of the deployment unit; wherein, if it is judged that the deployment unit is not related to the network connection, the first judgment unit executes the judgment based on the cluster tag in the information of the deployment unit corresponding to the alarm machine to determine whether the deployment unit belongs to a cluster;

[0044] The second feedback unit is used to feedback the target prompt information if it is determined that the deployment unit is related to the network connection.

[0045] Optionally, in the device for restarting a device after a fault is provided above, the acquiring unit includes:

[0046] The first acquisition unit is used to obtain the total number of machines in the cluster to which the alarm machine belongs and the maximum CPU usage and memory usage of each machine;

[0047] A third judgment unit is used to judge whether there is a special resource indicator identifier in the information of the deployment unit corresponding to the alarm machine;

[0048] The second obtaining unit is configured to obtain the maximum usage rate of the resource corresponding to the special resource indicator identifier of each of the machines if it is determined that the information of the deployment unit corresponding to the alarm machine contains a special resource indicator identifier.

[0049] A third aspect of the present application provides an electronic device, including:

[0050] A processor and a memory, wherein the memory is used to store program code and data for data update, and the processor is used to call program instructions in the memory to execute a fault restart method for a device as described in any one of the above.

[0051] A fourth aspect of the present application provides a storage medium, which includes a stored program, wherein when the program is running, the device where the storage medium is located is controlled to execute a device failure restart method as described in any one of the above.

[0052] The present application provides a method for restarting a device after a fault. The method obtains the total number of machines in the cluster to which the alarm machine belongs and the maximum utilization rate of each target resource of each machine. For each target resource, the sum of the maximum utilization rates of the target resources of each machine is calculated to obtain the total utilization rate corresponding to the target resource. The total utilization rate corresponding to each target resource is then divided by the target value to obtain the integer quotient corresponding to each target resource. The maximum value of the integer quotients calculated for each target resource is added to the target value to obtain the required number of machines, thereby obtaining the number of machines required to provide the maximum target resource requirement. The ratio of the number of machines is then divided by the integer quotient of the target value and added to the target value to obtain the number of groups. The ratio of the number of machines is equal to the quotient obtained by dividing the total number of machines by the difference between the total number of machines and the required number of machines. This method can thus obtain the number of groups required to ensure that the number of machines that have not been restarted is the required number of machines. Finally, the machines are grouped according to the number of groups, and each machine in each group is restarted in turn. This allows the machines to be restarted while ensuring resource requirements and without affecting the cluster's external business services. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0054] Figure 1 A flowchart of a method for restarting a device after a failure provided by an embodiment of the present application;

[0055] Figure 2 A flowchart of a method for obtaining the number of machines and the utilization rate of target resources of the machines provided in an embodiment of the present application;

[0056] Figure 3 A flowchart of another method for restarting a device after a failure provided by another embodiment of the present application;

[0057] Figure 4 A schematic diagram of the structure of a fault restart device for a device provided in an embodiment of the present application;

[0058] Figure 5 A schematic structural diagram of an acquisition unit provided in another embodiment of the present application;

[0059] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0060] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0061] In this application, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.

[0062] The present application embodiment provides a method for restarting a device due to a fault, such as Figure 1 As shown, the specific steps include:

[0063] S101: Obtain the total number of machines in the cluster to which the alarm machine belongs and the maximum usage rate of each target resource of each machine.

[0064] The alarm machine refers to a machine that generates a fault alarm, such as a server, etc. When a machine generates a fault alarm, it automatically enters the restart processing decision, that is, executes the method provided by the embodiment of the present application.

[0065] To ensure high availability and load balancing, a service is usually provided by a cluster of multiple machines. Therefore, the cluster to which the alarm machine belongs is composed of multiple machines that share the same service with the alarm machine.

[0066] It should be noted that target resources are designated resources, typically those required and highly utilized by the alarming machine to provide the corresponding service, such as CPU resources and memory resources. By setting the number of machines in the cluster and the maximum utilization of the target resources for each machine, the remaining machines in the cluster can ensure the effective provision of target resources during a restart, thereby ensuring that the services provided by the alarming machine can be provided normally after the restart.

[0067] Specifically, the obtained maximum utilization rate of each target resource of each machine may be the maximum utilization rate of each target resource of the machine recorded on the current machine within a period of time.

[0068] Optionally, when the alarm machine does not belong to a cluster, it cannot be restarted directly, otherwise the services it provides will be stopped. Therefore, more complex operations such as service transfer are required, which often require manual operation. Therefore, when the alarm machine does not belong to a cluster, a prompt message can be issued to the user, prompting the user to perform manual intervention.

[0069] Optionally, in the embodiment of the present application, a specific implementation of step S101 is as follows: Figure 2 As shown, the following steps are included:

[0070] S201: Obtain the total number of machines in the cluster to which the alarm machine belongs, as well as the maximum CPU usage and the maximum memory usage of each machine.

[0071] In the embodiment of the present application, the target resources mainly include CPU resources and memory resources, and of course may also include other resources, not limited to CPU resources and memory resources.

[0072] S202: Determine whether there is a special resource indicator identifier in the information of the deployment unit corresponding to the alarm machine.

[0073] It should be noted that when the alarming machine belongs to a cluster, its cluster constitutes the corresponding deployment unit, and the deployment unit information includes multiple pieces of information about the cluster. Machines providing special services may rely on special resources, and in this case, the deployment unit information will store the identifier of the special resource. Therefore, it is necessary to first determine whether the special resource identifier exists in the deployment unit information corresponding to the alarming machine. If so, the resource corresponding to the special resource identifier should also be the target resource, so step S203 must be further executed.

[0074] S203: Obtain the highest resource usage rate corresponding to the special resource indicator identifier of each machine.

[0075] S102 : For each target resource, calculate the sum of the highest utilization rates of the target resources of each machine to obtain the total utilization rate corresponding to the target resource.

[0076] By calculating the sum of the maximum utilization rates of the target resources of each machine, we can obtain the highest total utilization rate of each target resource when the cluster provides services, and then restart the cluster while ensuring the total utilization rate.

[0077] S103 , performing a quotient operation on the total usage rate corresponding to each target resource and the target value to obtain an integer quotient corresponding to each target resource.

[0078] Among them, the target value is 1.

[0079] It should be noted that taking the quotient of the total utilization rate corresponding to the target resource and the target value refers to the integer quotient obtained by dividing the total utilization rate corresponding to the target resource by the target value, that is, the remainder is not considered, and it can also be considered as the digits before the decimal point of the calculation result.

[0080] Since a target resource for each machine is 100%, that is, 1, the target value is 1, so that the number of machines required to provide the total utilization corresponding to the target resource can be determined.

[0081] S104. Calculate the sum of the maximum value of the integer quotients corresponding to each target resource and the target value to obtain the required number of machines.

[0082] It should be noted that when providing services, it is obviously necessary to meet the demand for each target resource. When the target resource with the maximum demand is met, the demand for all target resources can definitely be guaranteed. Therefore, the maximum value is determined from the integer quotients corresponding to each target resource, and then the calculation is performed.

[0083] Specifically, the total utilization rate for each target resource is divided by 1 to obtain the number of machines required to meet the target resource's total utilization rate. Considering the possibility of a remainder, an additional 1 is added to ensure the target resource's demand. This yields the minimum number of machines required to meet the target resource's demand. This is the required number of machines.

[0084] S105 . Divide the ratio of the number of machines by the target value, and add the integer quotient thereof to the target value to obtain the number of groups.

[0085] The machine quantity ratio is equal to the quotient obtained by dividing the total quantity of each machine by the difference between the total quantity of each machine and the required quantity of machines.

[0086] It's important to note that, because the number of machines not yet restarted must be at least the required number of machines, the maximum number of machines restarted at a time—that is, the number of machines in a group—can only be the difference between the total number of machines and the required number of machines. Therefore, if the number is divisible evenly, that is, when the ratio of the number of machines is an integer, the number of groups is the quotient of the total number of machines divided by the difference between the total number of machines and the required number of machines. This is the integer quotient of the ratio of the number of machines divided by the target value. However, given that the number is not divisible evenly in most cases, an additional group is added, with the number of machines in the final group being less than the difference between the total number of machines and the required number of machines.

[0087] Of course, you can also first determine whether the ratio of the number of machines is an integer. If it is an integer, you do not need to add the target value. In this case, the number of machines in each group is the difference between the total number of machines and the required number of machines. If the remainder is not zero, add the target value.

[0088] S106: Group the machines according to the number of groups.

[0089] It should be noted that when grouping, it is not only necessary to divide each machine into groups of the required number, but also the number of machines in each group should not exceed the difference between the total number of each machine and the required number of machines to ensure.

[0090] Alternatively, assuming the number of groups is N, the number of machines in each of the N-1 groups can be the difference between the total number of machines and the required number of machines, and one group can be the remaining machines. Of course, grouping can also be done using strategies.

[0091] S107: Restart each machine in each group in turn.

[0092] Specifically, for each divided group, the machines in the group are restarted simultaneously.

[0093] The embodiment of the present application provides a method for restarting a device after a fault, by obtaining the total number of machines in the cluster to which the alarm machine belongs and the maximum utilization rate of each target resource of each machine, and calculating the sum of the maximum utilization rates of the target resources of each machine for each target resource, to obtain the total utilization rate corresponding to the target resource, then taking the quotient of the total utilization rate corresponding to each target resource and the target value, to obtain the integer quotient corresponding to each target resource, and summing the maximum value of the integer quotients corresponding to each target resource with the target value to obtain the required number of machines, thereby obtaining the machines required to provide the maximum target resource demand. Then, the ratio of the number of machines is divided by the integer quotient of the target value and added to the target value to obtain the number of groups. The ratio of the number of machines is equal to the quotient obtained by dividing the total number of machines by the difference between the total number of machines and the required number of machines, thereby obtaining the number of groups under the condition that the number of machines that have not been restarted is the required number of machines. Finally, the machines are grouped according to the number of groups, and each machine in each group is restarted in turn, so that the machines can be restarted under the premise of ensuring resource demand and not affecting the external business services of the cluster.

[0094] Another embodiment of the present application provides another method for restarting a device after a failure, such as Figure 3 As shown, the specific steps include:

[0095] S301: Match the deployment unit corresponding to the alarm machine and obtain information about the deployment unit corresponding to the alarm machine.

[0096] It should be noted that each machine corresponds to a deployment unit, and a deployment unit can correspond to multiple machines. That is, a deployment unit can be composed of one machine or multiple machines to achieve availability.

[0097] Therefore, in the embodiment of the present application, the deployment unit corresponding to the alarm machine is first matched to obtain the information of the deployment unit corresponding to the alarm machine to determine whether the alarm machine has a cluster to which it belongs and whether it is network-related.

[0098] S302: Determine whether the deployment unit is related to the network connection according to the network connection related tag in the information of the deployment unit.

[0099] It should be noted that China National Clearing Corporation (CNS) is a platform managed by an organization and connected to the clearing platforms of various banks. Therefore, deployment units related to CNS cannot be restarted at will. Manual intervention is required to avoid problems. Therefore, if it is necessary to determine whether the deployment unit is related to CNS, if so, step S303 is executed. If not, the subsequent steps can be continued, that is, step S304 is executed.

[0100] S303: Feedback target prompt information.

[0101] The target prompt information is used to prompt the user to perform manual intervention.

[0102] S304: Determine whether the deployment unit belongs to a cluster based on the cluster tag in the information of the deployment unit corresponding to the alarm machine.

[0103] It's important to note that if a deployment unit consisting of a single machine is directly restarted, the service it provides will be interrupted because there are no other machines to back it up. This often requires complex operations such as service migration, which requires manual intervention. Therefore, it's necessary to first determine whether the deployment unit belongs to a cluster, that is, whether the alarming machine exists in a cluster.

[0104] If it is determined that the deployment unit does not belong to a cluster, step S303 is executed to feedback target prompt information. If it is determined that the deployment unit belongs to a cluster, step S305 is executed.

[0105] S305: Obtain the total number of machines under the deployment unit corresponding to the alarm machine and the maximum usage rate of each target resource of each machine.

[0106] Among them, each machine in the cluster to which the alarm machine belongs constitutes the deployment unit corresponding to the alarm machine, so obtaining the total number of machines under the deployment unit corresponding to the alarm machine and the maximum utilization rate of each target resource of each machine is to obtain the total number of machines in the cluster to which the alarm machine belongs and the maximum utilization rate of each target resource of each machine.

[0107] It should be noted that, for a more specific implementation of step S305, reference may be made to step S101 in the above method embodiment, which will not be described in detail here.

[0108] S306 : For each target resource, calculate the sum of the highest utilization rates of the target resources of each machine to obtain the total utilization rate corresponding to the target resource.

[0109] It should be noted that the specific implementation of step S306 may refer to step S102 in the above method embodiment, and will not be repeated here.

[0110] S307 , performing a quotient operation on the total usage rate corresponding to each target resource and the target value to obtain an integer quotient corresponding to each target resource.

[0111] Among them, the target value is 1.

[0112] It should be noted that the specific implementation of step S307 may refer to step S103 in the above method embodiment, and will not be repeated here.

[0113] S308. Calculate the sum of the maximum value of the integer quotients corresponding to each target resource and the target value to obtain the required number of machines.

[0114] It should be noted that the specific implementation of step S308 may refer to step S104 in the above method embodiment, and will not be repeated here.

[0115] S309: Divide the ratio of the number of machines by the target value, and add the integer quotient to the target value to obtain the number of groups.

[0116] The machine quantity ratio is equal to the quotient obtained by dividing the total quantity of each machine by the difference between the total quantity of each machine and the required quantity of machines.

[0117] It should be noted that the specific implementation of step S309 may refer to step S105 in the above method embodiment, and will not be repeated here.

[0118] S310: Group the machines according to the number of groups.

[0119] It should be noted that the specific implementation of step S310 may refer to step S106 in the above method embodiment, and will not be repeated here.

[0120] S311. Restart each machine in each group in turn.

[0121] It should be noted that the specific implementation of step S311 may refer to step S107 in the above method embodiment, and will not be repeated here.

[0122] Another embodiment of the present application provides a device for restarting a device after a failure, such as Figure 4 As shown, it includes the following units:

[0123] The acquisition unit 401 is configured to acquire the total number of machines in the cluster to which the alarm machine belongs and the maximum usage rate of each target resource of each machine.

[0124] The sum calculation unit 402 is used to calculate the sum of the highest utilization rates of the target resources of each machine for each target resource, and obtain the total utilization rate corresponding to the target resource.

[0125] The quotient calculation unit 403 is configured to perform a quotient calculation on the total usage rate corresponding to each target resource and the target value to obtain an integer quotient corresponding to each target resource.

[0126] Among them, the target value is 1.

[0127] The machine quantity determination unit 404 is configured to calculate the sum of the maximum value of the integer quotients corresponding to each target resource and the target value to obtain the required number of machines.

[0128] The group quantity determining unit 405 is configured to add the target value to the integer quotient of the ratio of the number of machines divided by the target value to obtain the number of groups.

[0129] The machine quantity ratio is equal to the quotient obtained by dividing the total quantity of each machine by the difference between the total quantity of each machine and the required quantity of machines.

[0130] The grouping unit 406 is configured to group the machines according to the group quantity.

[0131] The restart unit 407 is configured to restart each machine in each group in sequence.

[0132] Optionally, in the device failure restarting apparatus provided in another embodiment of the present application, the apparatus further includes:

[0133] The matching unit is used to match the deployment unit corresponding to the alarm machine and obtain the information of the deployment unit corresponding to the alarm machine.

[0134] The first determination unit is configured to determine, based on a cluster tag in the deployment unit information corresponding to the alarming machine, whether the deployment unit belongs to a cluster. If the deployment unit is determined to belong to a cluster, the acquisition unit acquires the total number of machines in the cluster to which the alarming machine belongs and the maximum utilization rate of each target resource of each machine.

[0135] The first feedback unit is configured to feed back target prompt information if it is determined that the deployment unit does not belong to a cluster.

[0136] The target prompt information is used to prompt the user to perform manual intervention.

[0137] The acquisition unit in the embodiment of the present application includes:

[0138] The acquisition subunit is used to obtain the total number of machines under the deployment unit corresponding to the alarm machine and the maximum utilization rate of each target resource of each machine. Among them, each machine in the cluster to which the alarm machine belongs constitutes the deployment unit corresponding to the alarm machine.

[0139] Optionally, in the device failure restarting apparatus provided in another embodiment of the present application, the apparatus further includes:

[0140] The second judgment unit is configured to determine whether the deployment unit is related to the network connection based on the network connection related tag in the deployment unit information. If it is determined that the deployment unit is not related to the network connection, the first judgment unit then determines whether the deployment unit belongs to a cluster based on the cluster tag in the deployment unit information corresponding to the alarming machine.

[0141] The second feedback unit is used to feedback target prompt information if it is determined that the deployment unit is related to the network connection.

[0142] Optionally, in the fault restart device of another embodiment of the present application, the acquisition unit, such as Figure 5 Shown, including:

[0143] The first obtaining unit 501 is configured to obtain the total number of machines in the cluster to which the alarm machine belongs and the maximum CPU usage and the maximum memory usage of each machine.

[0144] The third judgment unit 502 is used to judge whether there is a special resource indicator identifier in the information of the deployment unit corresponding to the alarm machine.

[0145] The second obtaining unit 503 is configured to obtain the maximum usage rate of the resource corresponding to the special resource indicator identifier of each machine if it is determined that the information of the deployment unit corresponding to the alarm machine contains a special resource indicator identifier.

[0146] It should be noted that the specific working process of each unit provided in the above embodiments of the present application can refer to the corresponding steps in the above method embodiments, and will not be repeated here.

[0147] Another embodiment of the present application provides an electronic device, such as Figure 6 Shown, including:

[0148] Processor 601 and memory 602 .

[0149] The memory 602 is used to store program codes and data for data updates, and the processor 601 is used to call program instructions in the memory 602 to execute a fault restart method for a device provided in any of the above embodiments.

[0150] Another embodiment of the present application provides a storage medium, which includes a stored program, wherein when the program is running, the device where the storage medium is located is controlled to execute a device failure restart method provided in any of the above embodiments.

[0151] Storage media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0152] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0153] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for restarting a device after a fault, characterized in that: The method comprises: Obtain the total number of machines in the cluster to which the alarm machine belongs and the maximum usage rate of each target resource of each machine; For each target resource, calculate the sum of the highest utilization rates of the target resource of each machine to obtain the total utilization rate corresponding to the target resource; Taking the quotient of the total usage rate corresponding to each target resource and the target value to obtain an integer quotient corresponding to each target resource; wherein the target value is 1; Calculate the sum of the maximum value of the integer quotients corresponding to each target resource and the target value to obtain the required number of machines; The number of groups is obtained by dividing the ratio of the number of machines by the target value and adding the target value to the integer quotient; wherein the ratio of the number of machines is equal to the quotient obtained by dividing the total number of the machines by the difference between the total number of the machines and the required number of machines; Grouping each of the machines according to the number of groups; Restart each of the machines in each group in turn.

2. The method according to claim 1, characterized in that Before obtaining the total number of machines in the cluster to which the alarm machine belongs and the maximum usage rate of each target resource of each machine, the method further includes: Matching the deployment unit corresponding to the alarm machine and obtaining information about the deployment unit corresponding to the alarm machine; wherein the deployment unit is composed of one machine or multiple machines; Determining whether the deployment unit belongs to a cluster based on the cluster tag in the information of the deployment unit corresponding to the alarm machine; If it is determined that the deployment unit belongs to a cluster, the step of obtaining the total number of machines in the cluster to which the alarm machine belongs and the maximum usage rate of each target resource of each machine is executed; If it is determined that the deployment unit does not belong to a cluster, target prompt information is fed back; wherein the target prompt information is used to prompt the user to perform manual intervention; The step of obtaining the total number of machines in the cluster to which the alarm machine belongs and the maximum usage rate of each target resource of each machine includes: Obtain the total number of machines under the deployment unit corresponding to the alarm machine and the maximum usage rate of each target resource of each machine; wherein, each machine in the cluster to which the alarm machine belongs constitutes the deployment unit corresponding to the alarm machine.

3. The method according to claim 2, characterized in that Before determining whether the deployment unit belongs to a cluster based on the cluster tag in the information of the deployment unit corresponding to the alarm machine, the method further includes: Determining whether the deployment unit is related to NetsUnion based on a NetsUnion-related tag in the deployment unit information; wherein NetsUnion is an organizational management platform that interfaces with the clearing platforms of various banks; If it is determined that the deployment unit is not related to the network connection, executing the cluster tag in the information of the deployment unit corresponding to the alarm machine to determine whether the deployment unit belongs to a cluster; If it is determined that the deployment unit is related to the network connection, the target prompt information is fed back.

4. The method according to claim 1, wherein The obtaining of the total number of machines in the cluster to which the alarm machine belongs and the maximum usage rate of each target resource of each machine includes: Get the total number of machines in the cluster to which the alarm machine belongs, as well as the maximum CPU usage and memory usage of each machine; Determine whether there is a special resource indicator identifier in the information of the deployment unit corresponding to the alarm machine; wherein the special resource indicator identifier is an identifier of a resource that the machine providing special services needs to rely on; If it is determined that a special resource indicator identifier exists in the information of the deployment unit corresponding to the alarm machine, the highest usage rate of the resource corresponding to the special resource indicator identifier of each of the machines is obtained.

5. A device for restarting a device after a fault, characterized in that: include: an acquisition unit, configured to acquire the total number of machines in the cluster to which the alarm machine belongs and the maximum usage rate of each target resource of each machine; a summation calculation unit, configured to calculate, for each target resource, the sum of the highest utilization rates of the target resource of each of the machines, to obtain a total utilization rate corresponding to the target resource; a quotient calculation unit, configured to perform a quotient calculation on the total usage rate corresponding to each target resource and the target value to obtain an integer quotient corresponding to each target resource; wherein the target value is 1; a machine quantity determination unit, configured to calculate the sum of the maximum integer quotient corresponding to each target resource and the target value to obtain the required number of machines; a grouping quantity determination unit, configured to obtain the grouping quantity by adding the target value to the integer quotient of the machine quantity ratio divided by the target value, wherein the machine quantity ratio is equal to the quotient obtained by dividing the total quantity of each of the machines by the difference between the total quantity of each of the machines and the required quantity of machines; A grouping unit, configured to group the machines according to the grouping quantity; The restart unit is used to restart each of the machines in each group in turn.

6. The device according to claim 5, characterized in that Also includes: A matching unit, configured to match a deployment unit corresponding to the alarm machine and obtain information about the deployment unit corresponding to the alarm machine; wherein the deployment unit is composed of one machine or multiple machines; A first determining unit is configured to determine, based on a cluster tag in the information of the deployment unit corresponding to the alarm machine, whether the deployment unit belongs to a cluster; wherein, if it is determined that the deployment unit belongs to a cluster, the obtaining unit executes the step of obtaining the total number of machines in the cluster to which the alarm machine belongs and the maximum usage rate of each target resource of each machine; A first feedback unit is configured to feedback target prompt information if it is determined that the deployment unit does not belong to a cluster; wherein the target prompt information is used to prompt a user to perform manual intervention; Wherein, the acquisition unit includes: An acquisition subunit is used to obtain the total number of machines under the deployment unit corresponding to the alarm machine and the maximum utilization rate of each target resource of each machine; wherein, each machine in the cluster to which the alarm machine belongs constitutes the deployment unit corresponding to the alarm machine.

7. The device according to claim 6, characterized in that Also includes: a second judgment unit, configured to judge whether the deployment unit is associated with NetsUnion based on a NetsUnion-related tag in the information of the deployment unit; wherein, if it is determined that the deployment unit is not associated with NetsUnion, the first judgment unit executes the method of judging whether the deployment unit belongs to a cluster based on a cluster tag in the information of the deployment unit corresponding to the alarming machine; NetsUnion is an organizational management platform that interfaces with the clearing platforms of various banks; The second feedback unit is used to feedback the target prompt information if it is determined that the deployment unit is related to the network connection.

8. The device according to claim 5, characterized in that The acquisition unit includes: The first acquisition unit is used to obtain the total number of machines in the cluster to which the alarm machine belongs and the maximum CPU usage and memory usage of each machine; A third judgment unit is configured to judge whether a special resource indicator identifier exists in the information of the deployment unit corresponding to the alarm machine; wherein the special resource indicator identifier is an identifier of a resource that the machine providing the special service needs to rely on; The second obtaining unit is configured to obtain the maximum usage rate of the resource corresponding to the special resource indicator identifier of each of the machines if it is determined that the information of the deployment unit corresponding to the alarm machine contains a special resource indicator identifier.

9. An electronic device, characterized in that: include: A processor and a memory, wherein the memory is used to store program code and data for data update, and the processor is used to call program instructions in the memory to execute a fault restart method for a device as described in any one of claims 1 to 4.

10. A storage medium, characterized in that: The storage medium includes a stored program, wherein when the program is running, the device where the storage medium is located is controlled to execute the device failure restart method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Multi-data center access method and system

    CN112380072A

  • Resource allocation apparatus, resource allocation program and recording media, and resource allocation method

    US20110225300A1