Method and device for recovering instance failure, electronic device and storage medium
Through automated grouping and evaluation of recovery sequence and parallel recovery methods, the problem of Helm instance failure recovery in the AI inference platform relying on manual intervention, achieving efficient and intelligent failure recovery.
Patent Information
- Application Number
- CN202510288610.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-12
AI Technical Summary
In the existing AI inference platform, the Helm instance failure recovery mechanism relies on manual intervention and cannot automatically and intelligently recover, and has a high error rate.
By determining the failure list of instances to be restored in the cluster, grouping the instances to be restored according to the preset instance call relationship, and determining the recovery order of each packet based on the preset evaluation attributes, and finally recovering the instance to be restored in parallel according to the recovery order and the cluster resource state.
The automated and intelligent Helm instance failure recovery is realized, which reduces the risk of errors in manual operations, shortens the overall recovery time of the instance, and improves the recovery efficiency.
Smart Images

Figure CN119806888B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a method and device for recovering from an instance failure, an electronic device, and a storage medium. Background Art
[0002] The Artificial Intelligence (AI) inference platform is a platform used to deploy AI images or models to actual application scenarios for real-time inference. The AI inference platform can create instances of AI images or models in a specified namespace through a package management tool (Helm) chart (a collection organized in a file directory).
[0003] The following failures may occur in the actual operation of the Helm instance in the AI reasoning platform, such as the uninstallation of the underlying Helm command, the accidental deletion of the namespace, the cleanup of the instance, etc. The existing instance recovery mechanism often relies on manual intervention, and cannot automatically and intelligently recover the Helm instance effectively, and the error rate is high. Summary of the invention
[0004] The present application provides a method, device, electronic device and storage medium for recovering from an instance failure.
[0005] According to a first aspect of the present application, a method for recovering from an instance failure is provided, comprising:
[0006] Determine the fault list of instances to be recovered in the cluster. The instances to be recovered are instances created in the specified namespace using a preset template file in the inference platform.
[0007] According to the preset instance call relationship, the instances to be recovered in the fault list are grouped to obtain a number of instance groups to be recovered;
[0008] Evaluate a plurality of instance groups to be restored according to preset evaluation attributes, and determine the restoration order of each instance group to be restored;
[0009] Restore the instances to be restored in parallel according to the restoration order and cluster resource status of each instance group to be restored.
[0010] According to a second aspect of the present application, a device for recovering from an instance failure is provided, comprising:
[0011] A first determining unit is used to determine a fault list of instances to be recovered in the cluster, where the instances to be recovered are instances created in a specified namespace using a preset template file in the inference platform;
[0012] A grouping unit, used to group the instances to be restored in the fault list according to a preset instance call relationship to obtain a plurality of instance groups to be restored;
[0013] A second determining unit, configured to evaluate a plurality of instance groups to be restored according to a preset evaluation attribute, and determine a restoration order of each instance group to be restored;
[0014] The recovery unit is used to recover the instances to be recovered in parallel according to the recovery order of each instance group to be recovered and the cluster resource status.
[0015] According to a third aspect of the present application, an electronic device is provided, including:
[0016] Memory for storing computer programs;
[0017] A processor is used to implement the steps of the method for recovering from a fault in the first aspect when executing a computer program.
[0018] According to a fourth aspect of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein when the computer program is executed by a processor, the steps of the method for recovering from a fault in the example of the first aspect are implemented.
[0019] According to a fifth aspect of the present application, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the steps of the method for recovering from a fault in the example of the first aspect.
[0020] The present application provides a method, device, electronic device and storage medium for recovering instance failures. First, after determining the failure list of instances to be recovered in the cluster, the instances to be recovered in the failure list are grouped according to the preset instance call relationship to obtain several groups of instances to be recovered. Secondly, the several groups of instances to be recovered are evaluated according to the preset evaluation attributes to determine the recovery order of each group of instances to be recovered. Finally, the instances to be recovered are recovered in parallel according to the recovery order of each group of instances to be recovered and the cluster resource status. Since the instances to be recovered in the group of instances to be recovered are recovered in a reasonable recovery order, unnecessary waiting is avoided and the overall recovery time of the instance is shortened. In addition, the recovery efficiency can be further improved by adopting a parallel recovery method. This solves the risk of errors in manual operations and the problem of low recovery efficiency.
[0021] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0023] Figure 1 A framework diagram of an example failure recovery system provided in an embodiment of the present application;
[0024] Figure 2 A flowchart of another example fault recovery method provided in an embodiment of the present application;
[0025] Figure 3 A flowchart of another example fault recovery method provided in an embodiment of the present application;
[0026] Figure 4 A flowchart of another example fault recovery method provided in an embodiment of the present application;
[0027] Figure 5 A flowchart of another example fault recovery method provided in an embodiment of the present application;
[0028] Figure 6 A flowchart of another example fault recovery method provided in an embodiment of the present application;
[0029] Figure 7 A schematic diagram of the structure of an example fault recovery device provided in an embodiment of the present application;
[0030] Figure 8 A schematic diagram of the structure of another example fault recovery device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0031] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0032] It should be noted that, in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0033] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0034] The embodiment of the present application is implemented based on the recovery system of the instance failure, such as Figure 1 As shown, the recovery system includes: an instance data collection module, an instance analysis module, an instance recovery strategy planning module, an instance recovery module, and an instance monitoring and feedback module. The specific implementation functions of each module are described in detail in the following embodiments.
[0035] An embodiment of the present application provides a method for recovering from an instance failure, and the method is described in detail in conjunction with the execution flow of the method for recovering from an instance failure. Figure 2 A flowchart of a method for recovering from an instance failure provided in an embodiment of the present application.
[0036] like Figure 2 As shown, the method comprises the following steps:
[0037] Step 101 : determine a fault list of instances to be recovered in the cluster, where the instances to be recovered are instances created in a specified namespace using a preset template file in an inference platform.
[0038] Step 101 by Figure 1 The example data acquisition module is shown.
[0039] The examples described in the embodiments of the present application include but are not limited to users using Helm template files in the AI reasoning platform to create Helm instances in a specified namespace, or instances in other formats. Specifically, the embodiments of the present application do not specifically limit the instance format and creation form.
[0040] After the Helm instance is created, the instance attribute data of the Helm instance is stored in the database, including the instance identifier (Identity document, id), instance name, namespace, namespace details, instance status (running, failed, offline), resource usage, Helm template file storage location, pull-up time, current average bandwidth, recent usage time, total number of requests, information about calling other instances, number of failures, user information, etc. The instance attribute data generated when the Helm instance is created is stored in the database. When performing instance failure recovery, there is no need to obtain it from the underlying system, which greatly saves retrieval time.
[0041] Since all instances in the system are recorded in the database, based on the difference between the recorded instances in the database and the instances in the cluster, it is possible to obtain which instances are to be restored. That is, when an instance exists in the database but not in the cluster, it means that the missing instance is the instance to be restored.
[0042] The instances to be restored are recorded in the fault list so that the instances to be restored can be repaired in sequence from the fault list.
[0043] Step 102: group the instances to be restored in the fault list according to the preset instance call relationship to obtain a plurality of instance groups to be restored.
[0044] Step 102 consists of Figure 1 The example analysis module execution is shown.
[0045] The purpose of grouping the instances to be restored based on the calling relationship of the instances is to restore the instances in a reasonable calling sequence and avoid unnecessary waiting.
[0046] According to the configuration file (such as values.yaml) in the Helm template file corresponding to the Helm instance, check whether there is any other instance configuration information called in the current instance to be restored. The calling relationship includes: no call, calling itself, and calling other instances. Calling other instances includes: direct call and indirect call.
[0047] After obtaining the call relationship of each instance to be restored, the instances to be restored with the call relationship are grouped into one group. Specifically, the number of instances to be restored and the number of groups are not limited.
[0048] Step 103 : Evaluate a plurality of instance groups to be restored according to preset evaluation attributes, and determine a restoration order for each instance group to be restored.
[0049] Step 103 by Figure 1 The instance recovery strategy planning module shown is executed.
[0050] The purpose of the assessment is to determine which instances among the instances to be recovered need to be restored first to shorten the overall recovery time. For example, instances among the instances to be recovered that are called by multiple instances need to be restored first.
[0051] As an implementation method of the embodiment of the present application, all the instance groups to be restored are scored based on preset evaluation attributes, and are sorted according to the size of the scores to determine the restoration order of each instance group to be restored.
[0052] Step 104 : Restore the instances to be restored in parallel according to the restoration order of each instance group to be restored and the cluster resource status.
[0053] Step 104 consists of Figure 1 The instance shown recovers the module execution.
[0054] The recovery method adopted in the embodiment of the present application is parallel recovery, which reduces the fault recovery time.
[0055] It should be noted that when performing instance recovery, it is necessary to perform recovery according to the cluster resource status, that is, to perform recovery within the operating range of the cluster resource status to improve the efficiency and reliability of fault recovery. When recovering an instance, its recovery resources exceed the cluster resource status, which increases the system load and causes other problems.
[0056] The method for recovering instance failures provided in the present application first determines the fault list of instances to be recovered in the cluster, then groups the instances to be recovered in the fault list according to the preset instance call relationship to obtain several groups of instances to be recovered, and then evaluates the several groups of instances to be recovered according to the preset evaluation attributes to determine the recovery order of each group of instances to be recovered, and finally, recovers the instances to be recovered in parallel according to the recovery order of each group of instances to be recovered and the cluster resource status. Since the instances to be recovered in the group of instances to be recovered are recovered in a reasonable recovery order, unnecessary waiting is avoided and the overall recovery time of the instance is shortened. In addition, the recovery efficiency can be further improved by adopting a parallel recovery method. This solves the error risk in manual operation and the problem of low recovery efficiency.
[0057] As a further limitation of step 101, when determining the fault list of instances to be restored in the cluster, the following methods may be used but are not limited to: Figure 3 As shown, including:
[0058] Step 201, obtain a list of all instances in a non-offline state recorded in a database.
[0059] The instance attribute data records the attribute information of all instances, and a list of all instances that are not offline can be obtained through the instance attribute data.
[0060] Step 202 : Match the instances in the cluster based on the list of all instances, search for instances that exist in the list of all instances but do not exist in the cluster, and determine them as instances to be restored.
[0061] To obtain a list of all Helm instances in the cluster, you can use the Helm command to obtain it. The specific execution logic based on the Helm command can be implemented by referring to any method in the relevant technology, and the embodiments of this application will not be repeated here.
[0062] Each piece of data in the Helm instance list in the database is matched with the data in the Helm instance list in the cluster. If the Helm instance exists in the database but does not exist in the Helm instance list in the cluster, the Helm instance is considered to be faulty. This process continues until all Helm instances in the database instance list are compared and the Helm instance that needs to be restored is obtained, which is the instance to be restored.
[0063] Step 203: record the instance to be restored in the fault list.
[0064] Step 204: Determine whether the namespace where the instance to be restored is located exists.
[0065] Because each instance to be restored must be created in a namespace, the instance cannot be created if the namespace does not exist.
[0066] Step 205: If it does not exist, rebuild the namespace corresponding to the instance to be restored according to the resource information recorded in the database.
[0067] According to the instance to be restored (Helm instance) in the fault list, check whether the namespace where the Helm instance is located exists. If the namespace does not exist, it is necessary to rebuild the scene (namespace) according to the resource information recorded in the database. For the implementation method of rebuilding the namespace, please refer to any method in the relevant technology, so it will not be repeated here.
[0068] In order to reduce the fault recovery time, the instances to be recovered in the fault list are grouped according to the preset instance call relationship to obtain several groups of instances to be recovered. The following method can be used to achieve this: Figure 4 As shown, including:
[0069] Step 301, obtaining a configuration file in the template file, wherein the configuration file includes configuration information between various instances.
[0070] According to the settings of the configuration file (such as values.yaml) in the Helm template file corresponding to the Helm instance, check whether there is configuration information of other instances called in the current instance. According to the settings of the configuration file, you can obtain information about the instance directly calling other instances.
[0071] Step 302: Based on the configuration file, determine the viscosity information of each to-be-recovered instance in the fault list directly calling and / or indirectly calling other instances.
[0072] The calling relationship between instances can be obtained according to the configuration file.
[0073] Step 303: group the instances to be restored in the fault list according to the viscosity information to obtain a plurality of groups of instances to be restored.
[0074] Assume that there are G instances to be restored, which are divided into M groups according to the calling relationship (viscosity) between instances. In practical applications, slices are used to store data, such as using var arrM [ ][ ] stores data, where len(arrM)=M, one-dimensional storage is 1 to M, and two-dimensional storage has multiple instances to be restored with call relationships. Assume that each group has Instances to be restored, The minimum value is 1 and the maximum value is G.
[0075] Each group uses a two-dimensional array var arrName[ ][ ] indicates the viscosity between Helm instances in a group. The instance was For viscosity Instance, Examples and The viscosity of an instance can be expressed as:
[0076]
[0077]
[0078] Indicates The viscosity of the instance to be restored. Since the instance to be restored does not have viscosity, therefore:
[0079]
[0080] Indicates The examples and The viscosity between the instances to be restored, that is, The instance to be restored calls Examples:
[0081]
[0082] in, Indicates The instance to be restored indirectly calls The number of paths that exist in the instance to be restored. len(path) indicates the length of the path.
[0083] The highest viscosity value for direct calls between instances to be restored is 1, the viscosity value for indirect calls between instances to be restored is moderate, and the lowest viscosity value for no calls between instances to be restored is 0. Direct calls and indirect calls may coexist between instances to be restored.
[0084]
[0085] Step 304 , sorting the instances to be restored in each group of instances to be restored according to the viscosity value to obtain sorted groups of instances to be restored.
[0086] According to the above calculation Instances value, The larger the value, the stronger the viscosity is. This instance is called by more instances and should be restored first. The values are sorted from large to small to obtain the order of the multiple instances to be restored in this group, and var arrM [ ][ ] in the order of multiple Helm instances stored in two dimensions.
[0087] To improve the efficiency and reliability of fault recovery, you can use the following method to recover instances in parallel according to the recovery order and cluster resource status of each instance group: Figure 5 The methods shown include:
[0088] Step 401, determining the number of parallel operations according to the usage status of processors and memory in the cluster, wherein cluster resources include processors and memory in the cluster.
[0089] When determining the number of parallel operations, you can do it in the following ways:
[0090] Step 4011, determining the first parallel number supported by the system according to the total number of processors in the cluster and the number already in use.
[0091] Determine the first parallel number When , it is calculated by the following formula:
[0092]
[0093] in, is the total number of processors in the cluster, is the number of processors already used in the cluster, It is a preset threshold, for example, 0.94, 0.92 or other values, which are not limited in the specific embodiments of the present application.
[0094] Step 4012: Determine the second parallel number supported by the system based on the total amount of memory in the cluster and the amount already used.
[0095] In determining the second parallel number When , it is calculated by the following formula:
[0096]
[0097] in, is the total amount of memory in the cluster, is the amount of memory used in the cluster. It is a preset threshold, for example, 0.86, 0.87 or other values, which are not limited in the specific embodiments of the present application.
[0098] Step 4013: determine the minimum value between the first parallel number and the second parallel number as the final parallel number.
[0099] The final number of parallel operations supported by the system: .
[0100] Step 402: Determine the grouping of instances to be restored that are executed in parallel according to the recovery order and the number of parallel operations.
[0101] In order to ensure efficient resource utilization during the recovery process and not cause resource overload, when determining the grouping of instances to be recovered for parallel execution, if the final number of parallel instances is less than the number of groupings of the instances to be recovered, the instances to be recovered with the parallel number are grouped according to the recovery order and determined as the grouping of instances to be recovered for parallel execution; if the final number of parallel instances is greater than or equal to the number of groupings of the instances to be recovered, all the instances to be recovered are grouped and determined as the grouping of instances to be recovered for parallel execution.
[0102] For example, if the parallel number number is less than the group number M, then it is determined that the number of groups of instances to be restored that are executed in parallel is number.
[0103] If the parallel number number is greater than or equal to the number of groups M, then the number of groups of instances to be restored that are executed in parallel is determined to be M.
[0104] Step 403 : From each group of instances to be restored executed in parallel, restore the instances to be restored according to the sorting result within the group.
[0105] If the number of parallel operations number is less than the number of groups M, then restore number of to-be-restored instance groups, starting with the first to-be-restored instance in each group, and move the restored number of to-be-restored instances from var arrM [ ][ ] is deleted.
[0106] If the number of parallel operations is greater than or equal to the number of groups M, restore the M groups of instances to be restored, starting from the first group of instances to be restored in each group, and restore the M Helm instances to be restored from var arrM [ ][ ] is deleted.
[0107] Step 404 , after the first instance to be restored is restored, the number of parallel operations is re-determined according to the usage status of the processors and the usage status of the memory in the cluster, and the next instance to be restored is restored.
[0108] Based on the final determined parallel number, after the instance to be restored is restored in parallel once, wait for the instance to be restored to return to normal, recalculate the parallel number of the instance to be restored, and then execute in parallel again until all the instances to be restored are restored.
[0109] Further, as a detailed description of step 103, when evaluating a plurality of to-be-restored instance groups according to preset evaluation attributes, a recovery order of each to-be-restored instance group is determined, such as Figure 6 As shown, including:
[0110] Step 501, score each instance group to be recovered according to the most recent usage time, requested resource size, total number of failures and expected pull-up time; wherein the preset evaluation attributes include the most recent usage time, requested resource size, total number of failures and expected pull-up time.
[0111] Step 5011, obtain instance attribute data generated when creating the instance to be restored, the instance attribute data at least includes: the most recent use time of each group of instances, the upper limit threshold of memory request resources, the upper limit threshold of processor request resources, the pull-up time, and the total number of failures of each group of instances.
[0112] For the process of obtaining instance attribute data, please refer to the detailed description of the above embodiment, which will not be described again here.
[0113] Step 5012, determining the requested resource size according to the memory requested resource upper limit threshold and the processor requested resource upper limit threshold.
[0114] Determine the upper limit threshold of the total memory request resource for each group It can be calculated by the following formula:
[0115]
[0116] in, The upper limit of memory resources requested for each instance to be recovered in the group.
[0117] When determining the upper threshold of the total processor resource request for each group It can be calculated by the following formula:
[0118]
[0119] in, The upper limit of processor resources requested for each instance to be recovered in the group.
[0120] Based on the memory request resource upper limit threshold and processor request resource upper limit threshold , determine the requested resource size for each group of instances to be restored.
[0121] Step 5013, determine the estimated pull-up time according to the pull-up time, the current average bandwidth, and the current bandwidth.
[0122] The estimated pull-up time for each group of instances to be restored can be calculated using the following formula:
[0123]
[0124] in, The pull-up time is recorded in the instance attribute data. is the current average bandwidth, recorded in the instance attribute data, is the current bandwidth, obtained through measurement.
[0125] Step 5014, query the instance attribute data to determine the most recent usage time of each group of instances and the total number of failures of each group of instances.
[0126] The most recent usage time for each group of instances:
[0127] Total number of failures per group of instances:
[0128] The most recent usage time is the last time the instance was accessed. Indicates the most recent usage time of each instance to be restored in the instance group to be restored.
[0129] Step 5015: Call a preset evaluation algorithm to calculate the score of each group of instances to be restored.
[0130] Each group of molecules to be restored is scored according to the following preset evaluation algorithm:
[0131]
[0132]
[0133] in, It is a constant value and tends to be a very small number, which avoids the influence of extreme values and makes the score more stable. is the value with the largest recent usage time of group M. is the minimum value of the number of processor resources requested in group M, is the maximum value of the number of processor resources requested in group M. The minimum value of the requested memory resources in group M. The maximum value of the requested memory resources in group M. is the maximum total number of instance failures in group M. is the minimum total number of instance failures in group M, is the sum of the estimated pull-up times of the M groups, The estimated pull-up time for the current group.
[0134] In an embodiment of the present application, the longer the most recent usage time of each group of instances, the higher the score; the smaller the memory request resource upper limit threshold of each group of instances, the higher the score; the smaller the processor request resource upper limit threshold of each group of instances, the higher the score; the shorter the pull-up time, the higher the score; the fewer the total number of failures of each group of instances, the higher the score.
[0135] Step 502: Determine the recovery order of each to-be-recovered instance group according to the score.
[0136] Presented by Figure 1 The system framework further includes an instance monitoring and feedback module, which is used to perform the following:
[0137] determining a first operating state of the restored instance;
[0138] comparing the first operating state with a second operating state recorded in a database;
[0139] If the first running state is consistent with the second running state, it is determined that the instance to be restored is restored successfully;
[0140] If the first running state is inconsistent with the second running state, it is determined that the instance to be restored has failed to be restored, and information of the failed instance to be restored is output.
[0141] After the instance to be restored is restored, the restored instance is obtained. It is necessary to monitor whether the restored instance has returned to normal and query whether the final state of the restored instance is consistent with the instance state stored in the database. If not, it means that the restoration of the instance to be restored has failed. The information of the failed instance to be restored is recorded and given, waiting for manual intervention and confirmation. The information of the failed instance to be restored includes but is not limited to the ID of the failed instance to be restored.
[0142] In some embodiments, a list of restored instances is obtained based on the instance recovery record, the list of restored instances is traversed, and the status of all minimum scheduling units (pods) and containers under the instance is obtained. If all are running, the restored instance is running. Otherwise, the instance status is failed and compared with the instance status queried in the database. If the status is consistent, the recovery is successful. If the status is inconsistent, the recovery fails. The recorded information is fed back to the operation and maintenance personnel for confirmation and manual recovery.
[0143] The embodiments of the present application can achieve the following beneficial effects:
[0144] 1. When an instance failure occurs, it can restore the instances to be restored in parallel according to the viscosity between the instances to be restored, the scores of the instances to be restored, and the cluster resource status, thereby reducing the failure recovery time and improving the efficiency and reliability of failure recovery.
[0145] 2. Restore the instances to be restored starting from the ones with the highest scores in a reasonable order (the order of restoration of each instance group to be restored), avoiding unnecessary waiting and shortening the overall recovery time. Avoid the failure of other instances to start due to unready instances. Support parallel recovery to further improve efficiency.
[0146] 3. Reduce the risk of errors in manual operations by standardizing the recovery sequence. This can improve the success rate of recovery and reduce the impact on user experience.
[0147] 4. After the instance to be restored fails to be restored, a clear recovery failure log is generated to quickly locate the problem.
[0148] Corresponding to the above-mentioned method for recovering from an example fault, the present application also proposes a device for recovering from an example fault. Since the device embodiment of the present application corresponds to the above-mentioned method embodiment, details not disclosed in the device embodiment can be referred to the above-mentioned method embodiment, and will not be repeated in the present application.
[0149] Figure 7 A schematic diagram of a structure of a device for recovering from an example fault provided in an embodiment of the present application is shown in FIG. Figure 7 As shown, including:
[0150] A first determining unit 61 is used to determine a fault list of instances to be recovered in the cluster, where the instances to be recovered are instances created in a specified namespace using a preset template file in the inference platform;
[0151] A grouping unit 62, configured to group the instances to be restored in the fault list according to a preset instance call relationship to obtain a plurality of instance groups to be restored;
[0152] A second determining unit 63 is used to evaluate a plurality of instance groups to be restored according to a preset evaluation attribute, and determine a restoration order of each instance group to be restored;
[0153] The recovery unit 64 is used to recover the instances to be recovered in parallel according to the recovery order of each group of instances to be recovered and the cluster resource status.
[0154] The recovery device for instance failure provided by the present application first determines the fault list of instances to be recovered in the cluster, then groups the instances to be recovered in the fault list according to the preset instance call relationship to obtain several groups of instances to be recovered, and then evaluates the several groups of instances to be recovered according to the preset evaluation attributes to determine the recovery order of each group of instances to be recovered, and finally, recovers the instances to be recovered in parallel according to the recovery order of each group of instances to be recovered and the cluster resource status. Since the instances to be recovered in the group of instances to be recovered are recovered in a reasonable recovery order, unnecessary waiting is avoided and the overall recovery time of the instances is shortened. In addition, the recovery efficiency can be further improved by adopting a parallel recovery method. This solves the error risk in manual operation and the problem of low recovery efficiency.
[0155] Furthermore, in a possible implementation of the embodiment of the present application, as Figure 8 As shown, the first determining unit 61 is further used for:
[0156] Get a list of all instances in the database that are not offline;
[0157] Match the instances in the cluster based on the list of all instances, find the instances that exist in the list of all instances but not in the cluster, and identify them as instances to be restored;
[0158] The instance to be restored is recorded in the fault list.
[0159] Furthermore, in a possible implementation of the embodiment of the present application, as Figure 8 As shown, the grouping unit 62 includes:
[0160] A first acquisition module 621 is used to acquire a configuration file in the template file, where the configuration file includes configuration information between various instances;
[0161] A first determination module 622 is used to determine, based on the configuration file, viscosity information of each to-be-recovered instance in the fault list directly calling and / or indirectly calling other instances;
[0162] The grouping module 623 is used to group the instances to be restored in the fault list according to the viscosity information to obtain a plurality of groups of instances to be restored.
[0163] Furthermore, in a possible implementation of the embodiment of the present application, as Figure 8 As shown, the grouping unit 62 also includes:
[0164] The sorting module 624 is used to sort the instances to be restored in each group of instances to be restored according to the viscosity value to obtain sorted groups of instances to be restored.
[0165] Furthermore, in a possible implementation of the embodiment of the present application, as Figure 8 As shown, the recovery unit 64:
[0166] A second determination module 641 is used to determine the number of parallel operations according to the usage status of the processors and the memory in the cluster, wherein the cluster resources include the processors and the memory in the cluster;
[0167] The third determination module 642 is used to determine the group of instances to be restored that are executed in parallel according to the recovery order and the number of parallel instances;
[0168] A recovery module 643, configured to recover the instances to be recovered from each group of instances to be recovered executed in parallel according to the sorting result within the group;
[0169] The processing module 644 is used to, after the first instance to be restored is restored, re-determine the number of parallel operations according to the usage status of the processors and the usage status of the memory in the group, and continue to restore the next instance to be restored.
[0170] Furthermore, in a possible implementation of the embodiment of the present application, as Figure 8 As shown, the second determination module 641 is also used for:
[0171] Determine the first number of parallel operations supported by the system based on the total number of processors in the cluster and the number of processors already in use;
[0172] Determine the number of second parallel operations supported by the system based on the total amount of memory in the cluster and the amount already in use;
[0173] The minimum value between the first parallel number and the second parallel number is determined as the final parallel number.
[0174] Furthermore, in a possible implementation of the embodiment of the present application, as Figure 8As shown, the third determination module 642 is further used for:
[0175] If the final number of parallel operations is less than the number of instances to be restored, the instances to be restored with the same number of parallel operations are grouped according to the recovery order to determine the instances to be restored to be executed in parallel.
[0176] If the final number of parallel operations is greater than or equal to the number of the to-be-recovered instance groups, all the to-be-recovered instance groups are grouped as the to-be-recovered instance groups to be executed in parallel.
[0177] Furthermore, in a possible implementation of the embodiment of the present application, as Figure 8 As shown, the second determination unit 63 is further used to score each to-be-recovered instance group according to the most recent use time, the requested resource size, the total number of failures, and the expected pull-up time; wherein the preset evaluation attributes include the most recent use time, the requested resource size, the total number of failures, and the expected pull-up time;
[0178] The recovery order of each instance group to be recovered is determined according to the size of the score.
[0179] Furthermore, in a possible implementation of the embodiment of the present application, as Figure 8 As shown, the second determining unit 63 includes:
[0180] The second acquisition module 631 is used to acquire instance attribute data generated when creating the instance to be restored, and the instance attribute data at least includes: the most recent use time of each group of instances, the upper threshold of memory request resources, the upper threshold of processor request resources, the pull-up time, and the total number of failures of each group of instances;
[0181] A fourth determination module 632, configured to determine a requested resource size according to a memory requested resource upper threshold and a processor requested resource upper threshold;
[0182] A fifth determining module 633, configured to determine an estimated pull-up time according to the pull-up time, the current average bandwidth, and the current bandwidth;
[0183] A query module 634 is used to query instance attribute data to determine the most recent usage time of each group of instances and the total number of failures of each group of instances;
[0184] The evaluation module 635 is used to call a preset evaluation algorithm to calculate the score of each group of instances to be restored.
[0185] Furthermore, in a possible implementation of the embodiment of the present application, as Figure 8 As shown, the device also includes:
[0186] A third determining unit 65 is used to determine a first running state of the restored instance after the restoration unit restores the instance to be restored in parallel according to the restoration order of each group of the instance to be restored and the cluster resource state;
[0187] a comparing unit 66, for comparing the first operating state with the second operating state recorded in a database;
[0188] A fourth determining unit 67, configured to determine that the instance to be restored is restored successfully if the first running state is consistent with the second running state;
[0189] The fifth determining unit 68 is configured to determine that the instance to be restored has failed to be restored if the first operating state is inconsistent with the second operating state, and output information of the instance to be restored that has failed to be restored.
[0190] Furthermore, in a possible implementation of the embodiment of the present application, as Figure 8 As shown, the device also includes:
[0191] A sixth determining unit 69 is configured to determine whether a namespace where the instance to be restored is located exists after the first determining unit determines the fault list of the instance to be restored in the cluster;
[0192] The reconstruction unit 610 is used to rebuild the namespace corresponding to the instance to be restored according to the resource information recorded in the database if it does not exist.
[0193] It should be noted that the above explanation of the method embodiment is also applicable to the device of the embodiment of the present application, the principle is the same, and it is no longer limited in the embodiment of the present application.
[0194] For the description of the features in the embodiment corresponding to the state recovery device, please refer to the relevant description of the embodiment corresponding to the example fault recovery method, which will not be repeated here.
[0195] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned example fault recovery method embodiments.
[0196] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned examples of the method for recovering from a fault when running.
[0197] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0198] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned examples of the method for recovering from a fault are implemented.
[0199] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned example fault recovery method embodiments are implemented.
[0200] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0201] The above is a detailed introduction to an example fault recovery method and device, electronic device and storage medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A method for recovering from an instance failure, characterized in that: include: Determine a fault list of instances to be recovered in the cluster, where the instances to be recovered are instances created in a specified namespace using a preset template file in the inference platform; According to a preset instance call relationship, the instances to be restored in the fault list are grouped to obtain a plurality of instance groups to be restored; Evaluate the plurality of instance groups to be restored according to preset evaluation attributes, and determine a restoration order for each instance group to be restored; Restore the instances to be restored in parallel according to the restoration order of each group of instances to be restored and the cluster resource status; The step of grouping the instances to be restored in the fault list according to the preset instance call relationship to obtain a plurality of instance groups to be restored includes: Based on the configuration file, determine the viscosity information of each to-be-recovered instance in the fault list directly calling and / or indirectly calling other instances The configuration file includes configuration information between each instance, and the viscosity information For example With examples The calling relationship between them; The instances to be restored in each group of instances to be restored are sorted according to the viscosity value to obtain sorted groups of instances to be restored; Wherein, the viscosity value is calculated by the following formula: in, Indicates the viscosity value, Indicates the number of the instances to be restored in each group of instances to be restored.
2. The method for recovering from instance failure according to claim 1, characterized in that: Determining the fault list of instances to be restored in the cluster includes: Get a list of all instances in the database that are not offline; Matching the instances in the cluster based on the list of all instances, searching for instances that exist in the list of all instances but do not exist in the cluster, and determining the instances as the instances to be restored; The instance to be restored is recorded in the fault list.
3. The method for recovering from instance failure according to claim 1, characterized in that: The step of grouping the instances to be restored in the fault list according to the preset instance call relationship to obtain a plurality of instance groups to be restored includes: Get the configuration file in the template file; Based on the configuration file, determine viscosity information of each to-be-recovered instance in the fault list directly calling and / or indirectly calling other instances; The instances to be restored in the fault list are grouped according to the viscosity information to obtain the plurality of groups of instances to be restored.
4. The method for recovering from instance failure according to claim 1, characterized in that: The recovering the instances to be recovered in parallel according to the recovery order of each group of instances to be recovered and the cluster resource status comprises: Determine the number of parallel operations according to the usage status of processors and memory in the cluster, wherein the cluster resources include processors and memory in the cluster; Determine the grouping of instances to be restored that are executed in parallel according to the recovery order and the number of parallel operations; From each group of instances to be restored executed in parallel, restore the instances to be restored according to the sorting result within the group; After the first instance to be restored is restored, the number of parallel operations is re-determined according to the usage status of the processors in the group and the usage status of the memory, and the next instance to be restored is restored.
5. The method for recovering from instance failure according to claim 4, characterized in that: The number of parallel operations is determined according to the usage status of the processors and the usage status of the memory in the cluster; Determining a first parallel number supported by the system according to the total number of processors in the cluster and the number of processors already in use; Determine the second parallel number supported by the system according to the total amount of the memory in the cluster and the amount that has been used; The minimum value between the first parallel number and the second parallel number is determined as the final parallel number.
6. The method for recovering from instance failure according to claim 5, characterized in that: Determining the grouping of instances to be restored to be executed in parallel according to the recovery order and the number of parallel operations includes: If the final number of parallel operations is less than the number of groups of the to-be-restored instance groups, the to-be-restored instances of the parallel number are grouped according to the recovery order to be determined as the to-be-restored instance groups executed in parallel; If the final number of parallel operations is greater than or equal to the number of the instance groups to be restored, all the instance groups to be restored are determined as the instance groups to be restored for parallel execution.
7. The method for recovering from instance failure according to claim 1, characterized in that: The step of evaluating the plurality of instance groups to be restored according to the preset evaluation attributes to determine the restoration order of each instance group to be restored includes: Scoring each instance group to be recovered according to the most recent usage time, the requested resource size, the total number of failures, and the expected pull-up time; wherein the preset evaluation attributes include the most recent usage time, the requested resource size, the total number of failures, and the expected pull-up time; The recovery order of each to-be-recovered instance group is determined according to the size of the score.
8. The method for recovering from instance failure according to claim 7, characterized in that: Scoring each instance group to be recovered according to the most recent usage time, requested resource size, total number of failures, and estimated pull-up time includes: Obtain instance attribute data generated when creating the instance to be restored, wherein the instance attribute data includes at least: the most recent usage time of each group of instances, the upper threshold of memory request resources, the upper threshold of processor request resources, the startup time, and the total number of failures of each group of instances; Determining the requested resource size according to the memory request resource upper limit threshold and the processor request resource upper limit threshold; Determine the estimated pull-up time according to the pull-up time, the current average bandwidth, and the current bandwidth; Querying the instance attribute data to determine the most recent usage time of each group of instances and the total number of failures of each group of instances; The preset evaluation algorithm is called to calculate the score of each instance group to be restored.
9. The method for recovering from instance failure according to claim 1, characterized in that: After the instances to be restored are restored in parallel according to the restoration order and cluster resource status of each group of instances to be restored, the method further includes: determining a first operating state of the restored instance; comparing the first operating state with a second operating state recorded in a database; If the first running state is consistent with the second running state, it is determined that the instance to be restored is restored successfully; If the first running state is inconsistent with the second running state, it is determined that the instance to be restored has failed to be restored, and information of the failed instance to be restored is output.
10. The method for recovering from instance failure according to claim 1, characterized in that: After determining the failure list of instances to be recovered in the cluster, the method further includes: Determine whether the namespace where the instance to be restored exists exists; If it does not exist, the namespace corresponding to the instance to be restored is rebuilt according to the resource information recorded in the database.
11. A device for recovering from instance failure, characterized in that: include: A first determining unit is used to determine a fault list of instances to be restored in the cluster, where the instances to be restored are instances created in a specified namespace using a preset template file in an inference platform; A grouping unit, used for grouping the instances to be restored in the fault list according to a preset instance call relationship to obtain a plurality of instance groups to be restored; A second determining unit, configured to evaluate the plurality of instance groups to be restored according to a preset evaluation attribute, and determine a restoration order of each instance group to be restored; A recovery unit, configured to recover the instances to be recovered in parallel according to the recovery order of each group of instances to be recovered and the cluster resource status; The grouping units include: The first determination module is used to determine, based on the configuration file, the viscosity information of each to-be-recovered instance in the fault list directly calling and / or indirectly calling other instances. The configuration file includes configuration information between each instance, and the viscosity information For example With examples The calling relationship between them; A sorting module, used to sort the instances to be restored in each group of instances to be restored according to the viscosity value, to obtain sorted groups of instances to be restored; Wherein, the viscosity value is calculated by the following formula: in, Indicates the viscosity value, Indicates the number of the instances to be restored in each group of instances to be restored.
12. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the method for recovering from an instance failure as claimed in any one of claims 1 to 10 when executing the computer program.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method for recovering from an instance failure as claimed in any one of claims 1 to 10.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method for recovering from a failure of an example as claimed in any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
Fault recovery method and device for micro-service, electronic equipment and medium
CN114138522A
Service fault recovery method and system for micro-service platform, and storage medium
CN119597521A