A fault tolerance processing method, device, medium and computer program product
By monitoring the status of resource nodes in a distributed system, determining task priorities and using the fault-tolerant processing action generation model for action evaluation, the problem of insufficient continuous operation capability of distributed system services in the prior art is solved, and higher fault-tolerant scheduling capabilities and system reliability are achieved.
Patent Information
- Application Number
- CN202411946519.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-12-27
AI Technical Summary
The prior art has failed to effectively ensure the continuous operation of services in distributed systems, especially when resource nodes fail. The fault tolerance strategy only considers the equipment failure handling at the current moment and fails to ensure the long-term reliability of services.
By monitoring the status of resource nodes of the distributed system, determining the task priority on the fault node, and using the trained fault-tolerant processing action generation model to generate multiple fault-tolerant processing actions, performing availability evaluation, and selecting actual fault-tolerant processing actions to ensure that tasks can be migrated or waited smoothly when a fault occurs, avoiding service interruption.
It improves the fault-tolerant scheduling capability of distributed systems when resource node failures, ensures the continuous operation of services, reduces service interruption time caused by node failures, and enhances the reliability of the system.
Smart Images

Figure CN119376994B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distributed computing, and particularly to a fault tolerance processing method, device, medium, and computer program product. Background Art
[0002] A distributed system refers to a system with multiple resource nodes, which may include but are not limited to computing resource nodes or storage resource nodes. To improve the reliability and availability of a distributed system, it is necessary for the system to have strong fault tolerance capabilities, which mainly include two aspects: fault avoidance and fault tolerance. Among them, fault tolerance includes two aspects: active fault tolerance and passive fault tolerance. Active fault tolerance means preprocessing faults through prediction. Passive fault tolerance means promptly detecting the occurred faults and adopting fault handling strategies for fault response.
[0003] Around the above-mentioned fault tolerance processing idea, there are many fault tolerance strategies for distributed systems in related technologies. However, these fault tolerance strategies only consider handling device faults at the current moment and do not consider the guarantee ability for the continuous operation of the services provided by the distributed system.
[0004] Improving the fault tolerance scheduling of a distributed system for the guarantee ability of the continuous operation of the services provided by the distributed system is a technical problem that needs to be solved by those skilled in the art. Summary of the Invention
[0005] The purpose of the present invention is to provide a fault tolerance processing method, device, medium, and computer program product for improving the guarantee ability of the fault tolerance scheduling of a distributed system for the continuous operation of the services provided by the distributed system.
[0006] To solve the above technical problem, the present invention provides a fault tolerance processing method, including:
[0007] Monitoring the status information of each resource node of the target distributed system;
[0008] After detecting a faulty resource node, determining the service tasks on the faulty resource node as tasks to be processed;
[0009] Determining the priorities of the tasks to be processed, and performing fault tolerance processing on the tasks to be processed in descending order of priorities;
[0010] When performing fault tolerance processing on the tasks to be processed, using a trained fault tolerance processing action generation model to generate multiple fault tolerance processing actions for the tasks to be processed according to the status information at the current moment, performing availability evaluation on each of the fault tolerance processing actions, determining the actual fault tolerance processing action according to the availability evaluation result, and performing the actual fault tolerance processing action on the tasks to be processed.
[0011] On the one hand, determining the priority of the task to be processed includes:
[0012] Determining the fault impact degree parameter of the faulty resource node according to the fault type of the faulty resource node;
[0013] Determining the task evaluation parameter of the task to be processed according to the load status of the task to be processed on the faulty resource node;
[0014] Determining the priority score of the task to be processed according to the fault impact degree parameter and the task evaluation parameter.
[0015] On the other hand, determining the fault impact degree parameter of the faulty resource node according to the fault type of the faulty resource node includes:
[0016] Determining the predicted value of the fault duration of the faulty resource node according to the historical operation data of the faulty resource node;
[0017] Taking the predicted value of the fault duration as the fault impact degree parameter.
[0018] On the other hand, determining the predicted value of the fault duration of the faulty resource node according to the historical data corresponding to the fault type of the faulty resource node includes:
[0019] Determining the expected value of the fault duration of the faulty resource node according to the fault type of the faulty resource node and the corresponding fault duration probability distribution model;
[0020] Taking the expected value of the fault duration as the predicted value of the fault duration;
[0021] Wherein, the fault duration probability distribution model corresponding to the fault type is determined according to the historical operation data of the resource node.
[0022] On the other hand, determining the task evaluation parameter of the task to be processed according to the load status of the task to be processed on the faulty resource node includes:
[0023] Taking the resource occupation parameter of the task to be processed on the faulty resource node and the remaining execution time of the task to be processed on the faulty resource node as the load status of the task to be processed on the faulty resource node, and determining the task evaluation parameter according to the resource occupation parameter and the remaining execution time.
[0024] On the other hand, the determining step of the resource occupation parameter includes:
[0025] Determining the resource occupation parameter according to the resource amounts of multiple types of resources occupied by the task to be processed on the faulty resource node;
[0026] Determine the task evaluation parameter according to the resource occupancy parameter and the remaining execution time, including:
[0027] Use the ratio of the resource occupancy parameter to the remaining execution time as the task evaluation parameter.
[0028] On the other hand, determine the task evaluation parameter of the task to be processed according to the load status of the task to be processed on the faulty resource node, including:
[0029] Determine the task evaluation parameter according to the load status of the task to be processed on the faulty resource node and the task type priority parameter of the task to be processed.
[0030] On the other hand, the fault impact degree parameter is the predicted value of the fault duration of the faulty resource node;
[0031] The task evaluation parameter is the resource consumption value per unit time of the task to be processed on the faulty resource node;
[0032] Determine the priority score of the task to be processed according to the fault impact degree parameter and the task evaluation parameter, including:
[0033] Use the ratio of the resource consumption value per unit time to the predicted value of the fault duration as the priority score.
[0034] On the other hand, the training steps of the fault tolerance processing action generation model include:
[0035] Input the sample status information into the fault tolerance processing action generation model, and output the sample fault tolerance processing action for the sample status information;
[0036] After executing the sample fault tolerance processing action, obtain the updated sample status information;
[0037] Determine the environmental reward value according to the updated sample status information and the reward function;
[0038] Use a set of the sample fault tolerance processing actions, the sample fault tolerance processing action, the updated sample status information, and the environmental reward value as a training sample;
[0039] Use the training sample to update the model parameters of the fault tolerance processing action generation model.
[0040] On the other hand, the reward function is a function of the fault tolerance processing time of the task.
[0041] On the other hand, the fault tolerance processing time includes a first processing time for the task migration action, a second processing time for the task waiting action, and a third processing time for the task scaling action;
[0042] Wherein, the task migration action is to migrate the to-be-processed task from the faulty resource node to the first resource node, and the first processing time includes the data transmission time and the service recovery time at the first resource node; the task waiting action is to set the to-be-processed task to a service interruption state at the faulty resource node, and the second processing time is the service interruption waiting time; the task scaling action is to start the to-be-processed task at the second resource node where there is a copy of the to-be-processed task and terminate the to-be-processed task at the faulty resource node, and the third processing time is the service recovery time at the second resource node.
[0043] On the other hand, perform an availability evaluation on each of the fault tolerance processing actions, and determine the actual fault tolerance processing action according to the availability evaluation result, including:
[0044] Perform an availability evaluation on the fault tolerance processing actions in descending order of the corresponding environmental reward values until the fault tolerance processing action that meets the availability evaluation conditions is obtained as the actual fault tolerance processing action.
[0045] To solve the above technical problems, the present invention also provides a fault tolerance processing device, including:
[0046] A memory for storing a computer program;
[0047] A processor for executing the computer program, and when the computer program is executed by the processor, the steps of the fault tolerance processing method as described in any one of the above are implemented.
[0048] To solve the above technical problems, the present invention also provides a non-volatile storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the fault tolerance processing method as described in any one of the above are implemented.
[0049] To solve the above technical problems, the present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the fault tolerance processing method as described in any one of the above are implemented.
[0050] The fault tolerance processing method provided by the present invention has the beneficial effect that by monitoring the status information of each resource node of the target distributed system, after detecting a faulty resource node, determining the service tasks on the faulty resource node as tasks to be processed, and determining the priorities of the tasks to be processed, so as to perform fault tolerance processing on the tasks to be processed in the order from high to low priority; when performing fault tolerance processing on the tasks to be processed, using the trained fault tolerance processing action generation model to generate multiple fault tolerance processing actions for the tasks to be processed according to the status information at the current moment, evaluating the availability of each fault tolerance processing action, determining the actual fault tolerance processing action according to the availability evaluation result, and performing the actual fault tolerance processing action on the tasks to be processed, so that when a resource node fails in the distributed system, the tasks on the resource node can be feasibly scheduled to ensure the continuous execution of these tasks, and there will be no long-term service interruption due to the failure of the node where they are located or the inability to continue providing services after migrating to an inappropriate node, thereby ensuring the continuous operation of the services provided by the distributed system and effectively improving the reliability of the distributed system.
[0051] The fault tolerance processing device, non-volatile storage medium and computer program product provided by the present invention have the above beneficial effects and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings.
[0053] Figure 1 It is a flowchart of a fault tolerance processing method provided by an embodiment of the present invention;
[0054] Figure 2 It is a schematic diagram of a distributed system fault tolerance scheduling architecture provided by an embodiment of the present invention;
[0055] Figure 3 It is a schematic diagram of the structure of a fault tolerance processing device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] The core of the present invention is to provide a fault tolerance processing method, device, medium and computer program product for improving the ability of the fault tolerance scheduling of a distributed system to ensure the continuous operation of the services provided by the distributed system.
[0057] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0058] To facilitate the understanding of the technical solutions provided by the embodiments of the present invention, some key terms used in the embodiments of the present invention will be explained here first.
[0059] A distributed system is a system formed by interconnecting multiple computers through communication lines. Therefore, each node in the distributed system contains its own processor and memory and has the function of independently processing data. A large task can be divided into several subtasks, which can be executed on different nodes in the distributed system respectively. Generally, for users, a distributed system is a unified whole, and its intelligent agent is used for task scheduling when executing user tasks.
[0060] According to different design goals, a distributed system can be designed as a distributed storage system or a distributed computing system, etc. The distributed storage system is mainly used to process storage service tasks, aiming to store and manage a large amount of data, and provide data persistence, reliability, and accessibility. The distributed computing system is mainly used to provide computing service tasks, aiming to process large-scale computing tasks in a distributed manner to improve computing efficiency and processing power.
[0061] In the fault tolerance processing solution provided by the embodiments of the present invention, since the nodes in the distributed system provide the resources required to execute service tasks, these nodes are called resource nodes. A resource node is usually a computer or a server and is configured as a computing resource node, a storage resource node, or other resource nodes with a particular emphasis according to service requirements.
[0062] An edge network is a distributed computing environment that deploys data storage and analysis capabilities near data generation endpoints or devices to achieve local data processing.
[0063] An edge data center is a newly emerging technology in recent years and has received extensive attention from industrial users including operators. It can provide resource support such as computing power, memory, and storage for mobile terminal users, and is interconnected through the communication network established by the operator. Within the same geographical area, edge data centers can cooperate with each other to achieve load balancing, which greatly expands the business scope of edge service providers and can quickly respond to users. For example, mobile communication operators place some communication services or network elements at the edge to reduce user response latency.
[0064] In a distributed system, especially an edge computing system (compared with a cloud data center, the environment where an edge computing system is located is usually more unreliable, including power outages and construction in the physical environment that can cause downtime, and management methods that prevent it from being inspected or repaired in a timely manner), to improve the reliability and availability of the system, it is necessary for it to have strong fault tolerance capabilities.
[0065] In related technologies, the fault tolerance technologies of distributed systems are mainly divided into two categories: fault avoidance and fault tolerance.
[0066] Among them, fault avoidance technologies involve probabilistic reliability-aware algorithms for task allocation, virtual machine placement, service scheduling, or resource provisioning in a distributed system. These technologies consider the failure probabilities of data, nodes, or communication links based on the Poisson distribution or by using historical data and combining the mean and variance of previous failures in the system. The goal of this fault tolerance method is usually to optimize latency, energy consumption, cost, and bandwidth.
[0067] Fault tolerance can be divided into active fault tolerance and passive fault tolerance. The main difference lies in the means of fault perception. Active fault tolerance means preprocessing faults through prediction. Passive fault tolerance means detecting the faults that have occurred in a timely manner and adopting fault handling strategies for fault response.
[0068] Whether it is active fault tolerance, passive fault tolerance, or fault avoidance, it mainly minimizes the impact of faults through mobile management and scheduling strategies. For the selection of fault tolerance strategies, the current main methods include resource redundancy, robust network design, and other immediate fault tolerance means based on fault perception.
[0069] Based on this, the main challenges faced by the current fault tolerance technologies of distributed systems are that the fault tolerance strategies only consider the reliability of the network and resources themselves, do not consider the continuous operation guarantee ability of the services during operation; the resource availability factor is not considered enough; they are highly reliable but lack flexibility.
[0070] To mitigate the problem of the unavailability of distributed systems caused by such faults, ensure that various services running in them do not experience large-scale interruptions and failures, and thus avoid significant losses to users, it is necessary to improve the fault tolerance ability of distributed systems to minimize the losses caused by faults, ensure the continuity of user services, and improve the reliability of distributed systems.
[0071] To improve the ability of fault-tolerant scheduling in a distributed system to ensure the continuous operation of the services provided in the distributed system, the fault-tolerant processing solution provided in the embodiments of the present invention monitors the status information of each resource node in the target distributed system. After detecting a faulty resource node, it determines the service tasks on the faulty resource node as tasks to be processed, and determines the priorities of the tasks to be processed, so as to perform fault-tolerant processing on the tasks to be processed in the order from highest to lowest priority; when performing fault-tolerant processing on the tasks to be processed, it uses the trained fault-tolerant processing action generation model to generate multiple fault-tolerant processing actions for the tasks to be processed according to the status information at the current moment, evaluates the availability of each fault-tolerant processing action, determines the actual fault-tolerant processing action according to the availability evaluation result, and performs the actual fault-tolerant processing action on the tasks to be processed, so that when a resource node fails in the distributed system, feasible scheduling can be performed on the tasks on the resource node to ensure the continuous execution of these tasks, and there will be no long-term service interruption due to the failure of the node where the tasks are located or the inability to continue providing services after migration to an inappropriate node, thus ensuring the continuous operation of the services provided by the distributed system and effectively improving the reliability of the distributed system.
[0072] The following describes the fault-tolerant processing method provided in the embodiments of the present invention with reference to the accompanying drawings.
[0073] Figure 1 It is a flowchart of a fault-tolerant processing method provided in the embodiments of the present invention; Figure 2 It is a schematic diagram of a distributed system fault-tolerant scheduling architecture provided in the embodiments of the present invention.
[0074] As Figure 1 shown, the fault-tolerant processing method provided in the embodiments of the present invention may include:
[0075] S101: Monitor the status information of each resource node in the target distributed system;
[0076] S102: After detecting a faulty resource node, determine the service tasks on the faulty resource node as tasks to be processed;
[0077] S103: Determine the priorities of the tasks to be processed, and perform fault-tolerant processing on the tasks to be processed in the order from highest to lowest priority;
[0078] S104: When performing fault-tolerant processing on the tasks to be processed, use the trained fault-tolerant processing action generation model to generate multiple fault-tolerant processing actions for the tasks to be processed according to the status information at the current moment, evaluate the availability of each fault-tolerant processing action, determine the actual fault-tolerant processing action according to the availability evaluation result, and perform the actual fault-tolerant processing action on the tasks to be processed.
[0079] In the embodiments of the present invention, as introduced in the above concept, the target distributed system can be any type of distributed system, such as a distributed computing system, a distributed storage system, an edge computing system, etc.
[0080] The fault tolerance processing method provided by the embodiments of the present invention can be applied to one or more resource nodes in the target distributed system, or can also be applied to one or more computing devices outside the target distributed system.
[0081] As Figure 2 shown, by deploying an agent decision network and an action filter in a computing device, by sensing the state information in the target distributed system, using the agent decision network to generate a fault tolerance processing action according to the state information, using the action filter to screen out the actual fault tolerance processing action from the fault tolerance processing actions, after executing the actual fault tolerance processing action in the target distributed system, continue to sense the state information in the target distributed system, calculate the environmental reward value using the reward function, and return to the agent decision network.
[0082] In a specific implementation, for S101, the real-time state information of the target distributed system can be collected through software and hardware detection devices deployed in the target distributed system, and potential or occurred system faults can be mined through data processing techniques to support the subsequent steps of fault tolerance processing for the fault location.
[0083] As Figure 2 shown, in the embodiments of the present invention, the state information may include but is not limited to basic environmental state information, task state information, and fault information.
[0084] Among them, the basic environmental state information may include the network topology information of the target distributed system, the resource state information of the resource nodes, and the network state information between the resource nodes.
[0085] The network topology information may include the connection relationships between the resource nodes in the target distributed system and between the resource nodes, and can be represented by a relationship graph wherein, represents the resource nodes in the target distributed system, represents the edges formed by the interconnection between the resource nodes.
[0086] The resource state information of the resource nodes may include information on the computing resources and storage resources on the resource nodes. Among them, the storage resources may include memory resources and persistent storage resources. Then the resource state information of the resource node can be as follows:
[0087] ;
[0088] wherein, , , respectively represent the total amount of computing resources, the available amount of computing resources, and the occupied amount of computing resources of the resource node at present, , , respectively represent the resource node in the total amount of memory resources, the available amount of memory resources, and the occupied amount of memory resources at present, , , respectively represent the resource node in the total amount of persistent storage resources, the available amount of persistent storage resources, and the occupied amount of persistent storage resources at present.
[0089] The network status information between resource nodes is used to represent the interconnection status between resource nodes. For the interconnection between resource nodes, can be used to represent the resource node and the resource node The interconnection status information between them, where represents the resource node to the resource node The actual bandwidth between them, represents the average delay level. The actual bandwidth capacity represents the data transmission capacity, while the delay level is affected by the position in the network topology between resource nodes, the link distance, and the data transceiver delay caused by other external interferences. Therefore, changes dynamically with the system.
[0090] The task status information may include information about various (user service) business tasks deployed on each resource node. A business task can be regarded as a task with a time duration. can be used to represent at the moment the set of various business tasks running on the resource node . represents the information of one of the business tasks, which may include:
[0091] ;
[0092] Among them, , , respectively represent the amount of computing resources, the amount of memory resources, and the amount of persistent storage resources occupied by a single copy of the business task . represents the number of copies of the business task running on the resource node where it is located. According to the service information, its total consumption on the resource node where it is located can be counted. Indicates a business task The occupation period of the current resource. Among them, the occupation period can refer to the time used for a business task to be completely executed once on the resource node where it is located.
[0093] The fault information can include the type of fault that occurs in the resource node and the duration of the fault. Resource node 's fault information Can be expressed as:
[0094] ;
[0095] Among them, Indicates the type of fault, The status quantity of can include that the resource node has no fault (which can be represented by 0), the resource node has a fault (which can be represented by n, and the specific value of n represents the specific type of fault), and the network related to the resource node has a fault (which can be represented by e, and the specific value of e represents the specific network or resource node pointed to). Among them, the network related to the resource node can include the network composed of all resource nodes that have a connection relationship with the resource node. For other resource nodes directly connected to the resource node, other resource nodes interconnected with the resource node and other resource nodes unidirectionally pointed to by the resource node can be included in the network related to the resource node; for other resource nodes not directly connected to the resource node, a hop count threshold can be preset, and other resource nodes with a connection to the resource node less than or equal to the hop count threshold can be included in the network related to the resource node, and other resource nodes with a connection to the resource node greater than the hop count threshold are not considered.
[0096] Indicates the duration of the fault, The status quantity of can include the predicted duration of the fault (which can be represented by T, and the specific value of T can be a specific duration) and the duration of the fault is unknown (which can be represented by TE). The duration of the fault being unknown means that the current fault occurring in the resource node cannot be recovered and the duration of the fault cannot be determined.
[0097] By statistically analyzing the historical data of each resource node, the law of the duration of the fault corresponding to various types of faults can be obtained. In the embodiments of the present invention, can be used To represent the set of actual fault types that occur on the resource node And use To represent the probability distribution model of the duration of the fault corresponding to the actual fault type that occurs on the resource node The latter can be obtained by statistically analyzing historical data.
[0098] Then in S101, by sensing the above state information, the quantified sensed target data can be obtained Among them, Represents the network topology information of the target distributed system, Represents a resource node of the resource status information, Represents a resource node and a resource node of the interconnection status information therebetween, Represents at a moment the set of various business tasks running on a resource node , Represents a resource node of the fault information.
[0099] The intelligent agent decision network manages the resource nodes with faults according to the sensed target data, that is, performs fault tolerance scheduling on the business tasks on the faulty resource nodes.
[0100] At the current moment, there may be multiple faulty resource nodes in the target distributed system, and there may be multiple business tasks on a faulty resource node. According to the device scheduling capability and system design, the to-be-processed tasks can be scheduled one by one, or multiple to-be-processed tasks can be scheduled at a time and the scheduling can be completed in multiple batches. Therefore, it is necessary to determine the priorities of the to-be-processed tasks and perform fault tolerance processing on the to-be-processed tasks in the order from high to low according to the priorities. The way to determine the priorities of the to-be-processed tasks can be determined according to at least one of the task type, workload, and user-set parameters of the to-be-processed tasks.
[0101] The fault tolerance processing action generation model can be trained by means of reinforcement learning or deep reinforcement learning. The trained fault tolerance processing action generation model is used to generate multiple fault tolerance processing actions for the to-be-processed tasks according to the state information at the current moment. During operation, the parameters of the fault tolerance processing action generation model can also be fine-tuned according to the operation state of the target distributed system.
[0102] To improve the availability of the fault tolerance scheduling result, in the embodiment of the present invention, an action filter is designed to screen the multiple fault tolerance processing actions output by the fault tolerance processing action generation model, and based on the pre-configured availability evaluation index, to ensure that the actual fault tolerance processing actions available in the target distributed system at the current moment are screened and executed.
[0103] When there are still to-be-processed tasks in the target distributed system for which fault tolerance scheduling has not been performed, then continue to execute S104 for the next or the next few to-be-processed tasks. When processing the next or the next few to-be-processed tasks, enter the next moment, that is, the state information of the target distributed system is updated, and at this time, execute S104 according to the state information at the next moment.
[0104] The fault tolerance processing method provided by the embodiments of the present invention monitors the status information of each resource node of the target distributed system. After detecting a faulty resource node, it determines the service tasks on the faulty resource node as tasks to be processed, and determines the priorities of the tasks to be processed, so as to perform fault tolerance processing on the tasks to be processed in the order from high to low priority; when performing fault tolerance processing on the tasks to be processed, it uses the trained fault tolerance processing action generation model to generate multiple fault tolerance processing actions for the tasks to be processed according to the status information at the current moment, evaluates the availability of each fault tolerance processing action, determines the actual fault tolerance processing action according to the availability evaluation result, and performs the actual fault tolerance processing action on the tasks to be processed. Therefore, when a resource node in the distributed system fails, feasible scheduling can be performed on the tasks on the resource node to ensure the continuous execution of these tasks, and there will be no long-term service interruption due to the failure of the node where the tasks are located or the inability to continue providing services after migration to an inappropriate node, thus ensuring the continuous operation of the services provided by the distributed system and effectively improving the reliability of the distributed system.
[0105] Based on the above embodiments, the embodiments of the present invention further describe the steps of determining the priorities of the tasks to be processed.
[0106] In some optional embodiments of the embodiments of the present invention, determining the priorities of the tasks to be processed in S103 may include: determining the fault impact degree parameter of the faulty resource node according to the fault type of the faulty resource node; determining the task evaluation parameter of the task to be processed according to the load status of the task to be processed on the faulty resource node; determining the priority score of the task to be processed according to the fault impact degree parameter and the task evaluation parameter.
[0107] The embodiments of the present invention measure the priorities of the tasks to be processed from two aspects: the fault impact degree of the faulty resource node and the load status of the tasks to be processed on the faulty resource node.
[0108] Among them, determining the fault impact degree parameter of the faulty resource node according to the fault type of the faulty resource node may include: determining the predicted value of the fault duration of the faulty resource node according to the historical operation data of the faulty resource node; using the predicted value of the fault duration as the fault impact degree parameter.
[0109] In the above embodiments of the present invention, by statistically analyzing the historical data of each resource node, the pattern of the fault duration corresponding to various fault types can be obtained. Then, determining the predicted value of the fault duration of the faulty resource node according to the historical data corresponding to the fault type of the faulty resource node may include: determining the expected value of the fault duration of the faulty resource node according to the fault type of the faulty resource node and the corresponding probability distribution model of the fault duration; using the expected value of the fault duration as the predicted value of the fault duration; wherein, the probability distribution model of the fault duration corresponding to the fault type is determined according to the historical operation data of the resource node.
[0110] Then, in the embodiments of the present invention, the expected value of the fault duration can be obtained through calculation, that is, under the condition of the fault type actually occurring on the resource node at a given moment, according to the probability distribution model of the fault duration corresponding to these fault types, the expected value of the fault duration of the resource node is determined.
[0111] In some alternative embodiments of the embodiments of the present invention, determining the task evaluation parameter of the task to be processed according to the load status of the task to be processed on the faulty resource node may include: using the resource occupancy parameter of the task to be processed on the faulty resource node and the remaining execution time of the task to be processed on the faulty resource node as the load status of the task to be processed on the faulty resource node, and determining the task evaluation parameter according to the resource occupancy parameter and the remaining execution time. That is to say, the task evaluation parameter of the task to be processed can be determined only by considering the workload.
[0112] Among them, the step of determining the resource occupancy parameter may include: determining the resource occupancy parameter according to the resource amounts of multiple types of resources occupied by the task to be processed on the faulty resource node. Determining the task evaluation parameter according to the resource occupancy parameter and the remaining execution time may include: using the ratio of the resource occupancy parameter to the remaining execution time as the task evaluation parameter. Then, the task evaluation parameter of the task to be processed can be calculated by the following formula:
[0113] ;
[0114] Among them, , , respectively represent the computing resource amount, memory resource amount and persistent storage resource amount occupied by a single copy of the service task , represents the number of copies of the service task running on the resource node where it is located, represents the occupation period of the service task for the current resource. Among them, is in the same unit, , , , , They can all be parameters after normalization processing. Then It can be used to represent the average value of resources consumed by the task to be processed on the faulty resource node per unit time.
[0115] In the embodiments of the present invention, the fault impact degree parameter can be the prediction of the fault duration of the faulty resource node, and the task evaluation parameter can be the resource consumption value per unit time of the task to be processed on the faulty resource node. Then, determining the priority score of the task to be processed according to the fault impact degree parameter and the task evaluation parameter may include: using the ratio of the resource consumption value per unit time to the predicted value of the fault duration as the priority score. Then, the priority score of the task to be processed can be calculated by the following formula:
[0116] ;
[0117] Wherein, represents the priority score of the task to be processed on the faulty resource node at time .
[0118] When comparing the priority scores, all tasks to be processed on all faulty resource nodes in the target distributed system can be sorted uniformly, or grouped according to the parallel ability of fault-tolerant scheduling, and the priority scores can be sorted within the groups.
[0119] In some other alternative embodiments of the embodiments of the present invention, determining the task evaluation parameter of the task to be processed according to the load status of the task to be processed on the faulty resource node may further include: determining the task evaluation parameter according to the load status of the task to be processed on the faulty resource node and the task type priority parameter of the task to be processed. That is to say, when determining the priority score of the task to be processed, the task type priority of the task to be processed can also be combined. Specifically, the task type priority score of the task to be processed can be added to the calculation formula of the priority score of the task to be processed.
[0120] Based on the above embodiments, the embodiments of the present invention further illustrate the fault-tolerant processing action generation model.
[0121] In the above embodiments of the present invention, it is introduced that the fault tolerance processing action generation model can be trained by means of reinforcement learning or deep reinforcement learning. Then, the training steps of the fault tolerance processing action generation model may include: inputting sample state information into the fault tolerance processing action generation model to output sample fault tolerance processing actions for the sample state information; after executing the sample fault tolerance processing actions, obtaining the updated sample state information; determining the environmental reward value according to the updated sample state information and the reward function; using a set of sample fault tolerance processing actions, sample fault tolerance processing actions, updated sample state information, and environmental reward value as a training sample; and updating the model parameters of the fault tolerance processing action generation model by using the training sample.
[0122] In the embodiments of the present invention, the environmental reward value may adopt the fault tolerance cost of executing the fault tolerance processing action. In some optional implementation manners of the embodiments of the present invention, the fault tolerance cost of executing the fault tolerance processing action may be measured from the perspective of the time cost required to execute the fault tolerance processing action, and the reward function may be a function of the fault tolerance processing time of the task.
[0123] As Figure 2 shown, the fault tolerance processing actions can be divided into three categories, namely task migration actions, task waiting actions, and task scaling actions. The time cost required to execute the fault tolerance processing action can be determined by combining the service interruption time and transmission time caused by executing the fault tolerance processing action. Then, for the above three types of actions:
[0124] Task migration action: Migrate the task to be processed on the current faulty resource node to a normally operating resource node. The main time cost generated in this process includes data transmission time and service recovery time.
[0125] Task waiting action: When the fault can be recovered in a short time, migration may not be performed, and the task to be processed is interrupted until the fault of the faulty resource node is recovered. The main cost of this process is the waiting time of service interruption.
[0126] Task scaling action: If there are identical copies of the task to be processed on other resource nodes, the service on other resource nodes can be scaled, and the service of the faulty resource node is immediately terminated. This process does not generate data migration, and the main cost is the service recovery time.
[0127] The fault tolerance processing time may include a first processing time for task migration actions, a second processing time for task waiting actions, and a third processing time for task scaling actions; wherein, the task migration action is to migrate the task to be processed from the faulty resource node to the first resource node, and the first processing time includes data transmission time and service recovery time at the first resource node; the task waiting action is to set the task to be processed to the service interruption state at the faulty resource node, and the second processing time is the service interruption waiting time; the task scaling action is to start the task to be processed at the second resource node where there is a copy of the task to be processed and terminate the task to be processed at the faulty resource node, and the third processing time is the service recovery time at the second resource node.
[0128] According to the training steps of the fault tolerance processing action generation model and the types of fault tolerance processing actions introduced in the embodiments of the present invention, the following introduces a specific training step of the fault tolerance processing action generation model.
[0129] Let The state parameters of the target distributed system observed at time be where represents the basic environmental state information of the target distributed system at time including network topology information, resource state information of resource nodes, and network state information between resource nodes, represents the set of various business tasks currently running on all resource nodes in the target distributed system at time where represents the fault information of all faulty resource nodes in the target distributed system at time
[0130] Let The decision variable at time be where represents the task migration action, represents the task waiting action,
[0131] where it can be configured that when and , are both 0, it means to migrate the task to be processed to the resource node . When and , are both 0, it means to keep the task to be processed on the faulty resource node for waiting without performing other operations. When and , are both 0, it means to scale the task to be processed on another resource node .
[0132] After performing the fault tolerance processing action, the state information in the target distributed system is collected again, and the environmental reward value that can be obtained represents the time cost required to execute the fault tolerance processing action to restore the service of the task to be processed. Among them, , represents the first processing time of the task migration action, represents the second processing time of the task waiting action, represents the third processing time of the task scaling action. Usually, since the decision variable is only one type of fault tolerance processing action can be executed within a period of time, and the time costs of the three types of fault tolerance processing actions are mutually exclusive.
[0133] Construct an online decision-making generation network for deep reinforcement learning, that is, an initial fault tolerance processing action generation model. The fault tolerance processing action generation model can adopt a deep Q-learning model (DQN). Denote the initial fault tolerance processing action generation model as , is the model parameter of the fault tolerance processing action generation model. Input the state parameter into the fault tolerance processing action generation model, output the fault tolerance processing action and after execution, calculate the environmental reward value of the current stage through the reward function , and the system enters the next observation state .
[0134] Collect multiple training samples through the above method and store them in the cache pool D. When the number of training samples stored in the cache pool D reaches the preset sample number, enter the parameter update step of the fault tolerance processing action generation model.
[0135] During the parameter update process of the fault tolerance processing action generation model, randomly select M training samples from the cache pool D, and calculate the target Q value according to the following formula:
[0136] ;
[0137] Among them, is a real coefficient between 0 and 1.
[0138] The loss function of the fault tolerance processing action generation model can be:
[0139] .
[0140] Update the network parameters according to gradient backpropagation , until the network converges. After convergence, the fault tolerance processing action generation model outputs a decision-making action, i.e., a fault tolerance processing action, according to the obtained state information of the target distributed system at the current moment, and makes scheduling management decisions for the tasks to be processed on the faulty resource nodes in different stages.
[0141] Based on the above embodiments, the embodiments of the present invention further illustrate the action filter.
[0142] To ensure the availability of the actually executed fault tolerance processing actions, as Figure 2 shown, the embodiments of the present invention first use the fault tolerance processing action generation model to generate multiple fault tolerance processing actions, and then use the action filter to screen out the available actual fault tolerance processing actions from them.
[0143] In the embodiments of the present invention, in S104, the availability of each fault tolerance processing action is evaluated, and the actual fault tolerance processing action is determined according to the availability evaluation result, which may include: evaluating the availability of the fault tolerance processing actions in the order of the corresponding environmental reward values from large to small until a fault tolerance processing action that meets the availability evaluation conditions is obtained as the actual fault tolerance processing action. That is to say, the fault tolerance processing action with the largest environmental reward value can be selected from the available fault tolerance processing actions as the actual fault tolerance processing action.
[0144] In a specific implementation, using the action filter to evaluate the availability of the fault tolerance processing actions may include evaluating the security of the fault tolerance processing actions. If the fault tolerance processing action meets the preset security conditions, it is determined that the fault tolerance processing action meets the availability evaluation conditions. Then, it can be set indicating the security action indicator at time, first sorting the fault tolerance processing actions output by the fault tolerance processing action generation model in the order of the environmental reward values from large to small, selecting the first K fault tolerance processing actions, and then checking one by one in the order of the environmental reward values from large to small whether they meet the availability evaluation conditions. If not, then , until the fault tolerance processing action is selected as the actual fault tolerance processing action and executed.
[0145] Based on the above embodiments, in the process of obtaining the actual fault tolerance processing actions by using the fault tolerance processing action generation model and the action filter, in addition to the time cost, factors such as energy consumption cost and service quality can also be considered.
[0146] Then, in some optional embodiments of the present invention, when the reward function of the fault tolerance processing action generation model is a function of time cost, using the action filter to evaluate the availability of the fault tolerance processing action may include: sorting the fault tolerance processing actions output by the fault tolerance processing action generation model from largest to smallest according to the environmental reward value, selecting the top K fault tolerance processing actions, respectively determining the energy consumption cost and service quality parameters for each fault tolerance processing action, and determining the availability evaluation result of the fault tolerance processing action according to the environmental reward value, energy consumption cost and service quality parameters. The fault tolerance processing action with the largest availability evaluation result is used as the actual fault tolerance processing action. Among them, the energy consumption cost can be determined according to the computing resources, storage resources, network resources and power consumption required to execute the fault tolerance processing action to recover the task to be processed. The service quality parameter can be determined according to the execution result of the task to be processed after executing the fault tolerance processing action to recover the task to be processed, and specifically can be the model accuracy.
[0147] Through the above solution of generating the actual fault tolerance processing action by the fault tolerance processing action generation model and the action filter, the training process and inference process of the fault tolerance processing action generation model can be simplified, and the optimal actual fault tolerance processing action can be quickly obtained under the multi-dimensional objectives of comprehensive time cost, energy consumption cost and service quality.
[0148] In some other optional embodiments of the present invention, the reward function of the fault tolerance processing action generation model can also be set as a function of time cost, energy consumption cost and service quality parameters. Then, using the action filter to evaluate the availability of the fault tolerance processing action may include: sorting the fault tolerance processing actions output by the fault tolerance processing action generation model from largest to smallest according to the environmental reward value, selecting the top K fault tolerance processing actions, and then checking one by one in the order of environmental reward value from largest to smallest whether it meets the preset safety conditions, then determining that the fault tolerance processing action meets the availability evaluation conditions. The preset safety conditions may include: if the fault tolerance processing action is a task migration action or a task scaling action, then when the task migrates or scales to another resource node to execute the fault tolerance processing action, it meets the requirements for executing the task to be processed, including normal running status and having the resources required to execute the task to be processed.
[0149] Through the above solution of generating the actual fault tolerance processing action by the fault tolerance processing action generation model and the action filter, on the basis of generating the fault tolerance processing action with multi-dimensional objectives of comprehensive time cost, energy consumption cost and service quality, by re-checking whether the fault tolerance processing action meets the preset safety conditions before the actual execution of the fault tolerance processing action, it is ensured to select the actual fault tolerance processing action with the optimal environmental reward and high availability.
[0150] It should be noted that, in the embodiments of each fault tolerance processing method of the present invention, some of the steps or features may be ignored or not executed. The hardware or software function modules divided for convenience of description are not the only implementation forms for implementing the fault tolerance processing method provided by the embodiments of the present invention.
[0151] The above details various embodiments corresponding to the fault tolerance processing method. On this basis, the present invention also discloses a fault tolerance processing device, a device, a non-volatile storage medium, and a computer program product corresponding to the above method.
[0152] The fault tolerance processing device provided by the embodiments of the present invention may include:
[0153] A monitoring unit, configured to monitor the status information of each resource node of the target distributed system; after detecting a faulty resource node, determine the service task on the faulty resource node as a task to be processed;
[0154] A task priority determination unit, configured to determine the priority of the task to be processed, and perform fault tolerance processing on the task to be processed in the order from high to low according to the priority;
[0155] A fault tolerance action generation unit, configured to, when performing fault tolerance processing on the task to be processed, use the trained fault tolerance processing action generation model to generate multiple fault tolerance processing actions for the task to be processed according to the status information at the current moment, perform availability evaluation on each fault tolerance processing action, determine the actual fault tolerance processing action according to the availability evaluation result, and perform the actual fault tolerance processing action on the task to be processed.
[0156] It should be noted that, in each implementation manner of the fault tolerance processing device provided by the embodiments of the present invention, the division of the units is only a logical functional division, and other division methods may be adopted. The connection manners between different units may adopt electrical, mechanical, or other connection manners. The separated units may be located at the same physical location or distributed on multiple network nodes. Each unit may be implemented in the form of hardware or may adopt the form of a software function unit. That is, according to actual needs, some or all of the units provided by the embodiments of the present invention may be selected and the corresponding connection manners or integration manners may be adopted to achieve the purpose of the solution of the embodiments of the present invention.
[0157] Since the embodiments of the device part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the device part, and details are not described here for the time being.
[0158] Figure 3 It is a schematic structural diagram of a fault tolerance processing device provided by the embodiments of the present invention.
[0159] As Figure 3As shown in the figure, the fault-tolerant processing device provided by the embodiment of the present invention includes: a memory 310 for storing a computer program 311; a processor 320 for executing the computer program 311, and when the computer program 311 is executed by the processor 320, it implements the steps of the fault-tolerant processing method provided by any of the above embodiments.
[0160] Among them, the processor 320 may include one or more processing cores, such as a 3-core processor, an 8-core processor, etc. The processor 320 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array. The processor 320 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 320 may be integrated with a graphics processing unit (GPU), and the graphics processing unit is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 320 may also include an artificial intelligence (AI) processor, and the artificial intelligence processor is used to process computational operations related to machine learning.
[0161] The memory 310 may include one or more non-volatile storage media, and the non-volatile storage media may be non-transitory. The memory 310 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory 310 is at least used to store the following computer program 311. After the computer program 311 is loaded and executed by the processor 320, it can implement the relevant steps in the fault-tolerant processing method disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 310 may also include an operating system 312 and data 313, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 312 may be Windows or other types of operating systems. The data 313 may include, but is not limited to, the data involved in the above method.
[0162] In some embodiments, the fault-tolerant processing device may further include a display screen 330, a power supply 340, a communication interface 350, an input / output interface 360, a sensor 370, and a communication bus 380.
[0163] Those skilled in the art can understand,Figure 3 The structure shown does not constitute a limitation on the fault-tolerant processing device, and may include more or fewer components than those shown in the figure.
[0164] The fault-tolerant processing device provided by an embodiment of the present invention includes a memory and a processor. When the processor executes the program stored in the memory, it can implement the steps of the fault-tolerant processing method provided in the above embodiment, and the effect is the same as above.
[0165] An embodiment of the present invention provides a non-volatile storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can implement the steps of the fault-tolerant processing method provided in any one of the above embodiments.
[0166] The non-volatile storage medium may include: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.
[0167] For the introduction of the non-volatile storage medium provided by an embodiment of the present invention, please refer to the above method embodiment, and the effect it achieves is the same as that of the fault-tolerant processing method provided by an embodiment of the present invention. The present invention will not be elaborated here.
[0168] An embodiment of the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the fault-tolerant processing method provided in any one of the above embodiments.
[0169] For the introduction of the computer program product provided by an embodiment of the present invention, please refer to the above method embodiment, and the effect it achieves is the same as that of the fault-tolerant processing method provided by an embodiment of the present invention. The present invention will not be elaborated here.
[0170] The above has provided a detailed introduction to a fault-tolerant processing method, device, medium, and computer program product provided by the present invention. The various embodiments in the specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices, equipment, non-volatile storage media, and computer program products disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple. For the relevant parts, please refer to the description in the method part. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can still be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.
[0171] It should also be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising that element.
Claims
1. A fault-tolerant processing method, characterized in that: include: Monitor the status information of each resource node of the target distributed system; After detecting a faulty resource node, determining that a service task on the faulty resource node is a task to be processed; Determine the priority of the pending tasks, and perform fault-tolerant processing on the pending tasks in order of priority from high to low; When performing fault-tolerant processing on the task to be processed, a trained fault-tolerant processing action generation model is used to generate multiple fault-tolerant processing actions for the task to be processed according to the state information at the current moment, an availability evaluation is performed on each of the fault-tolerant processing actions, an actual fault-tolerant processing action is determined according to the availability evaluation result, and the actual fault-tolerant processing action is performed on the task to be processed; The process of performing availability evaluation on each of the fault-tolerant processing actions and determining the actual fault-tolerant processing action according to the availability evaluation result includes: The fault-tolerant processing actions are sorted from large to small according to the environmental reward values calculated by the fault-tolerant processing action generation model, the first K fault-tolerant processing actions are selected, and the availability evaluation results of the fault-tolerant processing actions are determined according to the environmental reward values, energy consumption costs and service quality parameters of the fault-tolerant processing actions, and the fault-tolerant processing action with the largest availability evaluation result is taken as the actual fault-tolerant processing action; the energy consumption cost is determined according to the computing resources, storage resources, network resources and power consumption required for executing the fault-tolerant processing action to restore the pending task; the service quality parameter is determined according to the execution result of the pending task after executing the fault-tolerant processing action to restore the pending task; The environmental reward value is determined according to the time required to execute the fault-tolerant processing action to restore the pending task; the fault-tolerant processing action includes a task migration action, a task waiting action and a task scaling action.
2. The fault-tolerant processing method according to claim 1, characterized in that: Determine the priority of the pending tasks, including: Determine a fault impact parameter of the faulty resource node according to the fault type of the faulty resource node; Determine a task evaluation parameter of the task to be processed according to a load state of the task to be processed on the faulty resource node; The priority score of the task to be processed is determined according to the fault impact parameter and the task evaluation parameter.
3. The fault-tolerant processing method according to claim 2, characterized in that: Determining a fault impact parameter of the faulty resource node according to the fault type of the faulty resource node includes: Determine a predicted value of the fault duration of the faulty resource node according to historical operation data of the faulty resource node; The predicted value of the fault duration is used as the fault impact parameter.
4. The fault-tolerant processing method according to claim 3, characterized in that: Determining a predicted value of a fault duration of the faulty resource node according to historical data corresponding to the fault type of the faulty resource node includes: Determine an expected value of the fault duration of the faulty resource node according to the fault type of the faulty resource node and a corresponding fault duration probability distribution model; Using the expected value of the fault duration as the predicted value of the fault duration; The fault duration probability distribution model corresponding to the fault type is determined according to the historical operation data of the resource node.
5. The fault-tolerant processing method according to claim 2, characterized in that: Determining a task evaluation parameter of the task to be processed according to a load state of the task to be processed on the faulty resource node includes: The resource occupancy parameter of the task to be processed on the faulty resource node and the remaining execution time of the task to be processed on the faulty resource node are used as the load state of the task to be processed on the faulty resource node, so as to determine the task evaluation parameter according to the resource occupancy parameter and the remaining execution time.
6. The fault-tolerant processing method according to claim 5, characterized in that: The step of determining the resource occupancy parameter comprises: Determine the resource occupation parameter according to the amount of multiple types of resources occupied by the task to be processed on the faulty resource node; Determining the task evaluation parameter according to the resource occupancy parameter and the remaining execution time includes: The ratio of the resource occupancy parameter to the remaining execution time is used as the task evaluation parameter.
7. The fault-tolerant processing method according to claim 2, characterized in that: Determining a task evaluation parameter of the task to be processed according to a load state of the task to be processed on the faulty resource node includes: The task evaluation parameter is determined according to the load state of the task to be processed on the faulty resource node and the task type priority parameter of the task to be processed.
8. The fault-tolerant processing method according to claim 2, characterized in that: The fault impact parameter is a predicted value of the fault duration of the fault resource node; The task evaluation parameter is the resource consumption per unit time of the task to be processed at the faulty resource node; Determining the priority score of the task to be processed according to the fault impact parameter and the task evaluation parameter includes: The ratio of the resource consumption value per unit time to the predicted fault duration value is used as the priority score.
9. The fault-tolerant processing method according to claim 1, characterized in that: The training step of the fault-tolerant processing action generation model includes: Inputting sample state information into the fault-tolerant processing action generation model, and outputting a sample fault-tolerant processing action for the sample state information; After executing the sample fault tolerance processing action, obtaining updated sample status information; Determine an environment reward value according to the updated sample state information and reward function; Taking a group of the sample fault-tolerant processing actions, the sample fault-tolerant processing actions, the updated sample state information and the environmental reward value as a training sample; The model parameters of the fault-tolerant processing action generation model are updated using the training samples.
10. The fault-tolerant processing method according to claim 9, characterized in that: The reward function is a function of the fault-tolerant processing time of the task.
11. The fault-tolerant processing method according to claim 10, characterized in that: The fault-tolerant processing time includes a first processing time for a task migration action, a second processing time for a task waiting action, and a third processing time for a task scaling action; Among them, the task migration action is to migrate the pending task from the faulty resource node to the first resource node, and the first processing time includes the data transmission time and the service recovery time at the first resource node; the task waiting action is to set the pending task to a service interruption state at the faulty resource node, and the second processing time is the service interruption waiting time; the task scaling action is to start the pending task at the second resource node where a copy of the pending task exists and terminate the pending task at the faulty resource node, and the third processing time is the service recovery time at the second resource node.
12. A fault-tolerant processing device, characterized in that: include: Memory for storing computer programs; A processor is used to execute the computer program, and when the computer program is executed by the processor, the steps of the fault-tolerant processing method according to any one of claims 1 to 11 are implemented.
13. A non-volatile storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the fault-tolerant processing method according to any one of claims 1 to 11 are implemented.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the fault-tolerant processing method according to any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Lean management method and system for intelligent power distribution equipment
CN112862249A
Distributed automatic expansion method in high-performance computing
CN118838735A