Computing power operation task fault tolerance method and device of intelligent computing center
By monitoring the node status of the intelligent computing center, determining abnormal pods and restarting the task, the problem of task operation errors is solved, and the normal recovery of tasks and efficient utilization of resources is achieved.
Patent Information
- Application Number
- CN202510344706.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-18
AI Technical Summary
In the Intelligent Computing Center, how to recover the normal operation problems of tasks caused by task operation errors.
By monitoring the task running status of the node, determining the abnormal node and the Pod, and restarting the target computing power running task corresponding to the target Pod.
It realizes the fault tolerance mechanism of the intelligent computing center when a task error occurs, restores the normal operation of the error task, improves operation stability, and reduces waste of computing resources.
Smart Images

Figure CN120336086A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers, and computing power infrastructure, and particularly relates to a method and device for fault tolerance of computing power operation tasks in an intelligent computing center. Background Art
[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.
[0003] An "intelligent computing center" refers to a facility that provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios of artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power and intelligent computing power. The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0004] The "intelligent computing center" includes, but is not limited to, the "intelligent computing center".
[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications based on artificial intelligence theory and using an artificial intelligence computing architecture.
[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers", which is the ability of computer devices or computing / data centers to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of the target result by processing information data, and a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.
[0007] Currently, for an intelligent computing center, in the process of executing a computing power operation task, if the task runs into an error, how to restore the normal operation of the error-prone task is an urgent problem to be solved. Summary of the Invention
[0008] The present invention provides a method and device for fault tolerance of computing power operation tasks in an intelligent computing center to solve the problem of how to restore the normal operation of an error-prone task in an intelligent computing center.
[0009] To solve the above technical problems, the present invention is implemented as follows:
[0010] In a first aspect, the present invention provides a method for fault tolerance of computing power operation tasks in an intelligent computing center, including:
[0011] Step S1: Monitor and obtain the first information corresponding to each node, where the first information is used to indicate whether the task running status of the corresponding node is abnormal;
[0012] Step S2: When there is target information among the first information corresponding to all the nodes, determine the target node corresponding to the target information, where the target information indicates that the task running of the target node is abnormal;
[0013] Step S3: Based on the target node, determine the target Pod with an abnormal task running status;
[0014] Step S4: Restart the target computing power running task corresponding to the target Pod.
[0015] In one embodiment, step S4 includes:
[0016] Step S41: Delete the target Pod located on the target node;
[0017] Step S42: Create the target Pod on any other node except the target node;
[0018] Step S43: Instruct the target Pod to execute the target computing power running task.
[0019] In one embodiment, step 43 includes:
[0020] Step S431: Obtain the task running status data corresponding to the target Pod, where the task running status data is used to indicate the execution progress of the target computing power running task at the target moment, and the target moment is the nearest data backup moment from the moment when the target Pod runs abnormally;
[0021] Step S432: Based on the task running status data, instruct the target Pod to execute the target computing power running task.
[0022] In one embodiment, step S3 includes:
[0023] Step S31: Obtain the task label corresponding to the abnormal task of the target node;
[0024] Step S32: Based on the task label, determine N Pods corresponding to the task executed by the target node, where N is a positive integer;
[0025] Step S33: Obtain the second information corresponding to each of the N Pods, where the second information is used to indicate the running status of the corresponding Pod;
[0026] Step S34: Traverse the second information corresponding to each of the N Pods, and determine the target Pod with an abnormal running status.
[0027] In one embodiment, the first information includes at least one of the following:
[0028] The first status information, which is used to indicate that the corresponding node is in a normal state;
[0029] The second status information, which is used to indicate that the corresponding node is in a down state or a network disconnection state or a user suspension state;
[0030] The third status information, which is used to indicate that the corresponding node is in a task abnormal exit state.
[0031] In one embodiment, after the step S4, the method further includes:
[0032] Step S5: Generate a prompt message, where the prompt message is used to prompt that the computing power running task corresponding to the target Pod has been restarted.
[0033] In a second aspect, the present invention provides a computing power running task fault tolerance device for an intelligent computing center, including:
[0034] An acquisition module, configured to monitor and acquire the first information corresponding to each of all nodes, where the first information is used to indicate whether the task running status of the corresponding node is abnormal;
[0035] A first determination module, configured to determine a target node corresponding to the target information in the case that the target information exists in the first information corresponding to each of all nodes, where the target information indicates that the task running of the target node is abnormal;
[0036] A second determination module, configured to determine a target Pod with an abnormal running status based on the target node;
[0037] A restart module, configured to restart the target computing power running task corresponding to the target Pod.
[0038] In a third aspect, the present invention provides an electronic device, including: a processor, a memory, and a program stored on the memory and executable on the processor, where when the program is executed by the processor, the steps of the computing power running task fault tolerance method for the intelligent computing center as described in the first aspect above are implemented.
[0039] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the computing power running task fault tolerance method for the intelligent computing center as described in the first aspect above are implemented.
[0040] Fifth aspect, the present invention provides a computer program product, including computer instructions, which when executed by a processor implement the steps of the method for fault tolerance of computing power operation tasks of the intelligent computing center as described in the first aspect above.
[0041] In the present invention, by monitoring the first information corresponding to each node, when there is target information in the first information, the target node corresponding to the target information can be determined in a timely manner. The target node is the node with an abnormal task running state. Then, according to the target node, the target Pod with an abnormal running state is determined, and the target computing power operation task corresponding to the target Pod is restarted, so that the intelligent computing center can resume the normal operation of the failed task. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0043] Figure 1 is a schematic flowchart of a method for fault tolerance of computing power operation tasks of an intelligent computing center provided by an embodiment of the present invention;
[0044] Figure 2 is a schematic diagram of a node architecture provided by an embodiment of the present invention;
[0045] Figure 3 is another schematic diagram of a node architecture provided by an embodiment of the present invention;
[0046] Figure 4 is a schematic diagram of the structure of a device for fault tolerance of computing power operation tasks of an intelligent computing center provided by an embodiment of the present invention;
[0047] Figure 5 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] Next, the technical solutions in the present invention will be clearly and completely described in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0049] First, the technical terms related to the present invention will be briefly described below.
[0050] The "computing power" described in the present invention refers to: the ability of computer devices or computing / data centers to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of target results through processing information data, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.
[0051] The "computational power" (Computational Power, CP) described in the present invention refers to: the ability of electronic devices in a data center to process data and achieve result output, a comprehensive indicator for measuring the computing ability of a data center, including general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing ability. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream electronic device CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP_general + CP_intelligent + CP_super
[0052] The "network power" (Network Power, NP) described in the present invention refers to: the manifestation of the data transmission ability of computing power facilities, a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., involving network transmission within and between data centers, and a comprehensive indicator for measuring network transmission scheduling ability.
[0053] The "storage power" (Storage Power, SP) described in the present invention refers to: the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon, a comprehensive indicator for measuring the data storage ability of a data center, including external storage devices such as storage arrays and internal storage devices of electronic devices. The commonly used measurement unit for storage capacity is exabyte (EB, 1 EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.
[0054] The "computing power infrastructure" described in the present invention refers to: a new type of information infrastructure integrating information computing power, network carrying capacity, and data storage capacity, which can realize the centralized computing, storage, transmission, and application of information.
[0055] The "new information infrastructure" described in the present invention refers to: mainly including network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, satellite Internet, etc., computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, supercomputing centers, etc., and new technology facilities such as artificial intelligence, blockchain, and quantum computing.
[0056] The "computing power" described in the present invention includes: general computing power, intelligent computing power, and super computing power.
[0057] The "general computing power" described in the present invention refers to: the computing power provided by electronic devices based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0058] The "intelligent computing power" described in the present invention refers to: for various artificial intelligence innovation applications, a computing platform is deployed on a large scale based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing, machine vision, etc.
[0059] The "super computing power" described in the present invention refers to: mainly the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.
[0060] The "intelligent computing center" described in the present invention refers to: a facility that provides the required computing power, data, and algorithms mainly for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0061] The "intelligent computing center" described in the present invention includes but is not limited to the "intelligent computing center".
[0062] The "intelligent computing center" described in the present invention, that is, the artificial intelligence computing center, is a type of computing power infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications based on artificial intelligence theory and using an artificial intelligence computing architecture.
[0063] The "computing power center" described in the present invention refers to a facility mainly composed of infrastructure such as wind, fire, water, and electricity, and IT software and hardware devices, with computing power, transportation power, and storage power, including general data centers, intelligent computing centers, supercomputing centers, etc.
[0064] The "supercomputing center" described in the present invention refers to a supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters and can provide functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.
[0065] The "computing power resources" described in the present invention refer to technologies and facilities with information computing, transmission, storage, and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.
[0066] The "large language model" described in the present invention refers to a large language model (LLM), which is a language model with a relatively large number of parameters, aiming to understand and generate human language, trained through a large amount of text data, and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.
[0067] The "computing power operation task" described in the present invention refers to a specific workload or job that is executed on computing power resources and requires a certain amount of computing power support, usually involving scenarios such as complex data processing, numerical calculation, model training, or simulation.
[0068] In the present invention, by monitoring the first information corresponding to each node, when there is target information in the first information, the target node corresponding to the target information can be determined in a timely manner. The target node is the node with an abnormal task running state. Then, according to the target node, the target container set (Pod) with an abnormal running state is determined, and the target computing power operation task corresponding to the target Pod is restarted, so that the intelligent computing center can resume the normal operation of the faulty task. The above steps can be called the fault tolerance mechanism of the intelligent computing center. The fault tolerance mechanism can improve the stability of the operation of the intelligent computing center; it can also reduce the waste of computing power resources caused by the intelligent computing center continuing to execute tasks in the case of task errors.
[0069] Figure 1 It is a schematic flow chart of a method for fault tolerance of computing power operation tasks in an intelligent computing center provided by an embodiment of the present invention, as Figure 1 shown, the method includes the following steps:
[0070] Step S1: Monitor and obtain the first information corresponding to each node, where the first information is used to indicate whether the task running status of the corresponding node is abnormal.
[0071] In this step, after a certain cluster in the intelligent computing center is started, the status monitoring service is triggered, which can monitor the status of all nodes in the cluster, so as to obtain the first information indicating the task running status of the nodes.
[0072] A node is a working machine in the cluster (which can be a physical machine, a virtual machine or a cloud instance), and can be used to receive Pod scheduling instructions, provide computing, storage, and network resources for the Pod, monitor the status of the Pod, etc. A Pod is the smallest schedulable unit in the cluster, and a Pod can be created or deleted on a node at any time, or migrated between different nodes.
[0073] The task running status of the node indicated by the first information can include two types: the task running status is normal and the task running status is abnormal.
[0074] Step S2: When there is target information in the first information corresponding to each of the all nodes, determine the target node corresponding to the target information, where the target information indicates that the task running of the target node is abnormal.
[0075] The status indicated by the target information is abnormal. When there is target information in any first information, it means that there is a target node with an abnormal task running status in the cluster. Before the target node is restored to normal, if other nodes continue to execute tasks, it may cause waste of computing power resources. Therefore, it is necessary to process the target node in time.
[0076] Step S3: Based on the target node, determine the target Pod with an abnormal task running status.
[0077] In this step, determining the target Pod with an abnormal task running status from the Pods corresponding to the target node avoids the method of determining the target Pod with an abnormal running status from all nodes in the cluster, which can reduce the target quantity to be judged, thereby improving the efficiency.
[0078] Step S4: Restart the target computing power running task corresponding to the target Pod.
[0079] In this step, after determining the target Pod, restart the target computing power running task corresponding to the target Pod, so as to realize the process of the intelligent computing center restoring the normal operation of the error task.
[0080] In addition, the task type of the target computing power running task can be:
[0081] Artificial intelligence and machine learning, including tasks that require complex computations such as model training (e.g., training of large language models, training of computer vision models, etc.), model inference (e.g., speech recognition, graphic recognition, real-time decision-making for autonomous driving, etc.).
[0082] In the method for fault tolerance of computing power operation tasks in the intelligent computing center provided by the embodiments of the present invention, by monitoring the first information corresponding to each node, when there is target information in the first information, the target node corresponding to the target information can be determined in a timely manner. The target node is the node with an abnormal task running state. Then, according to the target node, the target Pod with an abnormal running state is determined, and the target computing power operation task corresponding to the target Pod is restarted. The above steps form a fault tolerance mechanism for the intelligent computing center, enabling the intelligent computing center to resume the normal operation of the faulty task.
[0083] In one embodiment, step S4 includes:
[0084] Step S41: Delete the target Pod located on the target node;
[0085] Step S42: Create the target Pod on any node other than the target node;
[0086] Step S43: Instruct the target Pod to execute the target computing power operation task.
[0087] Combined with Figure 2 and Figure 3 to illustrate this embodiment, as Figure 2 shown, the intelligent computing center includes Node 1 and Node 2. Node 1 includes Pod1, Pod2, and Pod3, and Node 2 includes Pod4. Node 1 is the target node, and the running state of Pod3 in Node 1 is abnormal. To resume the normal operation of the faulty task, step S41 is executed to delete Pod3 in Node 1. This step can prevent the further spread of abnormal problems and make Node 1 return to a normal node; and it can release the resources occupied by Pod3 for other Pods to use, thereby improving the utilization rate of computing power resources.
[0088] As Figure 3 shown, step S42 is executed to recreate Pod3 on Node 2, which can reset the state of Pod3; and Pod3 starts in the clean environment of Node 2, which can solve the abnormality caused by local buffering or damaged temporary files of Pod3.
[0089] After that, step S43 is executed to instruct Pod3 to execute the corresponding target computing power operation task, so that the computing power operation task is completely restored to normal.
[0090] In one embodiment, step 43 includes:
[0091] Step S431: Obtain the task running status data corresponding to the target Pod, where the task running status data is used to indicate the execution progress of the target computing power running task at the target moment, and the target moment is the nearest data backup moment to the moment when the target Pod runs abnormally;
[0092] Step S432: Based on the task running status data, instruct the target Pod to execute the target computing power running task.
[0093] In this embodiment, during the operation of the intelligent computing center, data backup is performed every preset time interval. The backup method can be incremental backup or full backup; incremental backup only backs up the data that has changed since the last backup, and full backup backs up all data each time.
[0094] Instructing the target Pod to execute the task according to the task running status data can restore the target computing power running task to the state closest to before the target Pod went abnormal, reducing data loss.
[0095] In one embodiment, the step S3 includes:
[0096] Step S31: Obtain the task label corresponding to the abnormal task of the target node;
[0097] Step S32: Based on the task label, determine N Pods corresponding to the task executed by the target node, where N is a positive integer;
[0098] Step S33: Obtain the respective second information corresponding to the N Pods, where the second information is used to indicate the running status of the corresponding Pod;
[0099] Step S34: Traverse the second information corresponding to the N Pods respectively, and determine the target Pod with an abnormal running status.
[0100] In this embodiment, the target node is a node with an abnormal task running status. Based on the task label, the Pods corresponding to the abnormal task of the target node, that is, N Pods, can be quickly determined without obtaining the second information of the Pods corresponding to the normal tasks of the target node, further reducing the amount of data to be processed, thereby further saving resources and improving efficiency.
[0101] In one embodiment, the first information includes at least one of the following:
[0102] The first status information, which is used to indicate that the corresponding node is in a normal state;
[0103] Second status information, used to indicate that the corresponding node is in a down state, a network-unreachable state, or a user-paused state;
[0104] Third status information, used to indicate that the corresponding node is in a state of abnormal task exit.
[0105] Under the fault tolerance mechanism provided by the embodiments of the present invention, although the task runs with errors, it can be resolved in a timely manner and will not affect the normal operation of the intelligent computing center. However, in this embodiment, since the first information corresponding to different abnormal states of the node is different, it is convenient for relevant personnel to determine the source of the abnormal operation, thereby reducing the probability of similar abnormal situations occurring subsequently.
[0106] In one embodiment, after the step S4, the method further includes:
[0107] Step S5: Generate a prompt message, where the prompt message is used to prompt that the computing power operation task corresponding to the target Pod has been restarted.
[0108] In this embodiment, a prompt message is generated after the target computing power operation task is restarted, which can remind relevant personnel to promptly review whether the target computing power operation task is completely restored after being released. If there are still problems, remedial measures can be taken in a timely manner, further enhancing the fault tolerance rate of the fault tolerance mechanism.
[0109] It should also be noted that after step S4, the execution steps of the fault tolerance mechanism return to step S1 until the target information appears again, and then the subsequent steps S2 to S4 are executed again, continuously repeating the above process to ensure the stability of the operation of the intelligent computing center.
[0110] Please refer to Figure 4 , embodiments of the present invention further provide a fault tolerance device 400 for the computing power operation task of an intelligent computing center. The device 400 includes:
[0111] An acquisition module 401, configured to monitor and acquire the first information corresponding to each node, where the first information is used to indicate whether the task running state of the corresponding node is abnormal;
[0112] A first determination module 402, configured to determine a target node corresponding to the target information in the case that the target information exists in the first information corresponding to each of the all nodes, where the target information indicates that the task running of the target node is abnormal;
[0113] A second determination module 403, configured to determine a target Pod with an abnormal task running state based on the target node;
[0114] A restart module 404, configured to restart the target computing power operation task corresponding to the target Pod.
[0115] In one embodiment, the restart module 404 includes:
[0116] A deletion sub-module, configured to delete the target Pod located at the target node;
[0117] A creation sub-module, configured to create the target Pod on any node other than the target node;
[0118] An indication sub-module, configured to instruct the target Pod to execute the target computing power operation task.
[0119] In one embodiment, the indication sub-module is further configured to:
[0120] Obtain task running status data corresponding to the target Pod, where the task running status data is used to indicate the execution progress of the target computing power operation task at a target moment, and the target moment is the nearest data backup moment from the moment when the target Pod runs abnormally;
[0121] Based on the task running status data, instruct the target Pod to execute the target computing power operation task.
[0122] In one embodiment, the second determination module 403 includes:
[0123] A first acquisition sub-module, configured to acquire a task label corresponding to an abnormal task of the target node;
[0124] A first determination sub-module, configured to determine N Pods corresponding to the task executed by the target node based on the task label, where N is a positive integer;
[0125] A second acquisition sub-module, configured to acquire second information corresponding to each of the N Pods, where the second information is used to indicate the running status of the corresponding Pod;
[0126] A second determination sub-module, configured to traverse the second information corresponding to each of the N Pods and determine a target Pod with an abnormal running status.
[0127] In one embodiment, the first information includes at least one of the following:
[0128] First status information, used to indicate that the corresponding node is in a normal state;
[0129] Second status information, used to indicate that the corresponding node is in a down state or a network disconnection state or a user suspension state;
[0130] Third status information, used to indicate that the corresponding node is in a task abnormal exit state.
[0131] In one embodiment, the apparatus 400 further includes:
[0132] An information generation module, configured to generate a prompt message, where the prompt message is used to prompt that the computing power operation task corresponding to the target Pod has completed restart.
[0133] The computing power operation task fault tolerance apparatus of the intelligent computing center provided by the present invention can implement each process of the above-mentioned intelligent computing center's computing power operation task fault tolerance method, with the technical features corresponding one by one and achieving the same technical effects. To avoid repetition, it will not be elaborated here.
[0134] Please refer to Figure 5 , the present invention also provides an electronic device 50, including a processor 51, a memory 52, and a computer program stored on the memory 52 and executable on the processor 51. When the computer program is executed by the processor 51, it implements each process of the above-mentioned intelligent computing center's computing power operation task fault tolerance method embodiment and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0135] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above-mentioned intelligent computing center's computing power operation task fault tolerance method embodiment and can achieve the same technical effects. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0136] The embodiment of the present application also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement each process of the above-mentioned Figure 1 shown intelligent computing center's computing power operation task fault tolerance method embodiment and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0137] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including that element.
[0138] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, electronic device, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0139] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the claims of the present invention, and all of them fall within the protection scope of the present invention.
Claims
1. A fault tolerance method for computing power operation tasks in an intelligent computing center, characterized in that, Including: Step S1: Monitor and obtain the first information corresponding to each node, where the first information is used to indicate whether the task running status of the corresponding node is abnormal; Step S2: When there is target information among the first information corresponding to each of the all nodes, determine the target node corresponding to the target information, where the target information indicates that the task running of the target node is abnormal; Step S3: Based on the target node, determine the target Pod with an abnormal task running status; Step S4: Restart the target computing power running task corresponding to the target Pod.
2. The method according to claim 1, wherein The step S4 includes: Step S41: Delete the target Pod located on the target node; Step S42: Create the target Pod on any other node except the target node; Step S43: Instruct the target Pod to execute the target computing power running task.
3. The method according to claim 2, wherein The step 43 includes: Step S431: Obtain the task running status data corresponding to the target Pod, where the task running status data is used to indicate the execution progress of the target computing power running task at the target moment, and the target moment is the nearest data backup moment from the moment when the target Pod runs abnormally; Step S432: Based on the task running status data, instruct the target Pod to execute the target computing power running task.
4. The method according to claim 1, wherein The step S3 includes: Step S31: Obtain the task label corresponding to the abnormal task of the target node; Step S32: Based on the task label, determine N Pods corresponding to the task executed by the target node, where N is a positive integer; Step S33: Obtain the second information corresponding to each of the N Pods, where the second information is used to indicate the running status of the corresponding Pod; Step S34: Traverse the second information corresponding to each of the N Pods and determine the target Pod with an abnormal running status.
5. The method according to any one of claims 1 to 4, characterized in that The first information includes at least one of the following: The first status information, which is used to indicate that the corresponding node is in a normal state; The second status information, which is used to indicate that the corresponding node is in a down state or a network connection failure state or a user suspension state; The third status information, which is used to indicate that the corresponding node is in a task abnormal exit state.
6. The method according to any one of claims 1 to 4, characterized in that, After the step S4, the method further includes: Step S5: Generate a prompt message, where the prompt message is used to prompt that the computing power running task corresponding to the target Pod has been restarted.
7. A fault tolerance device for computing power operation tasks in an intelligent computing center, characterized in that, Including: An acquisition module, which is used to monitor and obtain the first information corresponding to each node, where the first information is used to indicate whether the task running status of the corresponding node is abnormal; A first determination module, which is used to determine the target node corresponding to the target information when there is target information among the first information corresponding to each of the all nodes, where the target information indicates that the task running of the target node is abnormal; A second determination module, which is used to determine the target Pod with an abnormal task running status based on the target node; A restart module, which is used to restart the target computing power running task corresponding to the target Pod.
8. An electronic device, characterized in that, Including: A processor, a memory, and a program stored on the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the fault tolerance method for the computing power operation task of the intelligent computing center according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the fault tolerance method for the computing power operation task of the intelligent computing center according to any one of claims 1 to 6 are implemented.
10. A computer program product, characterized in that, It includes computer instructions, and when the computer instructions are executed by a processor, the steps of the fault tolerance method for the computing power operation task of the intelligent computing center according to any one of claims 1 to 6 are implemented.