Elastic fault tolerance method and device based on sidecar container

By introducing sidecar containers to monitor the status of the service container in the distributed training system, the problem of tight coupling of elastic fault tolerance mechanism and business logic in the existing technology is solved, efficient and flexible fault tolerance processing is achieved, and the success rate of training tasks and system stability are improved.

CN120336071AInactive Publication Date: 2025-07-18ZHEJIANG LAB
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510813204.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the existing distributed training technology, the elastic fault tolerance mechanism is tightly coupled with business logic, which increases development complexity and is difficult to maintain, resulting in training interruptions and inefficiency.

Method used

The sidecar container is introduced in the distributed system. By injecting auxiliary containers into the container group to monitor the status of the service container and reporting it to the main node in real time, it can achieve rapid response and processing of exceptions and avoid intrusion into the distributed system code.

Benefits of technology

It realizes flexible and flexible fault tolerance functions, reduces the learning and deployment costs of technicians, improves the success rate and stability of training tasks, and enhances the scalability and flexibility of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336071A_ABST
    Figure CN120336071A_ABST
Patent Text Reader

Abstract

The invention discloses an elastic fault tolerance method and device based on a sidecar container. When the elastic fault-tolerant method provided by the invention is adopted to realize an elastic fault-tolerant function in distributed training, the operation state of a service container can be monitored by additionally adding an auxiliary container in a container group for executing a training task, and the operation state is reported to a main node in real time; therefore, the purpose that the main node discovers and processes the exception of the service container at the first time is achieved. Meanwhile, according to the method, codes of the distributed system do not need to be intruded, decoupling with a distributed system framework is achieved, and the learning and deployment cost of technicians when the elastic fault-tolerant function is used is reduced reply.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of distributed training technologies, and particularly to an elastic fault tolerance method and device based on sidecar containers. Background Art

[0002] Distributed training is a technology for training neural network models using multiple nodes in a distributed system, aiming to solve the computational resource limitation problems during large-scale data and model training. However, in a distributed environment, there are problems such as training interruption and low efficiency caused by node failures or dynamic resource changes. To solve this problem, existing technologies have introduced elastic fault tolerance technologies in distributed training.

[0003] Existing elastic fault tolerance technologies usually integrate complex health checks, load balancing and other mechanisms based on the training framework itself. This not only increases the complexity for developers to develop the framework, but also may be difficult to maintain and expand due to tight coupling with business logic.

[0004] Therefore, how to achieve flexible elastic fault tolerance without increasing the complexity of the distributed system is an urgent problem to be solved. Summary of the Invention

[0005] This specification provides an elastic fault tolerance method and device based on sidecar containers to at least partially solve the above problems existing in the prior art.

[0006] This specification adopts the following technical solutions: This specification provides an elastic fault tolerance method, which is applied to the master node of a distributed system. The distributed system includes at least a master node and several slave nodes. The method includes: Receiving a configuration file input by a user, and embedding an injection instruction into the configuration file. The injection instruction is used to inject an auxiliary container into a container group; Creating a container group including a business container and an auxiliary container according to the configuration file. The auxiliary container is used to monitor the running status of the business container; Deploying each container group to a slave node, so that each slave node executes a training task in the business container; During the execution of the training task, continuously receiving the running status sent by the auxiliary containers included in each container group; In response to any business container being abnormal, performing a response process on the node where the abnormal business container is located.

[0007] Optionally, the running status at least includes the central processing unit utilization rate, memory usage, and network interface parameters.

[0008] Optionally, deploy each container group to a slave node, so that each slave node executes a training task in the service container. Specifically, it includes: For each container group, deploy the container group to an idle slave node, so that the auxiliary containers included in the container group detect the running environment of the idle slave node, and when the running environment of the idle slave node is normal, the service containers of the container group execute the training task.

[0009] Optionally, the running environment detection includes at least one of resource availability verification and network connection test.

[0010] Optionally, in response to any service container being abnormal, perform a response process on the node where the abnormal service container is located. Specifically, it includes: In response to any service container being abnormal, determine the type of the abnormality that the service container has. If the type of the abnormality is a specified abnormal type, re-select an idle slave node to replace the node where the service container is located.

[0011] Optionally, re-selecting an idle slave node to replace the node where the service container is located specifically includes: Evict the container group from the node where the abnormal service container is located. Create a replacement container group including a service container and an auxiliary container according to the evicted container group, and deploy the replacement container group to the re-selected idle slave node to continue executing the training task.

[0012] Optionally, the method further includes: After the training task is executed, use the auxiliary container to execute a cleaning process on all containers included in the container group where it is located.

[0013] An elastic fault tolerance device provided in this specification, the device is applied to the master node of a distributed system, the distributed system at least includes a master node and several slave nodes, and the device includes: An injection module, configured to receive a configuration file input by a user, and embed an injection instruction into the configuration file, where the injection instruction is used to inject an auxiliary container into a container group. A creation module, configured to create a container group including a service container and an auxiliary container according to the configuration file, where the auxiliary container is used to monitor the running state of the service container. A deployment module, configured to deploy each container group to a slave node, so that each slave node executes a training task in the service container. A receiving module, configured to continuously receive the running states sent by the auxiliary containers included in each container group during the execution of the training task. A processing module, configured to perform a response process on a node where a service container with an exception appears in response to an exception occurring in any service container.

[0014] This specification provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above elastic fault tolerance method is implemented.

[0015] This specification provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above elastic fault tolerance method is implemented.

[0016] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects: In the elastic fault tolerance method provided in this specification, a configuration file input by a user is received, and an injection instruction is embedded in the configuration file. The injection instruction is used to inject an auxiliary container into a container group; a container group including a service container and an auxiliary container is created according to the configuration file, and the auxiliary container is used to monitor the running state of the service container; each container group is deployed to a slave node, so that each slave node executes a training task in the service container; during the execution of the training task, the running states sent by the auxiliary containers included in each container group are continuously received; in response to an exception occurring in any service container, a response process is performed on the node where the service container with the exception appears.

[0017] When implementing the elastic fault tolerance function in distributed training by using the elastic fault tolerance method provided in this specification, the running state of the service container can be monitored by adding an additional auxiliary container to the container group executing the training task, and reported to the master node in real time, so as to achieve the purpose that the master node discovers and processes the exception of the service container in the first time. At the same time, this method does not need to invade the code of the distributed system, is decoupled from the distributed system framework, and greatly reduces the learning and deployment costs of technicians when using the elastic fault tolerance function. Description of the Drawings

[0018] The drawings described herein are used to provide a further understanding of this specification, and constitute a part of this specification. The illustrative embodiments and descriptions of this specification are used to explain this specification, and do not constitute an improper limitation to this specification. In the drawings: Figure 1 It is a schematic flowchart of an elastic fault tolerance method in this specification; Figure 2 It is a schematic diagram of the relationship between an Inception Controller and an Inception Agent provided in this specification; Figure 3Schematic diagram of an elastic fault tolerance device provided in this specification; Figure 4 Corresponding to what is provided in this specification Figure 1 Schematic diagram of an electronic device. Specific implementation manners

[0019] To make the purpose, technical solutions and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0020] The following will detail the technical solutions provided by each embodiment of this specification in conjunction with the drawings.

[0021] Figure 1 Flow schematic diagram of an elastic fault tolerance method in this specification. The method is applied to the master node of a distributed system. The distributed system includes at least a master node and several slave nodes, and specifically includes the following steps: S100: Receive a configuration file input by a user, and embed an injection instruction into the configuration file. The injection instruction is used to inject an auxiliary container into a container group.

[0022] All steps in the elastic fault tolerance method provided in this specification can be implemented by any electronic device with computing functions, such as devices like terminals and servers.

[0023] Elastic fault tolerance means that the system can dynamically adjust resource allocation and task scheduling in the face of hardware or software failures to ensure the continuity and reliability of services. Elastic fault tolerance combines elastic expansion and a fault tolerance mechanism, and can quickly recover and maintain the normal operation of the system when a failure occurs.

[0024] This method is mainly used to provide a flexible and efficient fault tolerance method for a distributed training system that does not require intrusion into the system code and is decoupled from the system. Therefore, when implementing this method, it is first necessary to build a distributed system capable of performing distributed training on large models.

[0025] The construction technology of distributed systems has been relatively mature, and this specification will not elaborate on it. To make the introduction of this method clearer, this specification will take the distributed system built based on the Kubernetes (K8S) cluster as an example for description.

[0026] Before starting to execute this method, it is necessary to ensure that the Kubernetes cluster has been correctly configured and all necessary components (such as kube-scheduler, kubelet, etc.) are running properly. Configure the network policy and service discovery mechanism so that each component can communicate with each other.

[0027] Subsequently, it is necessary to deploy the fault-tolerant controller (InceptionController) component in the Master layer of the cluster, which is responsible for sidecar container injection, command reception, task management, task monitoring, and fault diagnosis functions. At the same time, configure the integration of the Inception Controller with the cluster management system to ensure that it can monitor the cluster status and receive task scheduling instructions; and configure the connection between the Inception Controller and the Inception Controller to ensure that after the auxiliary container (Inception Agent) is generated in the Worker layer, the Inception Controller can receive task instructions and report node and task running status information.

[0028] Figure 2 This is a schematic diagram of the relationship between the Inception Controller and the Inception Agent provided in this specification. As Figure 2 shown, the Inception Controller and the Inception Agent proposed in this method are respectively deployed on the master node and the slave node, and the two are combined with each other to achieve the elastic fault-tolerant function.

[0029] The Inception Controller is a controller deployed in the Master layer designed based on the Sidecar architecture in this method, which is responsible for the coordination and management of the elastic fault-tolerant process of the entire system. It contains multiple functional modules for realizing comprehensive monitoring, resource management, and fault recovery of tasks. The functions of the Inception Controller can include fault command acceptance, task management, fault diagnosis, Sidecar injection, task monitoring, and automatic scaling. As the core controller in the Master layer, the Inception Controller is responsible for global elastic fault-tolerant management.

[0030] The Inception Agent is an agent program or module designed in this method for managing and executing tasks related to Inception elastic fault tolerance. The Inception Agent is loaded into the user's training task container group (pod) through the Sidecar injection function of the Inception Controller, monitors the running status of the business container in the form of a sidecar, and interacts with the Inception Controller in real time to achieve comprehensive support for system elastic fault tolerance. Its main functions include signal interaction, business container monitoring, fault reporting, fault tolerance execution support, etc.

[0031] The main application scenario of this method is the scenario where exceptions occur during distributed training. Based on this, the configuration file input by the user can be received in this step first to create a training task. For example, the user can initiate a PyTorch Job request by submitting a YAML file to the distributed system to start a training task.

[0032] Once the configuration file submitted by the user is received, the fault tolerance controller Inception Controller in the master node will automatically embed injection instructions into the configuration file. The role of the injection instructions is that when creating a pod for executing the training task subsequently, an additional auxiliary container Inception Agent will be created in each pod.

[0033] S102: Create a container group containing a business container and an auxiliary container according to the configuration file, where the auxiliary container is used to monitor the running status of the business container.

[0034] After receiving the configuration file and embedding the injection instructions in step S100, the creation of the container group pod for executing the training task can start in this step according to the configuration file.

[0035] The Inception Controller in the master node can schedule the Training Operator in the distributed system to create a pod according to the configuration file. The Training Operator is a custom controller included in Kubernetes, specifically used to manage and run distributed machine learning training tasks on the Kubernetes cluster, such as TensorFlow, PyTorch, MXNet, etc.

[0036] The Training Operator can start the corresponding Pods according to the updated configuration file. Kube-scheduler is responsible for allocating these Pods to appropriate nodes in the cluster, while Kubelet is responsible for actually creating and running these Pods. Among them, Kube-scheduler is one of the core components of the Kubernetes control plane, responsible for allocating newly created Pods to run on appropriate nodes in the cluster; Kubelet is the core component running on each worker node in the Kubernetes cluster, responsible for managing the Pod lifecycle and container runtime on the node.

[0037] In traditional distributed training technologies, when creating a container group for executing a training task, the container group only contains several business containers for computing; while in this method, auxiliary containers, namely Inception Agent, are additionally created in each container group through injection instructions.

[0038] In this method, the main role of the auxiliary container is to monitor the running status of business containers under the same container group, and the running status of business containers can include but is not limited to CPU utilization, memory usage, network interface parameters, etc.

[0039] Distributed training usually requires multiple nodes in the system to jointly execute the training task. Therefore, multiple container groups are usually created when creating the container group in this step, and each container group contains several business containers and one auxiliary container.

[0040] S104: Deploy each container group to the worker nodes so that each worker node executes the training task in the business container.

[0041] In this step, several container groups created in step S102 can be deployed to the worker nodes to start executing the training task.

[0042] Additionally, before the business container officially runs, Inception Agent can first perform a series of predefined pre-start checks. Specifically, for each container group, deploy the container group to an idle worker node, so that the auxiliary container included in the container group detects the running environment of the idle worker node, and when the running environment of the idle worker node is normal, make the business container of the container group execute the training task. Among them, the running environment monitoring can include resource availability verification, network connection testing, etc.

[0043] In this method, an idle slave node can be a slave node in a distributed system that is not currently executing other tasks and has sufficient available resources. Different container groups are deployed to different slave nodes, and only one container group is deployed on each slave node. After the deployment of each container group is completed, the business containers in each container group will start to execute the training task. At the same time, as an auxiliary container, Inception Agent will continuously monitor the running status of each business container in the same container group.

[0044] S106: During the execution of the training task, continuously receive the running status sent by the auxiliary containers included in each container group.

[0045] During the execution of the training task by each business container, Inception Agent will continuously monitor its health status and performance metrics, including but not limited to key parameters such as CPU utilization, memory usage, network I / O, etc. This continuous monitoring ensures that any abnormal behavior or potential problems can be detected in a timely manner. At the same time, in order to ensure that Inception Controller can understand the status of each business container in real time, Inception Agent will send heartbeat signals to InceptionController at a predetermined time interval.

[0046] Correspondingly, Inception Controller will continuously receive the information reported by Inception Agent, and after receiving the running status reported by each Inception Agent, process and analyze this information. InceptionController will summarize and statistically analyze the collected data according to different training tasks, so as to generate a comprehensive view of the entire training task. This approach helps to quickly identify any behavior that deviates from normal operation and take appropriate measures for intervention, such as adjusting resource allocation or triggering a fault recovery mechanism.

[0047] S108: In response to an exception occurring in any business container, perform response processing on the node where the business container with the exception is located.

[0048] Based on the continuous reception and analysis of the running status information of each business container in step S106, Inception Controller in the master node can detect and respond immediately when an exception occurs in a business container. Specifically, in response to an exception occurring in any business container, determine the type of exception that has occurred in this business container; if the type of the exception is a specified exception type, re-select an idle slave node to replace the node where this business container is located.

[0049] When an exception occurs in the business container of a certain node, the Inception Controller in the master node will receive a notice about the node failure immediately. This notice may come from the report of the auxiliary container included in the container group of the node, or it may be from the abnormal behavior detected by the observable monitoring system that may be carried in the distributed system. At this time, the Inception Controller can inform the faulty node to conduct a fault check.

[0050] During the execution of the training task, there are various types of exceptions that the business container may encounter. For different types of exceptions, the Inception Controller can adopt different ways to handle them. Generally, the types of exceptions that the business container encounters can be divided into specified exception types and unspecified exception types. Among them, the specified exception type means an exception that requires node replacement, such as hardware failure, network interruption, business container crash, etc.; the unspecified exception type means an exception that does not require node replacement.

[0051] When the type of exception that occurs in the business container is a specified exception type, the Inception Controller will handle it by replacing the node. Specifically, the container group can be evicted from the node where the abnormal business container is located; a replacement container group containing the business container and the auxiliary container is created according to the evicted container group, and the replacement container group is deployed to the newly selected idle slave node to continue executing the training task.

[0052] When an exception of the specified exception type is detected on a certain node, in order to prevent the spread of the fault and ensure the continuity of the service, the Inception Controller can further instruct the distributed system to evict all running pods from the faulty node to avoid the risk of service interruption. As the pods on the faulty node are evicted, the Training Operator will automatically create new Pods to replace those affected tasks. The newly created pods will be assigned to healthy idle nodes by the Kube-scheduler, and then the Kubelet is responsible for the actual startup and running work.

[0053] In each newly created pod, there will also be a business container and an auxiliary container. Whenever a new pod is started, it will be accompanied by an Inception Agent, which is used to perform necessary initialization checks and continuous monitoring. Before the business container officially runs, the Inception Agent will first conduct a comprehensive self-check of the node to ensure the health and stability of the environment. Once the self-check confirms that the node has returned to normal and is safe to use, the business container will start to execute its scheduled tasks. At the same time, the Inception Agent will continue to run in the background, periodically sending heartbeat signals to the Inception Controller to ensure the effective operation of the real-time monitoring and fault prevention mechanisms.

[0054] Furthermore, after the training task is completed, the auxiliary containers in each existing container group used to execute the training task will clean up all the containers in the container group. Specifically, after the training task is executed, the auxiliary container can be used to execute the cleaning process for all the containers included in its own container group.

[0055] When the business container successfully completes all the scheduled training tasks and exits in a normal state (for example, the exit code is 0), the Inception Agent will detect this event, which means that the training task has been successfully completed as expected without any abnormal interruption. After confirming the normal exit of the business container, the Inception Agent will record the corresponding information in the system log, including but not limited to the completion time of the training task, the final status, and any relevant performance metrics or result summaries. After completing all the above operations, the Inception Agent will execute the cleaning process for itself and the business container under the same container group and exit normally.

[0056] When implementing the elastic fault tolerance function in distributed training using the elastic fault tolerance method provided in this specification, the running state of the business container can be monitored by adding an additional auxiliary container to the container group executing the training task and reported to the master node in real time, so as to achieve the purpose of the master node discovering and handling the anomalies of the business container in a timely manner. At the same time, this method does not need to invade the code of the distributed system, is decoupled from the distributed system framework, and significantly reduces the learning and deployment costs of technicians when using the elastic fault tolerance function.

[0057] This method conducts elastic fault tolerance training decoupled from the framework by introducing the Sidecar mode, significantly improving the coupling between fault tolerance and the framework and the system flexibility during the model training process. The main beneficial effects are reflected in the following aspects: First, the decoupling of the service and the framework is achieved, enabling developers to focus on the implementation of the core business logic without having to concern themselves with non-business-related logics such as complex fault tolerance, elastic scaling, networking, and health monitoring. This design not only reduces the development cost but also enhances the scalability and flexibility of the system.

[0058] Second, by deploying a lightweight auxiliary container in each container group, the present invention provides the ability to monitor the training task container group in real time. The auxiliary container runs independently of the main application and is specifically responsible for collecting and analyzing key performance indicators such as heartbeat, CPU utilization, memory usage, and network I / O, ensuring that any abnormal behavior can be accurately captured even in high-concurrency training scenarios, thereby guaranteeing the efficiency and stability of the training process.

[0059] Furthermore, when a faulty instance is detected, the auxiliary container can immediately trigger the isolation mechanism, effectively preventing the spread of the fault from affecting other normally working pods, ensuring the stability and reliability of the system. In addition, the detailed monitoring data supports more accurate problem diagnosis and decision-making, providing strong data support for optimizing resource allocation strategies and predicting potential risk points.

[0060] Finally, since the Sidecar mode allows services with different architectures and technology stacks to share the same fault tolerance strategy, this promotes interoperability and consistency in the complex environment of large model training, laying a solid foundation for building an efficient and scalable distributed training system. In summary, the present invention provides a simple, efficient, and easy-to-maintain fault tolerance solution, solving the problems of tight coupling and difficult maintenance in the prior art, and significantly improving the success rate and stability of training tasks.

[0061] The above is the elastic fault tolerance method provided in this specification. Based on the same idea, this specification also provides a corresponding elastic fault tolerance device, such as Figure 3 shown.

[0062] Figure 3 FIG. is a schematic diagram of an elastic fault tolerance device provided in this specification. The device is applied to the master node of a distributed system, and the distributed system includes at least a master node and several slave nodes. Specifically, it includes: An injection module 200, configured to receive a configuration file input by a user and embed an injection instruction into the configuration file, where the injection instruction is used to inject an auxiliary container into a container group; A creation module 202, configured to create a container group including a business container and an auxiliary container according to the configuration file, where the auxiliary container is used to monitor the running state of the business container; A deployment module 204, configured to deploy each container group to a slave node, enabling each slave node to execute a training task in the business container; A receiving module 206, configured to continuously receive the running status sent by the auxiliary containers included in each container group during the execution of the training task. A processing module 208, configured to perform a response process on the node where the abnormal service container is located in response to an abnormality occurring in any service container.

[0063] Optionally, the running status at least includes the CPU utilization rate, the memory usage, and the network interface parameters.

[0064] Optionally, the deployment module 204 is specifically configured to, for each container group, deploy the container group to an idle slave node, enable the auxiliary containers included in the container group to detect the running environment of the idle slave node, and enable the service containers of the container group to execute the training task when the running environment of the idle slave node is normal.

[0065] Optionally, the running environment detection includes at least one of resource availability verification and network connection test.

[0066] Optionally, the processing module 208 is specifically configured to, in response to an abnormality occurring in any service container, determine the type of the abnormality that has occurred in the service container; if the type of the abnormality is a specified type of abnormality, re-select an idle slave node to replace the node where the service container is located.

[0067] Optionally, the processing module 208 is specifically configured to evict the container group from the node where the abnormal service container is located; create a replacement container group including service containers and auxiliary containers according to the evicted container group, and deploy the replacement container group to the re-selected idle slave node to continue executing the training task.

[0068] Optionally, the device further includes a cleaning module 210, specifically configured to, after the execution of the training task is completed, use the auxiliary containers to execute a cleaning process on all the containers included in the container group where the auxiliary containers are located.

[0069] This specification also provides a computer-readable storage medium storing a computer program, and the computer program can be used to execute the Figure 1 elastic fault tolerance method provided above.

[0070] This specification also provides Figure 4 a schematic structural diagram of the electronic device shown. As Figure 4 described above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and of course may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1The elastic fault tolerance method described above. Of course, in addition to the software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or logic devices.

[0071] The improvement of a technology can be clearly distinguished as either a hardware improvement (e.g., improvement of circuit structures such as diodes, transistors, switches, etc.) or a software improvement (improvement of method flows). However, with the development of technology, many improvements of method flows today can be regarded as direct improvements of hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement of a method flow cannot be implemented by a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. The designer can program by himself / herself to "integrate" a digital system on a piece of PLD, without the need to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL). And there is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that as long as the method flow is slightly logically programmed in the above-mentioned several hardware description languages and programmed into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0072] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that, in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same functions. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.

[0073] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0074] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0075] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0076] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the specified functions in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0077] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the specified functions in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0078] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0079] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0080] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0081] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0082] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0083] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, system or computer program product. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0084] This specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0085] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment.

[0086] The above description is only for the embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various modifications and changes can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this application.

Claims

1. An elastic fault tolerance method, characterized in that, The method is applied to the master node of a distributed system. The distributed system includes at least a master node and several slave nodes. The method includes: Receiving a configuration file input by a user, and embedding an injection instruction into the configuration file, where the injection instruction is used to inject an auxiliary container into a container group; Creating a container group including a service container and an auxiliary container according to the configuration file, where the auxiliary container is used to monitor the running status of the service container; Deploying each container group to a slave node, so that each slave node executes a training task in the service container; During the execution of the training task, continuously receiving the running status sent by the auxiliary containers included in each container group; In response to an exception occurring in any service container, performing a response process on the node where the service container with the exception is located.

2. The method according to claim 1, characterized in that, The running status at least includes the CPU utilization rate, the memory usage, and the network interface parameters.

3. The method according to claim 1, wherein Deploying each container group to a slave node, so that each slave node executes a training task in the service container, specifically including: For each container group, deploying the container group to an idle slave node, so that the auxiliary container included in the container group performs a running environment detection on the idle slave node, and when the running environment of the idle slave node is normal, enabling the service container included in the container group to execute the training task.

4. The method according to claim 3, wherein The running environment detection includes at least one of resource availability verification and network connection test.

5. The method according to claim 1, characterized in that, In response to an exception occurring in any service container, performing a response process on the node where the service container with the exception is located, specifically including: In response to an exception occurring in any service container, determining the type of the exception that has occurred in the service container; If the type of the exception is a specified exception type, reselecting an idle slave node to replace the node where the service container is located.

6. The method according to claim 5, wherein Reselecting an idle slave node to replace the node where the service container is located, specifically including: Evicting the container group from the node where the service container with the exception is located; Creating a replacement container group including a service container and an auxiliary container according to the evicted container group, and deploying the replacement container group to the reselected idle slave node to continue executing the training task.

7. The method according to claim 1, wherein The method further includes: After the training task is executed, using the auxiliary container to perform a cleaning process on all the containers included in the container group where the auxiliary container is located.

8. An elastic fault-tolerant device, characterized in that, The device is applied to the master node of a distributed system. The distributed system includes at least a master node and several slave nodes. The device includes: An injection module, configured to receive a configuration file input by a user, and embed an injection instruction into the configuration file, where the injection instruction is used to inject an auxiliary container into a container group; A creation module, configured to create a container group including a service container and an auxiliary container according to the configuration file, where the auxiliary container is used to monitor the running status of the service container; A deployment module, configured to deploy each container group to a slave node, so that each slave node executes a training task in the service container; A receiving module, configured to continuously receive the running status sent by the auxiliary containers included in each container group during the execution of the training task; A processing module, configured to perform a response process on the node where the service container with the exception is located in response to an exception occurring in any service container.

9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of the above claims 1 to 7 is implemented.

10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method described in any one of the above claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Container group instance termination method and device, container group instance creation method and device, electronic equipment and storage medium

    CN113923257A

  • MySQL high availability implementation method and system based on Kubernetes

    CN117827159A

  • Micro-service Sidecar injection method and system based on configuration file deployment

    CN119088430A

  • Method and apparatus with sidecar pattern checkpointing

    US20240095066A1