Assembly line work recovery method and device, electronic equipment and storage medium

By using the container scheduling adapter to destroy containers and trigger graceful shutdown during the Kubernetes rolling restart process, and storing pipeline information in real time and generating new containers, the pipeline interruption problem is solved, and precise pipeline continuation and automated recovery are achieved.

CN120909837APending Publication Date: 2025-11-07AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511016265.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

In existing technologies, the problem of pipeline job interruption caused by Kubernetes rolling restart has not been effectively solved.

Method used

The first container is destroyed by the container scheduling adapter, triggering a graceful shutdown mechanism to terminate the current pipeline and write the information to the data storage layer. The second container is then generated and the pipeline controller is started. Based on the data storage layer, the pipeline state is detected and restored to a pipeline that is currently executing.

Benefits of technology

It enables precise resume and fully automated recovery of interrupted operations in the production line, avoiding resource waste and duplicate execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909837A_ABST
    Figure CN120909837A_ABST
Patent Text Reader

Abstract

The invention discloses an assembly line work recovery method and device, electronic equipment and a storage medium. The method comprises the steps that when a container arrangement platform executes rolling restart, a first container is destroyed through a container scheduling adapter, and a signal processing engine triggers an elegant shutdown mechanism; closing a task receiving interface through a pipeline controller, terminating the currently executed pipeline, and writing information of the currently executed pipeline into a data storage layer; a second container is generated through the container scheduling adapter, an opening signal of the second container is sent to the signal processing engine through the container scheduling adapter, the signal processing engine starts the pipeline controller, and the pipeline state is detected to be the pipeline in execution in the data storage layer; and performing job recovery on the assembly line in the state of being executed based on the currently executed job step of the assembly line. According to the scheme, the problem of assembly line interruption caused by rolling restart is solved, and assembly line breakpoint accurate continuous transmission and full-automatic recovery are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, and in particular, to a pipeline job recovery method and device, electronic equipment and storage medium. BACKGROUND

[0002] In the software delivery process, pipeline job automation delivery is the core automation execution framework. With the popularity of containerization technology, such pipelines are usually run on cloud clusters such as Kubernetes, and are dynamically scheduled in the form of containers. However, when the software application version is upgraded, Kubernetes will perform a rolling restart, gradually replacing the old Pod, thereby causing the pipeline being executed in the Pod to be interrupted.

[0003] In the implementation of the present application, it is found that the prior art at least has the following technical problems: the prior art solution described above has the problem of pipeline interruption caused by rolling restart. SUMMARY

[0004] The present application provides a pipeline job recovery method and device, electronic equipment and storage medium to solve the problem of pipeline interruption caused by rolling restart.

[0005] According to an aspect of the present application, a pipeline job recovery method is provided, comprising:

[0006] In the case of performing a rolling restart on a container orchestration platform, a first container is destroyed through a container scheduling adapter;

[0007] A shutdown signal of the first container is sent to a signal processing engine through the container scheduling adapter, and the signal processing engine triggers a graceful shutdown mechanism;

[0008] A task receiving interface is closed through a pipeline controller, a currently executing pipeline is terminated, and information of the currently executing pipeline is written into a data storage layer, the information of the currently executing pipeline including a pipeline state and a job step currently executed by the pipeline;

[0009] A second container is generated through the container scheduling adapter, an opening signal of the second container is sent to the signal processing engine through the container scheduling adapter, and the signal processing engine starts the pipeline controller;

[0010] A pipeline state in execution is detected in the data storage layer through the pipeline controller, and a job recovery is performed on the pipeline state in execution based on a job step currently executed by the pipeline.

[0011] According to another aspect of the present application, a pipeline job recovery device is provided, comprising:

[0012] a container destroying module configured to destroy the first container via the container scheduling adapter in a case that a rolling restart is performed on a container orchestration platform;

[0013] a graceful shutdown module configured to send a shutdown signal of the first container to a signal processing engine via the container scheduling adapter, the signal processing engine triggering a graceful shutdown mechanism;

[0014] an executing pipeline information storing module configured to close a task receiving interface via the pipeline controller, terminate a currently executing pipeline, and write information of the currently executing pipeline into a data storage layer, the information of the currently executing pipeline including a pipeline state and a job step currently executed by the pipeline;

[0015] a container generating module configured to generate a second container via the container scheduling adapter, and send a start signal of the second container to the signal processing engine via the container scheduling adapter, the signal processing engine starting the pipeline controller;

[0016] a pipeline job resuming module configured to detect, via the pipeline controller, a pipeline in the data storage layer with a pipeline state of being executed, and resume a job of the pipeline in the pipeline state of being executed based on a job step currently executed by the pipeline.

[0017] According to another aspect of the present application, an electronic device is provided, the electronic device comprising:

[0018] at least one processor;

[0019] and a memory connected to the at least one processor in communication;

[0020] wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the pipeline job resuming method according to any one of the embodiments of the present application.

[0021] According to another aspect of the present application, a computer readable storage medium is provided, the computer readable storage medium storing computer instructions for enabling a processor to perform the pipeline job resuming method according to any one of the embodiments of the present application when executed by the processor.

[0022] The technical scheme of the embodiment of the application is that, in the case that the rolling restart is performed on the container orchestration platform, the first container is destroyed by the container scheduling adapter, the container scheduling adapter sends a shutdown signal of the first container to the signal processing engine, the signal processing engine triggers the graceful shutdown mechanism, the task receiving interface is closed by the pipeline controller, the currently executing pipeline is terminated, the information of the currently executing pipeline is written into the data storage layer, the second container is generated by the container scheduling adapter, the container scheduling adapter sends an opening signal of the second container to the signal processing engine, the signal processing engine starts the pipeline controller, and the pipeline controller detects the pipeline state as the pipeline being executed in the data storage layer, and the pipeline state as the pipeline being executed is recovered based on the job step currently executed by the pipeline. The above technical scheme stores the information of the currently executing pipeline in real time when the historical container is destroyed, detects the pipeline state as the pipeline being executed after the new container is generated, and then recovers the pipeline state as the pipeline being executed based on the job step currently executed by the pipeline, thereby solving the pipeline interruption problem caused by the rolling restart, and realizing the accurate breakpoint resuming and full-automatic recovery of the pipeline.

[0023] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the application, nor is it used to limit the scope of the application. Other features of the application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the application, the drawings needed in the embodiment description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0025] Figure 1 is a structural schematic diagram of a pipeline job recovery system provided by the application;

[0026] Figure 2 is a flowchart of a pipeline job recovery method according to the first embodiment of the application;

[0027] Figure 3 is a flowchart of a pipeline job recovery method according to the second embodiment of the application;

[0028] Figure 4 is a flowchart of a pipeline job recovery method according to the embodiment of the application;

[0029] Figure 5is a flow chart of a Redis lock according to an embodiment of the present application;

[0030] Figure 6 is a structural schematic diagram of a pipeline operation recovery device according to an embodiment three of the present application;

[0031] Figure 7 is a structural schematic diagram of an electronic device for implementing a pipeline operation recovery method according to an embodiment of the present application. DETAILED DESCRIPTION

[0032] In order to make the personnel in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the present application.

[0033] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices. The acquisition, storage, use, processing and the like of data in the technical solutions of the present application all comply with the relevant provisions of national laws and regulations.

[0034] Before introducing the specific embodiments, first introduce the application scenario of the pipeline operation recovery method, the pipeline operation recovery method can be applied to a pipeline operation recovery system, Figure 1 is a structural schematic diagram of a pipeline operation recovery system provided by the present application, such as Figure 1As shown, the pipeline operation recovery system includes a container scheduling adapter, a signal processing engine, a pipeline controller, and a data storage layer. The container scheduling adapter is responsible for managing the life cycle of containers (such as Pods, etc.), including destroying old Pods and generating new Pods during rolling restart. The signal processing engine is responsible for capturing the shutdown signal sent by the container scheduling adapter and managing the life cycle hook to perform cleanup operations such as releasing resources before the Java Virtual Machine (JVM) is closed. The pipeline controller is responsible for the creation, start-stop, and state monitoring of the pipeline. Depending on the content of the software delivery, it can be divided into program change pipeline, data change pipeline, or configuration change pipeline, etc. Each pipeline contains different job steps, for example, the data change pipeline includes data preparation job step, data execution job step, and data backup job step. The pipeline controller stores the information of the pipeline such as the current execution job step of the pipeline into the data storage layer. The data storage layer can be divided into persistent storage and non-persistent storage according to the database type. The persistent storage is an ORACLE or TDSQL database, and the non-persistent storage mainly uses REDIS. The persistent storage can be used to save the current execution job step of the pipeline, and the non-persistent storage can be used to add a REDIS distributed lock when recovering the interrupted pipeline after the container is restarted.

[0035] Embodiment one

[0036] Figure 2 A flowchart of a pipeline operation recovery method provided for the first embodiment of the application. This embodiment can be applied to the automatic recovery of pipeline operations after rolling restart of applications on the cloud. The method can be executed by a pipeline operation recovery device, which can be realized in the form of hardware and / or software, and can be configured in an electronic device such as a server. As shown, the method includes: Figure 2

[0037] S110, in the case of executing rolling restart on a container orchestration platform, destroying a first container through a container scheduling adapter.

[0038] Wherein the container orchestration platform can be Kubernetes or other container orchestration platforms, which are not specifically limited here. The first container can be an old Pod scheduled for destruction.

[0039] S120, sending a shutdown signal of the first container to a signal processing engine through the container scheduling adapter, and triggering a graceful shutdown mechanism by the signal processing engine.

[0040] Wherein the shutdown signal can be a SIGTERM signal or other signals used to request process termination.

[0041] ​Specifically, Kubernetes sends a SIGTERM signal to the signal processing engine through the container scheduling adapter, the signal processing engine uses the Spring container to capture the SIGTERM signal through the life cycle hook (@PreDestroy), and triggers the graceful shutdown mechanism. The delay shutdown time of the graceful shutdown mechanism can be set in advance.

[0042] It should be noted that the signal processing engine uses the Spring container to capture the SIGTERM signal through the life cycle hook (@PreDestroy), without relying on third-party components or modifying the cluster configuration, and adapts to the native operations such as Kubernetes rolling update and scaling.

[0043] S130, closing the task receiving interface through the pipeline controller, terminating the currently executing pipeline, and writing the information of the currently executing pipeline into the data storage layer, the information of the currently executing pipeline including the pipeline state and the job step currently executed by the pipeline.

[0044] Among them, the information of the currently executing pipeline can include one or more of the pipeline state, the job step currently executed by the pipeline, the pipeline identifier, the context data of the pipeline and the timestamp, and the pipeline state is running (Running) or execution completed (Completed).

[0045] Specifically, the pipeline controller can immediately close the task receiving interface to reject new pipeline requests, call the Interrupt method of the thread to terminate the currently executing pipeline, write the information of the currently executing pipeline into the data storage layer, and realize real-time state persistence of the execution state of each job step of the pipeline.

[0046] S140, generating a second container through the container scheduling adapter, sending an opening signal of the second container to the signal processing engine through the container scheduling adapter, and starting the pipeline controller by the signal processing engine.

[0047] Among them, the second container is a new Pod generated.

[0048] Specifically, Kubernetes generates a new Pod through the container scheduling adapter, sends an opening signal of the new Pod to the signal processing engine through the container scheduling adapter, and starts the initialization process of the pipeline controller by the signal processing engine.

[0049] S150, detecting the pipeline state as the pipeline in execution through the pipeline controller in the data storage layer, and performing job recovery on the pipeline in execution based on the job step currently executed by the pipeline.

[0050] Exemplarily, the pipeline controller can detect a pipeline with a pipeline state of "Running" in the data storage layer, and then can perform job recovery on the pipeline with the pipeline state of "Running" based on a job step currently executed by the pipeline, that is, execute a job step after the job step currently executed by the pipeline, to realize accurate continuation of the pipeline breakpoint. It should be noted that, by persisting the execution state of each job step of the pipeline in real time, the embodiment can ensure that the new Pod can accurately locate the last job step that has been verified to be successful when the container is rolled over and restarted, and seamlessly recover from the next job step, thereby avoiding resource waste caused by task re-execution in the traditional scheme.

[0051] On the basis of the above embodiment, optionally, the pipeline controller detects a pipeline with a pipeline state of "Running" in the data storage layer, including: the pipeline controller queries a pipeline with a pipeline state of "Running" in the data storage layer; and based on a preset filtering condition, filters the pipeline with the pipeline state of "Running" obtained by the query.

[0052] The preset filtering condition is to delete the pipeline with the pipeline state of "Running" when the difference between the current time and the update time of the pipeline is greater than a preset time length.

[0053] Exemplarily, the pipeline with the pipeline state of "Running" is deleted when the difference between the current time and the update time of the pipeline is greater than 24 hours, so that the timeout task can be excluded.

[0054] The technical scheme of the embodiment of the application, in the case of performing rolling restart on the container orchestration platform, destroys the first container through the container scheduling adapter, then sends a shutdown signal of the first container to the signal processing engine through the container scheduling adapter, triggers the graceful shutdown mechanism through the signal processing engine, then closes the task receiving interface through the pipeline controller, terminates the pipeline currently being executed, writes the information of the pipeline currently being executed into the data storage layer, then generates the second container through the container scheduling adapter, sends an opening signal of the second container to the signal processing engine through the container scheduling adapter, starts the pipeline controller through the signal processing engine, then detects a pipeline with a pipeline state of "Running" in the data storage layer through the pipeline controller, and performs job recovery on the pipeline with the pipeline state of "Running" based on a job step currently executed by the pipeline. The above technical scheme stores the information of the pipeline currently being executed in real time when the historical container is destroyed, automatically detects the pipeline with the pipeline state of "Running" after the new container is generated, and then performs job recovery on the pipeline with the pipeline state of "Running" based on the job step currently executed by the pipeline, thereby solving the pipeline interruption problem caused by rolling restart, realizing accurate continuation of the pipeline breakpoint, and achieving full-automatic recovery.

[0055] Embodiment two

[0056] Figure 3 A flow chart of a pipeline operation recovery method provided for embodiment two of the present application, the method of the present embodiment can be combined with each optional solution of the pipeline operation recovery method provided in the above embodiments. The pipeline operation recovery method provided in the present embodiment is further optimized. Optionally, the pipeline operation recovery method based on the current job step of the pipeline comprises: performing pipeline operation recovery on the pipeline in the executing state through a distributed lock based on the current job step of the pipeline.

[0057] As shown in Figure 3 , the method comprises:

[0058] S210, in the case of performing rolling restart on the container orchestration platform, destroying the first container through the container scheduling adapter.

[0059] S220, sending a shutdown signal of the first container to the signal processing engine through the container scheduling adapter, and triggering the graceful shutdown mechanism by the signal processing engine.

[0060] S230, closing the task receiving interface through the pipeline controller, terminating the pipeline currently being executed, and writing the information of the pipeline currently being executed into the data storage layer.

[0061] S240, generating a second container through the container scheduling adapter, sending an opening signal of the second container to the signal processing engine through the container scheduling adapter, and starting the pipeline controller by the signal processing engine.

[0062] S250, detecting the pipeline in the executing state in the data storage layer through the pipeline controller, and performing pipeline operation recovery on the pipeline in the executing state through a distributed lock based on the current job step of the pipeline.

[0063] Wherein, the distributed lock can be a Redis distributed lock or other distributed lock, which is not limited here. It should be noted that the use of the distributed lock can ensure that the same pipeline is only executed in one Pod, avoiding repeated initiation of the pipeline recovery process.

[0064] On the basis of the above-mentioned embodiment, optionally, the number of the second containers is multiple; the job recovery is performed on the pipeline in the executing state by the distributed lock based on the job step currently executed by the pipeline, including: for any current second container in the multiple second containers, the current second container judges whether the lock exists based on the pipeline identifier of the pipeline; if the pipeline identifier of the pipeline in the executing state in the current second container is different from the pipeline identifier of the pipeline in the other second containers, the pipeline in the executing state in the current second container is locked; the current second container calls the pipeline controller to perform the job recovery on the pipeline in the executing state based on the job step currently executed by the pipeline through the pipeline controller; the current second container releases the lock; if the pipeline identifier of the pipeline in the executing state in the current second container is the same as the pipeline identifier of the pipeline in the other second containers, the current second container does not perform the job recovery on the pipeline in the executing state.

[0065] Exemplarily, Figure 4 is a flowchart of a pipeline job recovery method according to an embodiment of the present application. In the case of performing rolling restart by the container orchestration platform, the old Pod is destroyed through the container scheduling adapter. Kubernetes sends a SIGTERM signal to the signal processing engine through the container scheduling adapter, the signal processing engine uses the Spring container to capture the SIGTERM signal through the life cycle hook, and triggers the graceful shutdown mechanism. The pipeline controller can stop accepting the signal pipeline, terminate the pipeline execution thread, and write the information of the pipeline in the executing state into the data storage layer to realize the pipeline real-time state persistence. Further, the new Pod is generated by Kubernetes through the container scheduling adapter, the signal processing engine is sent an opening signal of the new Pod through the container scheduling adapter, and the signal processing engine starts the pipeline controller. The pipeline controller monitors the unfinished pipeline in the data storage layer, that is, detects the pipeline in the executing state. Further, the pipeline context of the pipeline in the executing state is obtained through the persistent storage query. The Redis lock is acquired, the pipeline is reconstructed based on the Redis lock, the breakpoint is continued to execute, and the pipeline job step state is updated in real time. Specifically, the pipeline job recovery can execute the subsequent job steps of the pipeline from the job step currently executed by the pipeline + 1, for example, the job step currently executed by the pipeline saved by the old Pod is step 5, and the new Pod starts to execute from step 6.

[0066] Figure 5is a flowchart of a Redis lock provided by an embodiment of the present application. Specifically, to prevent multiple new Pods from detecting an unfinished pipeline from the database at the same time during a rolling restart, and repeatedly initiating the pipeline recovery process, the embodiment uses a Redis lock mechanism to achieve mutual exclusivity. As shown in Figure 5 , a distributed lock can be implemented using the Redis setnx operation. When multiple Pods attempt to acquire the lock for the pipeline recovery process, a setnx request is sent, with the lock key being the pipeline ID and the value being the Pod ID. If the lock already exists, it indicates that another Pod has already started executing the pipeline recovery process, and the current Pod does not need to repeat the execution. If the lock does not exist, the locking operation is successfully performed and a timeout period is set. If the lock is not released in time due to a crash, the Redis lock will be automatically released when the timeout period is reached. The Pod that successfully acquires the lock will call the pipeline controller to execute the pipeline recovery process through the pipeline controller, and release the lock after the process is completed.

[0067] The technical solution of the embodiment of the present application ensures that the same pipeline is only executed in one Pod, avoiding repeated initiation of the pipeline recovery process, by using a distributed lock.

[0068] Embodiment Three

[0069] Figure 6 is a structural schematic diagram of a pipeline job recovery device provided by Embodiment Three of the present application. As shown in Figure 6 , the device includes:

[0070] The container destruction module 310 is configured to destroy the first container through the container scheduling adapter in the case of a rolling restart of the container orchestration platform.

[0071] The graceful shutdown module 320 is configured to send a shutdown signal of the first container to a signal processing engine through the container scheduling adapter, and the signal processing engine triggers a graceful shutdown mechanism.

[0072] The executing pipeline information storage module 330 is configured to close a task receiving interface through a pipeline controller, terminate the currently executing pipeline, and write information of the currently executing pipeline into a data storage layer, the information of the currently executing pipeline including a pipeline state and a job step currently executed by the pipeline.

[0073] The container generation module 340 is configured to generate a second container through the container scheduling adapter, send an opening signal of the second container to the signal processing engine through the container scheduling adapter, and start the pipeline controller through the signal processing engine.

[0074] The pipeline operation recovery module 350 is configured to detect, by the pipeline controller, a pipeline in an execution state in the data storage layer, and perform operation recovery on the pipeline in the execution state based on a current operation step of the pipeline.

[0075] The technical solution of the embodiment of the application can achieve the following effects: in the case of performing rolling restart on the container orchestration platform, the first container is destroyed by the container scheduling adapter, the container scheduling adapter sends a shutdown signal of the first container to the signal processing engine, the signal processing engine triggers the graceful shutdown mechanism, the task receiving interface is closed by the pipeline controller, the current executing pipeline is terminated, the information of the current executing pipeline is written into the data storage layer, the second container is generated by the container scheduling adapter, the container scheduling adapter sends an opening signal of the second container to the signal processing engine, the signal processing engine starts the pipeline controller, the pipeline controller detects a pipeline in an execution state in the data storage layer, and performs operation recovery on the pipeline in the execution state based on a current operation step of the pipeline. The above technical solution can store the information of the current executing pipeline in real time when the historical container is destroyed, automatically detect a pipeline in an execution state after the new container is generated, and perform operation recovery on the pipeline in the execution state based on a current operation step of the pipeline, thereby solving the problem of pipeline interruption caused by rolling restart, and achieving accurate breakpoint resuming and full-automatic recovery of the pipeline.

[0076] In some optional embodiments, the pipeline operation recovery module 350 comprises:

[0077] The distributed lock pipeline recovery unit is configured to perform operation recovery on the pipeline in the execution state based on the current operation step of the pipeline by using the distributed lock.

[0078] In some optional embodiments, the number of the second containers is multiple.

[0079] The distributed lock pipeline recovery unit is specifically configured to:

[0080] For any current second container in the multiple second containers, the current second container is configured to determine whether the lock exists based on the pipeline identifier of the pipeline.

[0081] If the pipeline identifier of the pipeline in the current second container in the pipeline state of being executed is different from the pipeline identifier of the pipeline in the other second container, the pipeline in the pipeline state of being executed in the current second container is locked; the current second container calls the pipeline controller, and the pipeline controller performs job recovery on the pipeline in the pipeline state of being executed based on the job step currently executed by the pipeline; and the current second container releases the lock;

[0082] If the pipeline identifier of the pipeline in the current second container in the pipeline state of being executed is the same as the pipeline identifier of the pipeline in the other second container, the current second container does not perform job recovery on the pipeline in the pipeline state of being executed.

[0083] In some optional embodiments, the graceful shutdown module 320 is specifically configured to:

[0084] The signal processing engine uses the Spring container to capture the shutdown signal of the first container through a life cycle hook, and triggers the graceful shutdown mechanism.

[0085] In some optional embodiments, the executing pipeline information storage module 330 is specifically configured to:

[0086] The Interrupt method is called to terminate the currently executing pipeline.

[0087] In some optional embodiments, the information of the currently executing pipeline further includes one or more of a pipeline identifier, a job step sequence number of the pipeline, context data of the pipeline, and a timestamp, and the pipeline state is in execution or execution completion.

[0088] In some optional embodiments, the pipeline job recovery module 350 is specifically configured to:

[0089] The pipeline controller is called to query the pipeline in the pipeline state of being executed in the data storage layer;

[0090] The pipeline in the pipeline state of being executed obtained through the query is filtered based on a preset filtering condition, and the preset filtering condition is that the pipeline is deleted in the case that a difference between a current time and an update time of the pipeline is greater than a preset time length.

[0091] The pipeline job recovery apparatus provided in the embodiments of the present application can perform the pipeline job recovery method provided in any of the embodiments of the present application, and has the corresponding functional modules and beneficial effects of the execution method.

[0092] Embodiment four

[0093] Figure 7 A structural diagram of an electronic device 10 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the applications described and / or claimed in this document.

[0094] As shown in Figure 7 The electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., communicatively connected to the at least one processor 11, where the memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer programs stored in the read-only memory (ROM) 12 or loaded into the random access memory (RAM) 13 from the storage unit 18. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An I / O interface 15 is also connected to the bus 14.

[0095] Various components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc., an output unit 17, such as various types of displays, speakers, etc., a storage unit 18, such as a magnetic disk, an optical disk, etc., and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0096] The processor 11 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as the pipeline job recovery method, which includes:

[0097] destroying, by the container scheduling adapter, the first container in the case that the rolling restart is performed by the container orchestration platform;

[0098] sending a stop signal of the first container to a signal processing engine through the container scheduling adapter, the signal processing engine triggering a graceful shutdown mechanism;

[0099] closing a task receiving interface through the pipeline controller, terminating a currently executing pipeline, writing information of the currently executing pipeline into a data storage layer, the information of the currently executing pipeline including a pipeline state and a job step currently executed by the pipeline;

[0100] generating a second container through the container scheduling adapter, sending a start signal of the second container to the signal processing engine through the container scheduling adapter, the signal processing engine starting the pipeline controller;

[0101] detecting a pipeline in execution in the data storage layer through the pipeline controller, performing job recovery on the pipeline in execution based on a job step currently executed by the pipeline.

[0102] In some embodiments, the pipeline job recovery method can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as storage unit 18. In some embodiments, parts or all of the computer program can be loaded and / or installed onto electronic device 10 via, for example, ROM 12 and / or communication unit 19. When the computer program is loaded onto RAM 13 and executed by processor 11, one or more steps of the pipeline job recovery method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the pipeline job recovery method by way of other means (for example, by way of firmware).

[0103] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0104] Computer programs for implementing the methods of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer program, when executed, can cause instructions defined in the flow charts and / or block diagrams to be implemented. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, and partially on a remote machine or entirely on a remote machine or server.

[0105] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0106] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0107] The systems and techniques described herein can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described herein, or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0108] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.

[0109] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in series, or executed in different orders, as long as the desired results of the technical solutions of the present disclosure can be achieved, and the present disclosure is not limited herein.

[0110] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method of recovery of in- process, characterized by, Comprising: In the case of performing rolling restart on a container orchestration platform, destroying the first container through a container scheduling adapter; Sending a shutdown signal of the first container to a signal processing engine through the container scheduling adapter, the signal processing engine triggering a graceful shutdown mechanism; Closing a task receiving interface through a pipeline controller, terminating a currently executing pipeline, and writing information of the currently executing pipeline into a data storage layer, the information of the currently executing pipeline including a pipeline state and a job step currently executed by the pipeline; Generating a second container through the container scheduling adapter, and sending a startup signal of the second container to the signal processing engine through the container scheduling adapter, the signal processing engine starting the pipeline controller; Detecting a pipeline in an executing state in the data storage layer through the pipeline controller, and performing job recovery on the pipeline in the executing state based on the job step currently executed by the pipeline.

2. The method of claim 1, wherein, The job recovery on the pipeline in the executing state based on the job step currently executed by the pipeline comprises: Performing the job recovery on the pipeline in the executing state based on the job step currently executed by the pipeline through a distributed lock.

3. The method of claim 2, wherein, The number of the second containers is multiple; The job recovery on the pipeline in the executing state based on the job step currently executed by the pipeline comprises: For any current second container in the multiple second containers, the current second container determines whether a lock exists based on a pipeline identifier of the pipeline; If the pipeline identifier of the pipeline in the executing state in the current second container is different from pipeline identifiers of pipelines in other second containers, the pipeline in the executing state in the current second container is locked, the current second container calls the pipeline controller to perform the job recovery on the pipeline in the executing state based on the job step currently executed by the pipeline, and the current second container releases the lock; If the pipeline identifier of the pipeline in the executing state in the current second container is the same as the pipeline identifiers of the pipelines in the other second containers, the current second container does not perform the job recovery on the pipeline in the executing state.

4. The method of claim 1, wherein, The signal processing engine triggering the graceful shutdown mechanism comprises: The signal processing engine uses a Spring container to capture the shutdown signal of the first container through a life cycle hook, and triggers the graceful shutdown mechanism.

5. The method of claim 1, wherein, The termination of the currently executing pipeline comprises: Calling an Interrupt method to terminate the currently executing pipeline.

6. The method of claim 1, wherein, The information of the currently executing pipeline further comprises one or more of a pipeline identifier, context data of the pipeline, and a timestamp, and the pipeline state is in the executing state or an execution completion state.

7. The method of claim 1, wherein, The detection of the pipeline in the executing state in the data storage layer through the pipeline controller comprises: query, by the pipeline controller, a pipeline whose pipeline state is in execution in the data storage layer; filter, based on a preset filtering condition, the pipelines whose pipeline state is in execution obtained by the querying, wherein the preset filtering condition is that a difference between a current time and an update time of a pipeline is greater than a preset time length.

8. An apparatus for recovery of in- process work, characterized by comprise: a container destruction module configured to destroy the first container by a container scheduling adapter when the container orchestration platform performs a rolling restart; an elegant shutdown module configured to send, by the container scheduling adapter, a shutdown signal of the first container to a signal processing engine, and trigger an elegant shutdown mechanism by the signal processing engine; an executing pipeline information storage module configured to close, by a pipeline controller, a task receiving interface, terminate a currently executing pipeline, and write information of the currently executing pipeline into a data storage layer, the information of the currently executing pipeline comprising a pipeline state and a job step currently executed by the pipeline; a container generation module configured to generate, by the container scheduling adapter, a second container, send, by the container scheduling adapter, an opening signal of the second container to the signal processing engine, and start the pipeline controller by the signal processing engine; a pipeline job recovery module configured to detect, by the pipeline controller, a pipeline whose pipeline state is in execution in the data storage layer, and perform job recovery on the pipeline whose pipeline state is in execution based on a job step currently executed by the pipeline.

9. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the pipeline job recovery method in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to implement the pipeline job recovery method in any one of claims 1-7 when executed. The computer readable storage medium stores computer instructions for enabling the processor to implement the pipeline job recovery method in any one of claims 1-7 when executed.