Task state recovery method and device, equipment and storage medium
By recording the key information of the task when the worker node fails, the problem of inaccurate task status caused by worker node failure is solved, the accurate recovery of task status and determination of post-processing strategy are achieved, and repeated task execution and data competition are avoided.
Patent Information
- Application Number
- CN202510852340.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-10-10
AI Technical Summary
In the existing technology, when a worker node fails, it is difficult to accurately obtain the true status of the task, resulting in repeated or simultaneous execution of tasks, causing data competition and business logic confusion.
Key task files are used to record key information of tasks at different stages, including task status files, task process execution tracking files, and task execution files. The information in these files is combined to determine the actual status of the task after abnormal recovery, and reported to the scheduling node to determine the post-processing strategy.
When a worker node fails, accurately record the task execution status to avoid repeated and simultaneous execution, ensure the accuracy of task status recovery, and avoid data risks.
Smart Images

Figure CN120762837A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of script execution technology, and in particular to a task status recovery method, apparatus, device, and storage medium. Background Art
[0002] The architecture of a task scheduling product is generally divided into two parts: the master (scheduling) node, responsible for scheduling decisions, and the worker (worker) nodes, responsible for executing scheduled tasks. The master node consists of the master VM itself and the master application running on it. The master application is responsible for deciding when tasks can be scheduled. The worker node, similarly, consists of the worker VM itself and the worker application running on it. The worker application is responsible for receiving task execution requests from the master, initiating actual task execution, tracking the execution process, and ultimately reporting execution results, logs, and other information back to the master. When a worker node fails, it's impossible to accurately obtain information about the situation at the time of the failure, nor is it possible to accurately obtain the actual status of the task during the failure after recovery. Therefore, the task is often resubmitted. If the task at the time of the worker node failure has already completed but has not yet been updated, the task will be executed twice. Furthermore, if the worker application fails but the worker VM is still running, the task process may still be running on the VM. Blindly resubmitting the task without verifying the existence of the failed task process can cause two identical processes to be initiated simultaneously, leading to data races or business logic confusion.
[0003] The above content is only used to assist in understanding the technical solution of the present invention and does not constitute an admission that the above content is prior art. Summary of the Invention
[0004] The main purpose of this application is to provide a task status recovery method, device, equipment and storage medium, aiming to solve the technical problem in the prior art that it is difficult to accurately obtain the true status of the task when a working node fails, which easily leads to repeated execution of tasks or simultaneous execution of tasks.
[0005] To achieve the above objectives, the present application provides a task status recovery method, which includes:
[0006] Receive the target task issued by the scheduling node, and run the target task based on the task message corresponding to the target task;
[0007] Recording key information of the target task at different stages of execution through key task files, wherein the key task files include at least one of a task status file, a task process execution tracking file, and a task execution file;
[0008] After abnormal recovery, determining the actual status of the target task based on the key information recorded in the key task file;
[0009] The real state of the target task is reported to the scheduling node, so that the scheduling node determines a post-processing strategy for the target task based on the real state of the target task.
[0010] In one embodiment, the critical information includes at least an execution status and an operating system return code. After abnormal recovery, the step of determining the actual status of the target task based on the critical information recorded in the critical task file includes:
[0011] After abnormal recovery, scan the preset status directory;
[0012] When the task status file exists in the preset status directory, obtaining the execution status recorded in the execution status field of the task status file;
[0013] When the execution status is execution success or execution failure, the execution status is used as the real status of the target task;
[0014] When the execution state is the submitted state or the running state, the real state of the target task is determined based on the operating system return code recorded in the task execution file and the return code field in the task process execution tracking file.
[0015] In one embodiment, when the execution state is the submitted state or the running state, the step of determining the actual state of the target task based on the task process execution tracking file and the task execution file includes:
[0016] When the execution state is a submission state or a running state, executing a preset viewing command;
[0017] When the process corresponding to the task execution file cannot be obtained through the preset viewing command, scanning the preset tracking directory;
[0018] When the task process execution tracking file exists in the preset tracking directory, obtaining the operating system return code recorded in the return code field of the task process execution tracking file;
[0019] When the operating system return code is empty, determining that the real state of the target task is an unknown state;
[0020] When the operating system return code is not empty, the real state of the target task is determined based on the operating system return code.
[0021] In an embodiment, the step of determining the real status of the target task based on the operating system return code comprises:
[0022] When the operating system return code meets the preset value, the real status of the target task is determined as execution success.
[0023] When the operating system return code does not meet the preset value, the real status of the target task is determined as execution failure.
[0024] In an embodiment, the step of recording the key information of the target task at different task stages in the execution process through the key task file comprises:
[0025] When the target task enters the task initialization stage, the task message of the target task is consumed to determine the task information of the target task.
[0026] A task status file is newly created under a preset state directory, and the fields of the task status file are initialized based on the task information of the target task. The name prefix of the task status file at least includes the task identification number and the task execution times of the target task, and the fields of the task status file at least include a task identification code field, a task name field, a task type field, an execution times field, an execution date field, an execution status field, an actual start running time field, an actual end running time field, a business return code field, a business return log field, an operating system return code field, an operating system return log field, an execution agent name field, a post-processing rule field, and a command line field. The execution status recorded in the execution status field is initialized as a submission state.
[0027] Based on the initialization result of the task status file, a result message is sent to a target message queue.
[0028] In an embodiment, the step of sending a result message to a target message queue based on the initialization result of the task status file further comprises:
[0029] When the initialization result of the task status file is initialization success, it is determined that the target task enters the task execution stage, and the execution command in the task message is parsed.
[0030] The execution command is encapsulated to form the task execution file.
[0031] Based on the task execution file, a local task process is started, and the execution status recorded in the execution status field of the task status file is updated to a running state.
[0032] The task process execution tracking file is generated in a preset tracking directory, and the fields of the task process execution tracking file are updated. The name prefix of the task process execution tracking file includes at least the task identification number and the number of task executions of the target task. The fields of the task process execution tracking file include at least the task type field, the process number field and the return code field.
[0033] In one embodiment, after the step of generating the task process execution tracking file in the preset tracking directory and updating the fields of the task process execution tracking file, the step further includes:
[0034] When the local task process is correctly executed and completed, determining that the target task enters the task status write-back phase, and updating the obtained operating system return code to the return code field of the task process execution tracking file;
[0035] updating the execution status of the target task according to the operating system return code, and writing the updated execution status back to the scheduling node;
[0036] When the write-back is successful, the task status file is moved to a preset write-back success directory.
[0037] In addition, to achieve the above-mentioned purpose, the present application also proposes a task status recovery device, which includes:
[0038] An information recording module is used to receive a target task issued by a scheduling node and run the target task based on a task message corresponding to the target task;
[0039] The information recording module is further configured to record key information of the target task at different task stages during execution through key task files, wherein the key task files include at least one of a task status file, a task process execution tracking file, and a task execution file;
[0040] A status confirmation module is used to determine the real status of the target task based on the key information recorded in the key task file after abnormal recovery;
[0041] The status confirmation module is further configured to report the real status of the target task to the scheduling node, and determine a post-processing strategy for the target task based on the real status of the target task.
[0042] In addition, to achieve the above-mentioned purpose, the present application also proposes a task status recovery device, which includes: a memory, a processor, and a computer program stored in the memory and runnable on the processor, and the computer program is configured to implement the steps of the task status recovery method as described above.
[0043] In addition, to achieve the above-mentioned purpose, the present invention also proposes a storage medium, which is a computer-readable storage medium and stores a computer program. When the computer program is executed by a processor, the steps of the task state recovery method as described above are implemented.
[0044] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the task state recovery method as described above are implemented.
[0045] The present application provides a task status recovery method, which receives a target task issued by a scheduling node, executes the target task based on a task message corresponding to the target task; records key information of the target task at different task stages during execution through a key task file, wherein the key task file includes at least one of a task status file, a task process execution tracking file, and a task execution file; after abnormal recovery, determines the true status of the target task based on the key information recorded in the key task file; reports the true status of the target task to the scheduling node, so that the scheduling node determines the post-processing strategy of the target task based on the true status of the target task. The present application can accurately record the task execution status when a worker node fails, and perform on-site task recovery when the worker service is restored. By combining the true execution status at the time of the failure and the time of failure recovery, a comprehensive judgment is made to accurately obtain the true status of the task, avoid repeated execution and simultaneous execution, avoid data risks, and solve the technical problem that it is difficult to accurately obtain the true status of the task when a working node fails, which easily leads to repeated execution of tasks or simultaneous execution of tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0047] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0048] Figure 1 This is a flowchart of the first embodiment of the task status recovery method of this application;
[0049] Figure 2 This is a flow chart of the second embodiment of the task status recovery method of this application;
[0050] Figure 3A task execution diagram of the task state recovery method provided in Example 2 of the present application;
[0051] Figure 4 A schematic diagram of a brief process of the task status recovery method provided in Example 2 of the present application;
[0052] Figure 5 This is a schematic diagram of the module structure of the task state recovery device according to an embodiment of the present application;
[0053] Figure 6 This is a schematic diagram of the device structure of the hardware operating environment involved in the task state recovery method in the embodiment of the present application.
[0054] The realization of the objectives, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0055] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0056] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0057] The main solution of the embodiment of the present application is: receiving the target task issued by the scheduling node, and executing the target task based on the task message corresponding to the target task; recording the key information of the target task at different task stages during the execution process through the key task file, and the key task file includes at least one of the task status file, the task process execution tracking file and the task execution file; after the abnormality is recovered, the real status of the target task is determined based on the key information recorded in the key task file; the real status of the target task is reported to the scheduling node, so that the scheduling node determines the post-processing strategy of the target task based on the real status of the target task.
[0058] Currently, when a worker node fails, it's impossible to accurately obtain on-site information at the time of the failure, nor is it possible to accurately obtain the actual status of tasks during the failure period after recovery. Consequently, tasks are often simply resubmitted. If the task at the time of the worker node failure has already completed but the status hasn't been written back in time, the task will be repeated. Furthermore, if a worker application fails while the worker VM is still running, the task's process may still be running on the VM. Blindly resubmitting the task without verifying the existence of the failed task's process can cause two identical processes to be initiated simultaneously, leading to data races or confusing business logic.
[0059] The application provides a solution, which accurately records task execution when a worker node fails, and performs task site recovery when the worker service recovers, accurately obtains the real state of the task by combining the real execution at the time of failure and the real execution at the time of failure recovery, and avoids the situations of repeated execution and simultaneous execution, solves the technical problem that it is difficult to accurately obtain the real state of the task when the worker node fails, and easily leads to repeated execution of the task or simultaneous execution of the task.
[0060] The application embodiment provides a task state recovery method, referring to Figure 1 , Figure 1 The application embodiment provides a task state recovery method, referring to
[0061] In the embodiment, the task state recovery method includes steps S10-S40:
[0062] Step S10, receiving a target task issued by a scheduling node, and running the target task based on a task message corresponding to the target task;
[0063] It should be noted that the execution subject of the embodiment is a worker node, that is, a worker node, which can also be recorded as a node / agent node, and the embodiment takes the worker node as an example for description. The scheduling node is a selector node, which can also be recorded as a master node, and the embodiment takes the selector node as an example for description. In addition to the selector node and the worker node, the embodiment also sets a newday node.
[0064] In addition, it should be noted that the whole scheduling logic architecture includes a plurality of newday nodes, a plurality of selector nodes and a plurality of worker nodes. Among them, the newday node is mainly responsible for the daily cut disk, and is responsible for the preparation of the next day's task scheduling and the cleaning of the previous day's task; the selector node is responsible for the scheduling selection and judgment of the task; the worker node is responsible for the real running of the task script. When the task running precondition is ready, the selector node will issue the task responsible for scheduling to the corresponding worker node for running. The selector node and the worker node communicate through RabbitMQ (an open source message middleware based on the Advanced Message Queuing Protocol), and the task issuing and state writing back are all through RabbitMQ.
[0065] It is understandable that this embodiment can achieve disaster recovery of task status after worker failure under a distributed architecture scheduling cluster. In the specific implementation, the task will switch from the initial init (initialization) state to different waiting states after continuous selection and judgment by the selector node. For example: when the upstream dependent service is not ready, it will switch to the waitDependency (wait dependency) state. When the upstream time is not met, it will switch to the waitTime (wait time) state. When all the prerequisites of the task are ready, the selector node will send the task to the corresponding worker node through RabbitMQ for execution.
[0066] It should be understood that the target task is the task that the selector node currently sends to the worker node for execution, and the task message is the message containing information related to the target task. There can be many types of task types for target tasks, including at least script tasks (such as shell, python, etc.), database programs (such as Sql, stored procedures, etc.), and Http interface calls. In the task issuance stage, the selector node will send the target task that meets the scheduling conditions to the corresponding worker node through RabbitMQ, waiting for the worker node to consume. The scheduling conditions can be set according to actual conditions, and this embodiment does not make specific restrictions on this.
[0067] Step S20, recording key information of the target task at different task stages during execution using a key task file, wherein the key task file includes at least one of a task status file, a task process execution tracking file, and a task execution file;
[0068] It should be noted that the true status of the target task is the true execution status of the target task. If an exception / failure occurs in the worker node, after recovery, it is necessary to be able to accurately obtain the true execution status of the target task so as to take corresponding post-processing measures to achieve task disaster recovery. In this embodiment, key information of the target task at different task stages during the execution process is recorded through key task files. This key information can be used to determine the true status of the target task. There are usually two situations when an exception / failure occurs in the worker node. Situation one: the worker application is abnormal and the worker virtual machine is normal. Situation two: both the worker application and the worker virtual machine are abnormal. This embodiment is compatible with both situations.
[0069] It is understood that key task files include at least one of the following: a task status file, a task process execution tracking file, and a task execution file. The task status file, the task process execution tracking file, and the task execution file are generated / updated during different task phases. Therefore, the true status of the target task can be accurately determined based on the key information recorded in the task status file, the task process execution tracking file, and the task execution file. During specific implementation, due to the occurrence of failures / abnormalities, the target task may not be able to complete all task phases. Therefore, the following situations may occur: only the task status file exists; the task status file and the task execution file exist; or the task status file, the task process execution tracking file, and the task execution file all exist.
[0070] It should be understood that the key information, i.e., the information used to determine the actual status, includes at least the execution status and the operating system return code. The execution status refers to the execution status of the target task recorded in the key file, and the operating system return code refers to the direct return code after the program is executed at the operating system level. In this embodiment, the execution status is recorded in the execution status field of the task status file, and the operating system return code is recorded in the operating system return code field (sysCod) of the task status file. Simultaneously, the operating system return code is also recorded in the return code field (exitcode) of the task process execution tracking file.
[0071] Step S30, after abnormality recovery, determining the real status of the target task based on the key information recorded in the key task file;
[0072] It should be noted that exception recovery refers to the recovery of a worker node after an exception or failure. At this point, the worker node can generally be considered to have successfully restarted. After exception recovery, the actual status of the target task is determined using key information recorded in the task status file, task process execution tracking file, and task execution file.
[0073] In a feasible implementation, step S30 may include steps S301 to S304:
[0074] Step S301, after abnormality recovery, scan the preset status directory;
[0075] It should be noted that the preset status directory is the pre-set directory for generating task status files, namely the " / status / " directory. After abnormal recovery, the " / status / " directory is first scanned to traverse all task status files therein.
[0076] Step S302: when the task status file exists in the preset status directory, obtaining the execution status recorded in the execution status field of the task status file;
[0077] It is understandable that if a task status file exists in the preset status directory, the execution status of the target task recorded in the execution status field is obtained from the task status file.
[0078] Furthermore, when the execution status recorded in the execution status field in the task status file cannot be obtained, it is determined that the real status of the target task is an unknown status.
[0079] It should be understood that if the execution status cannot be obtained normally, the actual status can be considered to be an unknown state.
[0080] Step S303: When the execution status is success or failure, the execution status is used as the real status of the target task;
[0081] It should be noted that when the execution status is success (successful execution) or error (failed execution), it means that the target task has been completed at the time of the failure and has a final status. However, because the status file is still in the " / status / " directory, it means that the status writeback is not completed. At this time, the recorded execution status is directly returned to the selector node as the real status. It should be noted that if the status writeback has been completed when the failure occurs, but there is no time to migrate the task status file from the " / status / " directory to the " / backup / " directory, duplicate status writebacks may occur at this time. However, the selector node implements idempotence through the state machine. Even if two identical status writeback messages are received, the selector node can handle them correctly. Among them, the " / backup / " directory is the preset writeback success directory, that is, the directory where the task status file is stored after the status writeback is successful.
[0082] Step S304 , when the execution state is the submitted state or the running state, the actual state of the target task is determined based on the operating system return code recorded in the return code field in the task execution file and the task process execution tracking file.
[0083] It should be noted that if the execution status is "submitted," the target task has been submitted but has not yet begun execution. If the execution status is "running," the target task has begun execution. If the execution status is "submitted" or "running," the actual status of the target task must be determined by further checking the operating system return code recorded in the return code field of the task execution file and the task process execution tracking file.
[0084] In a feasible implementation, step S304 may include steps S3041 to S3045:
[0085] Step S3041, when the execution state is the commit state or the running state, executing a preset viewing command;
[0086] It should be noted that the preset viewing command is "ps-ef|grep Runfile", and the process corresponding to the task execution file is obtained by executing the "ps-ef|grep Runfile" command in the virtual machine. The task execution file is the Runfile file.
[0087] Step S3042, when the process corresponding to the task execution file cannot be obtained through the preset viewing command, scanning a preset tracking directory;
[0088] It can be understood that the preset tracking directory is a directory for generating a task process execution tracking file, that is, the " / procid / " directory. If the process corresponding to the task execution file cannot be obtained, it means that the target task is not currently executing. At this time, the " / procid / " directory is scanned to obtain the task process execution tracking file.
[0089] Further, when the process corresponding to the task execution file is obtained through the preset viewing command, it is determined that the real state of the target task is the running state.
[0090] It should be understood that if the process corresponding to the task execution file can be obtained, it means that the target task is still executing. At this time, the task state file is ignored, and the real state is determined again by waiting for the next time to reiterate the " / status / " directory.
[0091] Step S3043, when the task process execution tracking file exists in the preset tracking directory, obtaining the operating system return code recorded in the return code field of the task process execution tracking file;
[0092] It should be noted that if the task process execution tracking file exists in the " / procid / " directory, the operating system return code recorded in the return code field (exitcode) of the task process execution tracking file is obtained.
[0093] Further, when the task process execution tracking file does not exist in the preset tracking directory and the execution state is the commit state, it is determined that the real state of the target task is the not started state; when the task process execution tracking file does not exist in the preset tracking directory and the execution state is the running state, it is determined that the real state of the target task is the unknown state.
[0094] It can be understood that if the task process execution tracking file does not exist and the execution state recorded in the task state file is the submit state, it indicates that the target task has not been formally initiated for execution at the time of failure, at which time the real state can be considered to be the unstarted state, and subsequent re-initiation of execution can be performed; if the task process execution tracking file does not exist and the execution state recorded in the task state file is the running state, the real state cannot be determined, at which time the real state can be considered to be the unknow state.
[0095] Step S3044, when the operating system return code is empty, determining that the real state of the target task is the unknow state.
[0096] It can be understood that the operating system return code is empty (null), that is, the operating system return code has no value. If the operating system return code has no value, it indicates that no valid operating system return code has been generated, that is, the target task has not been executed to completion, and the real state cannot be determined, at which time the real state can be considered to be the unknow state.
[0097] Step S3045, when the operating system return code is not empty, determining the real state of the target task based on the operating system return code.
[0098] It should be noted that the operating system return code is not empty, that is, the operating system return code has a value. If the operating system return code has a value, it indicates that the target task has been executed to completion, at which time the recorded execution state can be returned to the selector node as the real state.
[0099] Specifically, when the operating system return code conforms to the preset value, it is determined that the real state of the target task is execution success; when the operating system return code does not conform to the preset value, it is determined that the real state of the target task is execution failure.
[0100] It can be understood that the preset value is a value corresponding to execution success that is preset in advance, and is usually 0. If the operating system return code is 0, it indicates that the real state of the target task is execution success, and if the operating system return code is not 0, it indicates that the real state of the target task is execution failure.
[0101] Step S40, reporting the real state of the target task to the scheduling node, so that the scheduling node determines the post-processing strategy of the target task based on the real state of the target task.
[0102] It should be noted that the post-processing strategy is a post-processing measure. The real state of the target task needs to be reported to the selector node in order to determine the corresponding post-processing strategy.
[0103] In a feasible implementation, when the actual state of the target task is the running state, the post-processing strategy is determined to be to wait for a preset time and then rescan the preset state directory; when the actual state of the target task is the unstarted state, the post-processing strategy is determined to be to re-execute the target task; when the actual state of the target task is the unknown state, the post-processing strategy is determined to be to notify manual intervention; when the actual state of the target task is successful execution, the post-processing strategy is determined to execute a new task; when the actual state of the target task is failed execution, the post-processing strategy is determined to re-execute the target task.
[0104] It is understood that the actual status in this embodiment includes five situations: not started, running, successful, failed, and unknown. If the actual status of the target task is not started, the target task is re-executed. If the actual status of the target task is running, the " / status / " directory is traversed again next time and step S301 is re-executed. If the actual status of the target task is successful, the next new task can be executed. If the actual status of the target task is failed, the target task can be re-executed. If the actual status of the target task is unknown, the monitoring system is triggered to actively notify manual intervention.
[0105] The present embodiment provides a task status recovery method, which receives a target task issued by a scheduling node, executes the target task based on a task message corresponding to the target task; records key information of the target task at different task stages during execution through a key task file, wherein the key task file includes at least one of a task status file, a task process execution tracking file, and a task execution file; after abnormal recovery, determines the true state of the target task based on the key information recorded in the key task file; reports the true state of the target task to the scheduling node, so that the scheduling node determines the post-processing strategy of the target task based on the true state of the target task. The present embodiment can accurately record the task execution status when a worker node fails, and perform on-site task recovery when the worker service is restored. By combining the true execution status at the time of the failure and the time of failure recovery, a comprehensive judgment is made to accurately obtain the true state of the task, thereby avoiding repeated execution and simultaneous execution.
[0106] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above embodiment 1 can be referred to the above introduction and will not be described in detail later. Figure 2 , step S20 may include steps S201 to S203:
[0107] Step S201, when the target task enters the task initialization phase, consuming the task message of the target task to determine the task information of the target task;
[0108] It should be noted that the task information of a target task refers to the relevant information of the target task. During the task initialization phase, after the worker node consumes the task message, it will obtain the task information of the target task in order to generate the task status file. In this embodiment, the task phase includes the task issuance phase, the task initialization phase, the task execution phase, and the task status writeback phase.
[0109] Step S202: creating a new task status file in a preset status directory, and initializing fields of the task status file based on the task information of the target task;
[0110] It should be noted that this embodiment will create a new task status file in the " / status / " directory. The name prefix of the task status file includes at least the task identification number of the target task and the task execution count. The task identification number is the unique ID of the target task, and the task execution count is the current execution count of the target task (a task may be executed multiple times). In a specific implementation, the name prefix of the task status file can use the form of "[task identification number]_[task execution count]", for example: 12423455314_2.status, where 12423455314 is the task identification number and 2 is the task execution count (indicating the second execution).
[0111] In addition, it should be noted that the task status file mainly records various types of task information and execution status changes. The fields of the task status file include at least the task identification code field (orderId), the task name field (jobName), the task type field (jobType), the execution count field (runCount), the execution date field (odate), the execution status field (status), the actual start time field (startTime), the actual end time field (endTime), the service return code field (rtnCod), the service return log field (rtnMsg), the operating system return code field (sysCod), the operating system return log field (sysMsg), the execution agent name field (agentName), the post-processing rule field (onDoList), and the command line field (cmdline).
[0112] It can be understood that the task identification code field is used to record the task identification code of the target task, the task name field is used to record the task name of the target task, the task type field is used to record the task type of the target task, the execution count field is used to record the execution count of the target task, the execution date field is used to record the execution date of the target task, the execution status field is used to record the execution status of the target task, the actual start running time field is used to record the actual start running time of the target task, the actual end running time field is used to record the actual end running time of the target task, the business return code field is used to record the business return code of the target task, the business return log field is used to record the business return log of the target task, the operating system return code field is used to record the operating system return code of the target task (the direct return code after the program is executed from the operating system level), the operating system return log field is used to record the operating system return log of the target task (the direct return log after the program is executed from the operating system level), the execution agent name field is used to record the name of the execution agent where the target task is running, the post-processing rule field is used to record the post-processing rule of the target task, and the command line field is used to record the command line of the target task. The execution status recorded in the execution status field is initialized to the submit state during the task initialization phase, indicating that the task is in the submitted state and has not yet been actually executed. The execution status will be updated in subsequent task phases.
[0113] For example, the information contained in the fields in the task status file is: {"orderId":"oderid_123321","jobName":"TEST_JOBNAME_1","jobType":"OS","runCount":1,"odate":"200807","status":"Submit","startTime":"20200807170620","endTime":"20200807170808","rtnCod":null,"rtnMsg":null,"sysCod":null,"sysMsg":null,"agentName":"Agent1","onDoList":null,"cmdline":"echo 1"}.
[0114] Step S203: Sending a result message to a target message queue based on the initialization result of the task status file.
[0115] It is understandable that the target message queue is RabbitMQ. The initialization results include two situations: initialization success and initialization failure. If the initialization result of the task status file is initialization success, it means that the task status file has completed initialization. At this time, the worker node will send an ACK (Acknowledgment) confirmation message to RabbitMQ, and the result message is the ACK confirmation message. If the initialization result of the task status file is initialization failure, an UNACK (Unacknowledged) message will be sent to RabbitMQ, and the result message is the UNACK message. In this way, even if the local disk write is slow or the text write is abnormal, the task execution message (task MQ message) sent by RabbitMQ will not be lost.
[0116] In a feasible implementation manner, after step S203, steps A11 to A14 may be included:
[0117] Step A11, when the initialization result of the task status file is initialization success, determining that the target task enters the task execution phase, and parsing the execution command in the task message;
[0118] It should be noted that after the task status file is initialized, the next phase begins: task execution. The worker node parses the execution command in the task message and starts the task process locally on the virtual machine.
[0119] Step A12: encapsulate the execution command to form the task execution file;
[0120] It is understandable that the task execution file (Runfile) needs to be used to start the task process. The Runfile file is a layer of encapsulation of the execution command. When the worker application fails but the virtual machine is normal, it can ensure that the task can continue to execute and the final execution result of the task can be obtained. For example, assuming that the actual execution command of the task is "sh sum_business_count.sh", the worker application will package the execution command into "(sh sum_business_count.sh)" >> [task execution log file]; echo'EXITCODE'$? >> [task status tracking file]". There are two main additions here. The first part is to add redirection to the script execution output log, and the second part will monitor the execution result of the script and obtain the operating system return code after the script is executed through "$?". When the worker executes the task, the Runfile file is actually initiated. The Runfile file can be named "[task identification number]_[task execution times].runfile".
[0121] Step A13: Based on the task execution file, start the local task process and update the execution status recorded in the execution status field in the task status file to the running status;
[0122] It is understandable that the target task actually starts to execute at this time, and the worker application will update the execution status field in the task status file to running, that is, the running status.
[0123] Step A14, generates the task process execution tracking file under the preset tracking directory, and updates the fields of the task process execution tracking file, the name prefix of the task process execution tracking file includes at least the task identification number and the number of task executions of the target task, and the fields of the task process execution tracking file include at least the task type field, the process number field and the return code field.
[0124] It should be noted that this embodiment generates a task process execution tracking file in the " / procid / " directory. The name prefix of the task process execution tracking file includes at least the task identification number of the target task and the number of task executions. In a specific implementation, the name prefix of the task status file can use the form of "[task identification number]_[task execution number]", for example: 12423455314_2.runFile, where 12423455314 is the task identification number and 2 is the number of task executions (indicating the second execution).
[0125] It is understood that the fields in the task process execution tracking file include at least the task type field (jobtype), the process ID field (pid), and the return code field (exitcode). The task type field is used to record the task type of the target task, the process ID field is used to record the operating system process ID of the started operating system, and the return code field is used to record the operating system return code after the process execution is completed. The value of the return code field is only updated when the process execution ends successfully.
[0126] In a feasible implementation manner, after step A14, steps B11 to B13 may be included:
[0127] Step B11, when the local task process is correctly executed and completed, determining that the target task enters the task status write-back phase, and updating the obtained operating system return code to the return code field of the task process execution tracking file;
[0128] It should be noted that after a task completes, the corresponding operating system process ends and the next task phase, the task status writeback phase, begins. The worker application obtains the operating system return code of the process and updates it to the exitcode field in the task process execution trace file. If the task completes successfully, the exitcode is 0. If the task fails, the exitcode is updated to the corresponding system return code, for example, 255. This embodiment does not impose specific restrictions on this.
[0129] Step B12: updating the execution status of the target task according to the operating system return code, and writing the updated execution status back to the scheduling node;
[0130] It can be understood that the current execution state is determined according to the operating system return code, and then the current execution state is written back to the selector node.
[0131] Step B13: When the write-back is successful, the task status file is moved to a preset write-back success directory.
[0132] As you can understand, the default directory for successful writebacks is usually the / backup / directory, and the default directory for failed writebacks is usually the / exception / directory. If the writeback succeeds, the task status file is moved to the / backup / directory; if the writeback fails, it is moved to the / exception / directory. A process within the worker service scans the task status files in the / exception / directory and attempts to rewrite the task status to the selector node.
[0133] refer to Figure 3 The selector node sends tasks through RabbitMQ, the worker node receives task MQ messages (task execution messages), initializes the task status file in the task initialization phase, updates the task status file in the task execution phase, and generates a task process execution tracking file. In the task status writeback phase, the task status file and the task process execution tracking file are updated. According to the execution status recorded in the updated task status file and the operating system return code recorded in the task process execution tracking file, the task status is determined and reported to the selector node.
[0134] This embodiment provides a task status recovery method. When a target task enters the task initialization phase, the task message of the target task is consumed to determine the task information of the target task. A new task status file is created in a preset status directory. Based on the task information of the target task, the fields of the task status file are initialized. Based on the initialization result of the task status file, a result message is sent to the target message queue. This embodiment can accurately record the task execution status when a worker node fails and perform on-site task recovery when the worker service is restored. By combining the actual execution status at the time of the failure and the failure recovery, a comprehensive judgment is made to accurately obtain the actual status of the task, avoiding repeated execution and simultaneous execution.
[0135] For example, in order to help understand the implementation process of the task state recovery method obtained by combining this embodiment with the above embodiment 2, please refer to Figure 4 , Figure 4 A brief flowchart of a task status recovery method is provided, specifically:
[0136] 1) After the worker node recovers from a fault, it first scans the status directory, traverses all task status files in it, and obtains the status field of each task. If it cannot be obtained normally, the task status is updated to unknow and returned to the selector node.
[0137] 2) If the task status can be obtained from the task status file, when the task status is success (execution success) or error (execution failure), it means that the task has been completed at the time of the failure and has a final status. However, because the status file is still in the status directory, it means that the status writeback is not completed. At this time, the task status is directly returned to the selector node.
[0138] 3) If the task status is submit (the task has been submitted but has not yet started) or running (the task has started), execute the "ps-ef | grep Runfile" command in the virtual machine to determine whether the process is still running in the virtual machine. If the process exists, it means that the task is still running. Ignore the task status file and wait for the next time to traverse the status directory and re-determine.
[0139] 4) If the process does not exist, determine whether the process execution tracking file exists. If the process execution tracking file exists, determine whether the exit code in the file has a value. If so, it means that the task has been completed. At this time, the task status can be reported based on the execution result. If there is no value, the task is updated to unknow. If the process execution tracking file does not exist, and the task status in the task status file is submit, it means that the task has not been officially initiated for execution at the time of the failure, and the task can be re-initiated for execution. If the process execution tracking file does not exist, and the task status in the task status file is running, it means that the status cannot be determined, and the task status is updated to unknow.
[0140] 5) For tasks in the unknown state, the monitoring system will be triggered to proactively notify human intervention.
[0141] It should be noted that the above examples are only used to understand this application and do not constitute a limitation on the task state recovery method of this application. More simple transformations based on this technical concept are all within the scope of protection of this application.
[0142] This application also provides a task status recovery device, please refer to Figure 5 , the task state recovery device includes:
[0143] The information recording module 10 is used to receive the target task issued by the scheduling node and run the target task based on the task message corresponding to the target task;
[0144] The information recording module 10 is further configured to record key information of the target task at different stages of execution through key task files, wherein the key task files include at least one of a task status file, a task process execution tracking file, and a task execution file;
[0145] A status confirmation module 20 is used to determine the actual status of the target task based on the key information recorded in the key task file after abnormal recovery;
[0146] The status confirmation module 20 is further configured to report the real status of the target task to the scheduling node, and determine a post-processing strategy for the target task based on the real status of the target task.
[0147] In a feasible implementation manner, the state confirmation module 20 is further configured to scan a preset state directory after abnormality recovery;
[0148] When the task status file exists in the preset status directory, obtaining the execution status recorded in the execution status field of the task status file;
[0149] When the execution status is execution success or execution failure, the execution status is used as the real status of the target task;
[0150] When the execution state is the submitted state or the running state, the real state of the target task is determined based on the operating system return code recorded in the task execution file and the return code field in the task process execution tracking file.
[0151] In a feasible implementation manner, the state confirmation module 20 is further configured to execute a preset viewing command when the execution state is a submission state or a running state;
[0152] When the process corresponding to the task execution file cannot be obtained through the preset viewing command, scanning the preset tracking directory;
[0153] When the task process execution tracking file exists in the preset tracking directory, obtaining the operating system return code recorded in the return code field of the task process execution tracking file;
[0154] When the operating system return code is empty, determining that the real state of the target task is an unknown state;
[0155] When the operating system return code is not empty, the real state of the target task is determined based on the operating system return code.
[0156] In a feasible implementation manner, the status confirmation module 20 is further configured to determine that the actual status of the target task is successful execution when the operating system return code meets a preset value;
[0157] When the operating system return code does not conform to the preset value, it is determined that the actual status of the target task is execution failure.
[0158] In a feasible implementation manner, the information recording module 10 is further configured to consume the task message of the target task when the target task enters the task initialization phase to determine the task information of the target task;
[0159] A new task status file is created in a preset status directory, and based on the task information of the target task, the fields of the task status file are initialized, wherein the name prefix of the task status file includes at least the task identification number of the target task and the number of task executions, and the fields of the task status file include at least a task identification code field, a task name field, a task type field, an execution number field, an execution date field, an execution status field, an actual start time field, an actual end time field, a business return code field, a business return log field, an operating system return code field, an operating system return log field, an execution agent name field, a post-processing rule field, and a command line field, wherein the execution status recorded in the execution status field is initialized to a submission state;
[0160] Based on the initialization result of the task status file, a result message is sent to the target message queue.
[0161] In a feasible implementation manner, the information recording module 10 is further configured to, when the initialization result of the task status file is that initialization is successful, determine that the target task enters the task execution phase, and parse the execution command in the task message;
[0162] Encapsulating the execution command to form the task execution file;
[0163] Based on the task execution file, start the local task process and update the execution status recorded in the execution status field of the task status file to the running status;
[0164] The task process execution tracking file is generated in a preset tracking directory, and the fields of the task process execution tracking file are updated. The name prefix of the task process execution tracking file includes at least the task identification number and the number of task executions of the target task. The fields of the task process execution tracking file include at least the task type field, the process number field and the return code field.
[0165] In a feasible embodiment, the information recording module 10 is further configured to determine that the target task has entered the task status write-back phase when the local task process is correctly executed, and update the obtained operating system return code to the return code field of the task process execution tracking file;
[0166] updating the execution status of the target task according to the operating system return code, and writing the updated execution status back to the scheduling node;
[0167] When the write-back is successful, the task status file is moved to a preset write-back success directory.
[0168] The task state recovery device provided in this application, which utilizes the task state recovery method of the aforementioned embodiment, can resolve the technical problem of difficulty in accurately obtaining the true state of a task when a working node fails, which can easily lead to repeated or simultaneous task execution. Compared to the prior art, the beneficial effects of the task state recovery device provided in this application are the same as those of the task state recovery method provided in the aforementioned embodiment, and the other technical features of the task state recovery device are the same as those disclosed in the aforementioned embodiment method, and are not further described here.
[0169] The present application provides a task state recovery device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the task state recovery method in the above-mentioned embodiment one.
[0170] Reference below Figure 6 , which shows a schematic diagram of the structure of a task state recovery device suitable for implementing an embodiment of the present application. The task state recovery device in the embodiment of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The task status recovery device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0171] like Figure 6As shown, the task state recovery device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in ROM (Read Only Memory) 1002 or programs loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the task state recovery device. Processing device 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input device 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output device 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage device 1003 including, for example, a magnetic tape, hard disk, etc.; and communication device 1009. The communication device 1009 can allow the task state recovery device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a task state recovery device with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems can be implemented or provided instead.
[0172] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0173] The task state recovery device provided in this application, which utilizes the task state recovery method of the aforementioned embodiment, can resolve the technical problem of difficulty in accurately obtaining the true state of a task when a working node fails, which can easily lead to repeated or simultaneous task execution. Compared to the prior art, the beneficial effects of the task state recovery device provided in this application are the same as those of the task state recovery method provided in the aforementioned embodiment, and the other technical features of the task state recovery device are the same as those disclosed in the method of the previous embodiment, and are not further described here.
[0174] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0175] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0176] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, a computer program) stored thereon, wherein the computer-readable program instructions are used to execute the task state recovery method in the above-mentioned embodiment.
[0177] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0178] The computer-readable storage medium may be included in the task status recovery device; or may exist independently without being assembled into the task status recovery device.
[0179] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the task status recovery device, the task status recovery device: receives the target task issued by the scheduling node, and executes the target task based on the task message corresponding to the target task; records the key information of the target task at different task stages during the execution process through the key task file, and the key task file includes at least one of the task status file, the task process execution tracking file and the task execution file; after the abnormal recovery, determines the true status of the target task based on the key information recorded in the key task file; reports the true status of the target task to the scheduling node, so that the scheduling node determines the post-processing strategy of the target task based on the true status of the target task.
[0180] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0181] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0182] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0183] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., a computer program) for executing the above-mentioned task state recovery method. This can solve the technical problem that it is difficult to accurately obtain the true state of a task when a working node fails, which easily leads to repeated or simultaneous execution of tasks. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the task state recovery method provided in the above-mentioned embodiment, and will not be repeated here.
[0184] The present application also provides a computer program product, including a computer program, which implements the steps of the task state recovery method as described above when the computer program is executed by a processor.
[0185] The computer program product provided in this application can resolve the technical problem of difficulty in accurately obtaining the true status of tasks when a worker node fails, which can easily lead to repeated or simultaneous task execution. Compared to the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the task status recovery method provided in the above-mentioned embodiment, and will not be elaborated here.
[0186] The above are only some embodiments of the present application and are not intended to limit the patent scope of the present application. All equivalent structural transformations made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A task status recovery method, characterized in that: The method comprises: Receive the target task issued by the scheduling node, and run the target task based on the task message corresponding to the target task; Recording key information of the target task at different stages of execution through key task files, wherein the key task files include at least one of a task status file, a task process execution tracking file, and a task execution file; After abnormal recovery, determining the actual status of the target task based on the key information recorded in the key task file; The real state of the target task is reported to the scheduling node, so that the scheduling node determines a post-processing strategy for the target task based on the real state of the target task.
2. The method according to claim 1, wherein The key information includes at least an execution status and an operating system return code. After the abnormal recovery, the step of determining the actual status of the target task based on the key information recorded in the key task file includes: After abnormal recovery, scan the preset status directory; When the task status file exists in the preset status directory, obtaining the execution status recorded in the execution status field of the task status file; When the execution status is execution success or execution failure, the execution status is used as the real status of the target task; When the execution state is the submitted state or the running state, the real state of the target task is determined based on the operating system return code recorded in the task execution file and the return code field in the task process execution tracking file.
3. The method according to claim 2, wherein When the execution state is the submitted state or the running state, the step of determining the actual state of the target task based on the task process execution tracking file and the task execution file includes: When the execution state is a submission state or a running state, executing a preset viewing command; When the process corresponding to the task execution file cannot be obtained through the preset viewing command, scanning the preset tracking directory; When the task process execution tracking file exists in the preset tracking directory, obtaining the operating system return code recorded in the return code field of the task process execution tracking file; When the operating system return code is empty, determining that the real state of the target task is an unknown state; When the operating system return code is not empty, the real state of the target task is determined based on the operating system return code.
4. The method according to claim 3, wherein The step of determining the actual status of the target task based on the operating system return code includes: When the operating system return code meets a preset value, determining that the actual status of the target task is successful execution; When the operating system return code does not conform to the preset value, it is determined that the actual status of the target task is execution failure.
5. The method according to any one of claims 1 to 4, characterized in that The step of recording key information of the target task at different task stages during execution through key task files includes: When the target task enters the task initialization phase, the task message of the target task is consumed to determine the task information of the target task; A new task status file is created in a preset status directory, and based on the task information of the target task, the fields of the task status file are initialized, wherein the name prefix of the task status file includes at least the task identification number of the target task and the number of task executions, and the fields of the task status file include at least a task identification code field, a task name field, a task type field, an execution number field, an execution date field, an execution status field, an actual start time field, an actual end time field, a business return code field, a business return log field, an operating system return code field, an operating system return log field, an execution agent name field, a post-processing rule field, and a command line field, wherein the execution status recorded in the execution status field is initialized to a submission state; Based on the initialization result of the task status file, a result message is sent to the target message queue.
6. The method according to claim 5, wherein After the step of sending a result message to a target message queue based on the initialization result of the task status file, the step further includes: When the initialization result of the task status file is that the initialization is successful, determining that the target task enters the task execution phase, and parsing the execution command in the task message; Encapsulating the execution command to form the task execution file; Based on the task execution file, start the local task process and update the execution status recorded in the execution status field of the task status file to the running status; The task process execution tracking file is generated in a preset tracking directory, and the fields of the task process execution tracking file are updated. The name prefix of the task process execution tracking file includes at least the task identification number and the number of task executions of the target task. The fields of the task process execution tracking file include at least the task type field, the process number field and the return code field.
7. The method according to claim 6, wherein After the step of generating the task process execution tracking file in the preset tracking directory and updating the fields of the task process execution tracking file, the method further includes: When the local task process is correctly executed and completed, determining that the target task enters the task status write-back phase, and updating the obtained operating system return code to the return code field of the task process execution tracking file; updating the execution status of the target task according to the operating system return code, and writing the updated execution status back to the scheduling node; When the write-back is successful, the task status file is moved to a preset write-back success directory.
8. A task status recovery device, characterized in that: The device comprises: An information recording module is used to receive a target task issued by a scheduling node and run the target task based on a task message corresponding to the target task; The information recording module is further configured to record key information of the target task at different task stages during execution through key task files, wherein the key task files include at least one of a task status file, a task process execution tracking file, and a task execution file; A status confirmation module is used to determine the real status of the target task based on the key information recorded in the key task file after abnormal recovery; The status confirmation module is further configured to report the real status of the target task to the scheduling node, and determine a post-processing strategy for the target task based on the real status of the target task.
9. A task status recovery device, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the task state recovery method according to any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the task state recovery method according to any one of claims 1 to 7 are implemented.