Method, device, computer equipment and storage medium for executing computing tasks
By introducing a fault tolerance mechanism in the MPI task execution system, re-execute failed tasks according to the preset fault tolerance level, the problem of MPI tasks crashing when nodes crash or network disconnection is solved, the task success rate and execution efficiency are improved, and the operation and maintenance costs are reduced.
Patent Information
- Application Number
- CN202111015422.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-31
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-08-31
AI Technical Summary
In Linux, when MPI tasks are running in a cluster, if the nodes crash or the network is disconnected, it will cause the death of all MPI processes, reduce the success rate of MPI tasks, and increase the execution time and operation and maintenance costs of MPI tasks.
It provides an execution method of an operation task, obtains execution instructions through the first main library node, carries the operation task to be executed and the preset fault tolerance level, and re-executes the failed task according to the preset fault tolerance level. Specific measures include derived child processes to execute tasks, listen for exit codes, re-execute them in a timely manner, or recovering by generating key point files and migrating tasks to the slave library node.
It improves the success rate of MPI tasks, improves the efficiency of MPI tasks execution, and reduces operation and maintenance costs.
Smart Images

Figure CN114036005B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and particularly to a method, apparatus, computer device, and storage medium for executing an operation task. Background Art
[0002] MPI (Message Passing Interface) defines an interface method, including protocols and semantic descriptions. The goal of MPI is high performance, large scale, and portability. MPI remains the main model for high-performance computing today.
[0003] When an MPI task runs in a cluster, each machine (physical machine or virtual machine) can be regarded as a node (process). Currently, in Linux, if an MPI process encounters an exception, such as a node crashing or the network disconnecting, the default error handling mechanism of MPI will be adopted, that is, as long as any one MPI process makes an error, the process manager will send a process termination signal, and all processes will receive this signal, which will cause the death of all MPI processes, resulting in a reduction in the success rate of executing MPI tasks, an increase in the execution time of MPI tasks, and an increase in operation and maintenance costs. Summary of the Invention
[0004] Embodiments of this application provide a method, apparatus, computer device, and storage medium for executing an operation task, which can improve the success rate of executing MPI tasks, improve the execution efficiency of MPI tasks, and reduce operation and maintenance costs.
[0005] In a first aspect, an embodiment of this application provides a method for executing an operation task, which is applied to an execution system of an operation task. The execution system of the operation task includes a first main library node. The method for executing the operation task includes:
[0006] The first main library node obtains an execution instruction, and the execution instruction carries an operation task to be executed and a preset fault tolerance level;
[0007] The first main library node executes the operation task to be executed according to the preset fault tolerance level;
[0008] If the first main library node fails to execute the operation task to be executed, the operation task to be executed is re-executed according to the preset fault tolerance level.
[0009] In the method for executing an operation task provided by the embodiment of this application, if the preset fault tolerance level is the first level, the first main library node executes the operation task to be executed according to the preset fault tolerance level, including:
[0010] The first main library node spawns a child process;
[0011] The subprocess executes the to-be-executed computing task, and the first master database node monitors the execution exit code of the subprocess.
[0012] In the method for executing a computing task provided in an embodiment of the present application, before the first master database node fails to execute the to-be-executed computing task, it includes:
[0013] If the first master database node monitors that the execution exit code is not the preset exit code, it is determined that the first master database node fails to execute the to-be-executed computing task;
[0014] Re-executing the to-be-executed computing task according to the preset fault tolerance level includes: returning to execute the step of the first master database node spawning a subprocess until the execution exit code is monitored as the preset exit code.
[0015] In the method for executing a computing task provided in an embodiment of the present application, the execution system of the computing task further includes a slave database node corresponding to the first master database node. If the preset fault tolerance level is the second level, the first master database node executes the to-be-executed computing task according to the preset fault tolerance level, including:
[0016] The first master database node executes the to-be-executed computing task, generates a key point file, and sends the key point file to the slave database node.
[0017] In the method for executing a computing task provided in an embodiment of the present application, re-executing the to-be-executed computing task according to the preset fault tolerance level includes:
[0018] If the first master database node fails to execute the to-be-executed computing task, resume executing the to-be-executed computing task according to the key point file.
[0019] In the method for executing a computing task provided in an embodiment of the present application, the execution system of the computing task further includes a main control node, and the main control node is used to manage the first master database node and the slave database node. The slave database nodes corresponding to the first master database node include multiple slave database nodes. The method for executing the computing task further includes:
[0020] If the main control node determines that the first master database node fails, the main control node migrates the to-be-executed computing task to a target slave database node among the multiple slave database nodes, and upgrades the target slave database node to a second master database node. The target slave database node is the node corresponding to the target key value among the multiple slave database nodes;
[0021] The second master database node resumes executing the to-be-executed computing task according to the key point file.
[0022] In the method for executing an operation task provided in the embodiment of the present application, before the master node determines that the first primary database node fails, the method further includes:
[0023] The first primary database node periodically sends process status information to the master node;
[0024] If the process status information received by the master node is a preset string, it is determined that the first primary database node is valid;
[0025] If the process status information received by the master node is not a preset string, it is determined that the first primary database node fails.
[0026] In a second aspect, an embodiment of the present application further provides an apparatus for executing an operation task, which is applied to an operation task execution system. The operation task execution system includes a first primary database node. The apparatus for executing an operation task includes:
[0027] An acquisition module, configured to acquire an execution instruction for the first primary database node, where the execution instruction carries an operation task to be executed and a preset fault tolerance level;
[0028] An execution module, configured to execute the operation task to be executed by the first primary database node according to the preset fault tolerance level;
[0029] A re - execution module, configured to re - execute the operation task to be executed according to the preset fault tolerance level if the first primary database node fails to execute the operation task to be executed.
[0030] In a third aspect, an embodiment of the present application further provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above - mentioned method are implemented.
[0031] In a fourth aspect, an embodiment of the present application further provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above - mentioned method are implemented.
[0032] An embodiment of the present application provides a method, an apparatus, a computer device, and a storage medium for executing an operation task, which are applied to an operation task execution system. The execution system includes a first primary database node. The execution method includes: the first primary database node acquires an execution instruction, where the execution instruction carries an operation task to be executed and a preset fault tolerance level; the first primary database node executes the operation task to be executed according to the preset fault tolerance level; if the first primary database node fails to execute the operation task to be executed, it re - executes the operation task to be executed according to the preset fault tolerance level, thereby improving the success rate of executing the MPI task, improving the efficiency of MPI task execution, and reducing the operation and maintenance cost. Description of the Drawings
[0033] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0034] Figure 1 It is a schematic flowchart of an execution method for an operation task provided by an embodiment of the present application.
[0035] Figure 2 It is a schematic structural diagram of an execution system for an operation task provided by an embodiment of the present application.
[0036] Figure 3 It is a schematic module diagram of an execution system for an operation task provided by an embodiment of the present application
[0037] Figure 4 It is a schematic diagram of a first application scenario of an execution method for an operation task provided by an embodiment of the present application;
[0038] Figure 5 It is a schematic diagram of a second application scenario of an execution method for an operation task provided by an embodiment of the present application;
[0039] Figure 6 It is a schematic diagram of a third application scenario of an execution method for an operation task provided by an embodiment of the present application;
[0040] Figure 7 It is a schematic structural diagram of an execution device for an operation task provided by an embodiment of the present application;
[0041] Figure 8 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Specific embodiments
[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0043] The embodiments of the present application provide an execution method, device, computer device, and storage medium for an operation task. Specifically, this embodiment provides an execution method for an operation task applicable to an execution device for an operation task, and the execution device for the operation task can be integrated in a computer device.
[0044] Please refer to Figure 1 ,Figure 1 A flowchart of a method for executing a computing task provided in an embodiment of the present application is provided. The method for executing a computing task is applied to an execution system for a computing task. The execution system for the computing task may include a master library node. The method may mainly include steps 101 to 103. The description of each step is as follows:
[0045] Step 101: The first master library node obtains an execution instruction, which carries a computing task to be executed and a preset fault tolerance level.
[0046] The computing task to be executed may be an MPI task. When the MPI task runs in a cluster, each machine (physical machine or virtual machine) may be regarded as a node. The first main library node may be a machine in the cluster.
[0047] For easier understanding, see Figure 2 , Figure 2 The schematic diagram of the structure of the execution system of the computing task provided in the embodiment of the present application first describes the architecture of the execution system of the computing task. The execution system of the computing task mainly includes a user layer, a task scheduling layer, a master node layer and a data storage layer, wherein the user can submit the MPI task to the task scheduling layer through the user layer. For example, the user layer can be a request program (sender) for the user to submit the task. Then, the task scheduling layer detects whether there is an available working node (worker) in the data storage layer. Specifically, the worker can be divided into a master library node and a slave library node. Each master library node can correspond to multiple slave library nodes. The task scheduling layer can be a server program (recver_ex), which can interact with the user and the master control node. If there are still resources, the SQL statement in the MPI task is parsed, and the selected master-slave worker is allowed to execute the SQL statement at the same time. If the user does not need to perform a secondary comparison, the task scheduling layer will reduce the results of the comparison of multiple workers and feed them back to the user. If the task submitted by the user requires a secondary comparison, after the master and slave workers have executed the SQL query, they will write the query results to the temporary file on the disk for the secondary comparison. The initiator of the secondary comparison is the task scheduling layer, which will select one of the available master nodes (master) from the master node layer, and then start an MPI task execution command for it, and the master will select the corresponding master library node from the available workers to execute the MPI task. Every time a worker completes its own task, it will feed back the results to the master. When the master has collected all the worker results, it will perform a final reduction and feed back the reduction results to the task scheduling layer, and then the task scheduling layer will pass the results to the user who submitted the task request.
[0048] Specifically, the master can send an execution instruction to the first master database node. The execution instruction carries the MPI task to be executed and a preset fault tolerance level specified by the user. In subsequent steps, the master will determine which fault tolerance strategy should be selected for recovery when the MPI task execution fails based on the preset fault tolerance level.
[0049] Please refer to Figure 3 , Figure 3 which is a schematic diagram of the modules of the execution system for computing tasks provided by the embodiments of the present application, where:
[0050] Client, the user module. The user can submit task jobs through the user module.
[0051] Job-manager, the job management module, integrated in the master. It is responsible for accepting jobs remotely submitted by users, and in a non-blocking manner, saving the job requests submitted by users into a queue, and is also responsible for job scheduling.
[0052] FT, the fault tolerance module. In this system, it is mainly divided into two types of fault tolerance functions. One is the fault tolerance for socket communication, and the other is the fault tolerance for MPI communication, to ensure that the jobs submitted by users will not be unable to be completed due to external factors such as process errors, network anomalies, and system crashes.
[0053] Middler-layer, the middleware module, which is one of the core modules of the system. It mainly processes operations such as SQL statement parsing, user request connection processing, and redirection of data addition, deletion, modification, etc. And after each worker finishes executing the task, it finally needs to feedback to the master, and finally be passed back to the middleware module for reduction processing, and then feedback the result to the user.
[0054] Monitor, the monitoring module, which can be integrated in the master and also belongs to the core module of the system. It is used to monitor the survival status of all workers in the cluster, the master-slave switch, and the list of machines on which the MPI program depends.
[0055] Step 102: The first master database node executes the to-be-executed computing task according to the preset fault tolerance level.
[0056] It is easy to understand that when the user uploads an MPI task, a preset fault tolerance level can be specified for the MPI task. For example, the preset fault tolerance level can include the first level, the second level, etc. Correspondingly, the execution system of the computing task is set with multiple fault tolerance strategies. When the execution of the to-be-executed computing task fails, the execution system of the computing task can select the fault tolerance strategy corresponding to the preset fault tolerance level for processing.
[0057] In Linux, for the failure of MPI processes on worker nodes, by default, the default error handling mechanism of MPI will be adopted. As long as any one MPI process encounters an error, the process manager will send a process termination signal, and all processes will receive this signal, which will cause the communication domain composed of the entire master and workers to collapse, resulting in the death of all MPI processes.
[0058] The execution method of this computing task selects different execution methods for different computing tasks according to the preset fault tolerance level, and at the same time adopts different fault tolerance strategies, effectively improving the success rate of executing MPI tasks.
[0059] Step 103: If the first master database node fails to execute the to-be-executed computing task, re-execute the to-be-executed computing task according to the preset fault tolerance level.
[0060] In some embodiments, if the preset fault tolerance level is the first level, step 102 may specifically include: the first master database node spawns a child process; the child process executes the to-be-executed computing task, and the first master database node listens to the execution exit code of the child process.
[0061] It should be noted that a node is a process that performs a computing task. If the preset fault tolerance level is the first level, the first master database node can spawn a child process through the fork function provided by the Linux system. After the child process calls the execl function, it will break away from the code area shared with the original first master database node (parent process), and then the child process executes the to-be-executed computing task, while the parent process can call the wait function to listen to the execution exit code of the child process. For example, if the return value is 0, it means that the child process has executed successfully; if the return value is greater than 0, it indicates that the child process has exited abnormally.
[0062] In this embodiment, before "the first master database node fails to execute the to-be-executed computing task", the method may include: if the first master database node monitors that the execution exit code is not the preset exit code, it is determined that the first master database node fails to execute the to-be-executed computing task; step 103 may include: return to execute the step of "the first master database node spawns a child process" until the execution exit code is the preset exit code.
[0063] For example, the preset exit code can be any character or string indicating the successful execution of the child process. For example, the preset exit code can be 0. After the first master database node detects that the return value of the child process is not 0, it will call the fork function again, that is, spawn a child process again to re-execute the previously failed MPI task.
[0064] For example, please refer toFigure 4 The master sends an execution instruction to the first master database node. The execution instruction carries the operation task to be executed and indicates that the preset fault tolerance level is the first level. The first master database node spawns a child process, and the child process executes the operation task to be executed. The first master node then monitors the exit code of the child process. If the exit code returned by the child process is not 0, it indicates that the child process is abnormal. Then the first master node spawns a new child process, and the new child process re-executes the operation task to be executed. The first master database node continues to monitor the exit code. If the exit code is 0, it can be considered that the operation task of the child process is successfully executed.
[0065] In this embodiment, when the child process fails to execute the operation task, it will be detected by the parent process, thus triggering the spawning of a new child process to execute the operation task again, and there is no need to set the error handling mechanism of the process manager.
[0066] In some embodiments, the method may further include: monitoring the number of times of "the step of the first master database node spawning a child process until the execution exit code is the preset exit code" is monitored. If the number of times exceeds the maximum threshold, it is determined that the operation task to be executed fails.
[0067] It is easy to understand that, in order to avoid an infinite loop caused by the task itself, a maximum threshold can be set. When the maximum threshold is exceeded, the result of the failure of the operation task to be executed is returned to the master.
[0068] Specifically, in this embodiment, once the operation task to be executed fails, it will be re-executed from the beginning. This may not seem to have much advantage for smaller tasks, but for some larger comparison tasks, the time cost is very high. For example, a task that takes about an hour to execute exits abnormally after 90% of the time. If it is re-executed, a large amount of time will be wasted on repetitive work, which is unacceptable to users. Therefore, this case sets a fault tolerance level. For smaller tasks, the user can set the preset fault tolerance level to the first level, and for larger tasks, the user can set the preset fault tolerance level to the second level.
[0069] In some embodiments, the execution system of the operation task further includes a slave database node corresponding to the first master database node. If the preset fault tolerance level is the second level, step 102 may include: the first master database node executes the operation task to be executed, generates a key point file, and sends the key point file to the slave database node.
[0070] The second level makes improvements to address the deficiencies of the first level. Among them, a key point file is introduced, enabling users to automatically and regularly import the key data they need into the disk during task execution according to their actual needs, thereby generating a key point file (checkpoint). It should be noted that for the same task, whenever a new key point file is generated, the old key point file will be replaced. In this way, it can be ensured that if the main library node fails subsequently, there will be a key point file for the migrated task to load the intermediate result stored in the disk last time.
[0071] In this embodiment, step 103 mainly may include: if the first main library node fails to execute the to-be-executed operation task, then resume the execution of the to-be-executed operation task according to the key point file.
[0072] Specifically, in the current execution system of the operation task, the worker periodically sends some key data to the master, and the master makes a global checkpoint. Although this strategy is feasible, it also implies a relatively large problem, that is: when the number of cluster nodes is large, all workers send the intermediate results back to the same master, which will cause the communication of the master node to be too frequent and increase its burden, and even the situation of channel blockage may occur.
[0073] To solve the above situation and reduce the overhead of the master, the execution of the MPI task is only performed on the main library node, and the periodically generated checkpoint is not sent back to the master, but the main library node passes the checkpoint to the slave library nodes belonging to it. If the main library node fails subsequently, the MPI task can be migrated to the slave library node belonging to it, and the corresponding checkpoint file can be loaded to enable the task to continue to execute downward.
[0074] For example, please refer to Figure 5 , the master sends an execution instruction to the first main library node. The execution instruction carries the to-be-executed operation task and indicates that the preset fault tolerance level is the second level. Then the first main library node executes the to-be-executed operation task and periodically sends the key point file to the slave library node. Each time the slave library node receives the key point file, it uses the new key point file to replace the old key point file to update the key point file. If the first main library node fails to execute the to-be-executed operation task, it loads the key point file and resumes the progress of the to-be-executed operation task according to the key point file and continues to execute, without having to re-execute the to-be-executed operation task from the beginning. Compared with the first level, the time for task execution can be saved.
[0075] Specifically, the checkpoint is transmitted by socket multicast. The master database node obtains the dictionary of the master-slave node relationship generated by the monitoring program, and based on this, the master database node transmits the checkpoint to its corresponding slave database node. Here, a strategy is also set to handle extreme cases: if all the slave database nodes corresponding to the master database fail, the master database will not send the checkpoint to the failed slave database nodes and will only return a transmission failure.
[0076] For the checkpoint transfer program, the applicant once tried to send it in the MPI way, that is, each master database node acts as a sender and calls the MPI_Bcast function provided in MPI3 to send the checkpoint to the corresponding slave database in a broadcast manner. For the master database node, the 0th process is started in MPI_COMM_WORLD, while the slave database node will start a non-0th process. The broadcast method is that the 0th process sends data and the non-0th process receives data. However, a conclusion is drawn through experiments: when running other MPI programs, the separate call of this MPI file transfer program can run successfully, but if this MPI file transfer program is called by the system main program as an executable file, problems will occur. Here, two solutions have been tried: the first is to call the fork function through the system, and let the child process execute the checkpoint file transfer program, but it will prompt a running error. The MPI documentation mentions that in an MPI program with dynamically spawned processes (that is, MPI processes are dynamically started through MPI_Comm_spawn), calling the fork function will result in duperror, that is, an error will occur when copying the file descriptors shared by the parent and child processes; the second is to start the checkpoint file transfer program by calling MPI_Comm_spawn. However, the MPI_Comm_spawn to spawn child processes depends on the MPI_INFO parameter, which will determine on which IP the spawned child processes should run. MPI3 states regarding MPI_Comm_spawn - for the MPI_INFO parameter, it is only valid for the root process (that is, the 0th process). And in this system, at the beginning, MPI_Comm_spawn is used to spawn child processes to execute each subtask, and the process numbers in the child process communication domain range from 0 to N - 1 (N is the total number of processes in the communication domain). Then, if the file transfer program is called through MPI_Comm_spawn in these child processes, only one process can call successfully, while the other non-0th child processes simply don't know on which nodes to start receiving. In summary, the checkpoint is transmitted by socket multicast.
[0077] In this embodiment, the execution system of the operation task further includes a master node, which is used to manage the first master database node and the slave database nodes. The slave database nodes corresponding to the first master database node include multiple slave database nodes. The method for executing the operation task further includes: if the master node determines that the first master database node fails, the master node migrates the operation task to be executed to a target slave database node among the multiple slave database nodes, and upgrades the target slave database node to a second master database node. The target slave database node is the node corresponding to the target key value among the multiple slave database nodes; the second master database node resumes the execution of the operation task to be executed according to the key point file.
[0078] Among them, the master node is the master. If the master node determines that the first master database node fails, since the corresponding slave database nodes store the intermediate files (key point files) of the first master database node for executing the operation task to be executed, therefore, the master can migrate the operation task to be executed to the corresponding slave database nodes, load the intermediate files, resume the progress of the first master database node for executing the operation task to be executed, and continue the execution.
[0079] Specifically, the master will determine which slave database node the MPI task should be started in according to the MPI_INFO parameter of MPI_Comm_spawn. MPI_INFO is a data type provided by MPI3. It actually provides a way to bind the key values provided by the user together. For example, by calling MPI_Info_set, the user can set the list of machines to be bound, so that after calling MPI_Comm_spawn, the new child processes will be started on the IP nodes that appear in the list of machines. For example, MPI_Info_set(&mpi_info, "host", ". / mf"), where "host" is a fixed name provided by MPI3, and "mf" is the file name of the available node list of the user, that is, the target key value. When the MPI_Comm_spawn function is called, the IP addresses in the mf file will be parsed, and the operation task to be executed will be migrated to the slave database node corresponding to this IP to execute the remaining tasks of the previously failed master database node.
[0080] Specifically, before the step "the master node determines that the first master database node fails", it further includes: the first master database node periodically sends process status information to the master node; if the process status information received by the master node is a preset string, it is determined that the first master database node is valid; if the process status information received by the master node is not the preset string, it is determined that the first master database node fails.
[0081] In this system, in order for the master to detect a failed worker in a timely manner, the master needs to maintain a heartbeat function. That is, the worker executing the task needs to periodically send a specific string to the master. If the master can receive this content normally, it indicates that the corresponding first master database node executing the task is still valid. If the first master database node fails, it will cause the master's recv operation for this channel to fail and return a content that is not the specific string.
[0082] At the initial state when the system executes tasks, each worker is in the same communication domain. The master completes heartbeat detection by performing inter-domain communication with the worker. Once a worker fails, the master migrates the MPI task to another worker for execution. The process ID of this worker is 0, and the communication domain it forms with other workers is not the same area. Therefore, the master needs to modify the communication domain parameter in the heartbeat detection function again. Otherwise, due to the failure of heartbeat detection, the master will misjudge that this worker has failed, resulting in an abnormal phenomenon where nodes continuously fail and migrate.
[0083] The method for executing an arithmetic task provided by an embodiment of the present application is applied to an execution system for an arithmetic task. The execution system includes a first master database node. An execution instruction is obtained through the first master database node. The execution instruction carries an arithmetic task to be executed and a preset fault tolerance level. Then, the first master database node executes the arithmetic task to be executed according to the preset fault tolerance level. If the first master database node fails to execute the arithmetic task to be executed, the arithmetic task to be executed is re-executed according to the preset fault tolerance level, thereby improving the success rate of executing the MPI task, improving the efficiency of executing the MPI task, and reducing the operation and maintenance cost.
[0084] To facilitate better implementation of the method for executing an arithmetic task according to an embodiment of the present application, an embodiment of the present application further provides an apparatus for executing an arithmetic task. Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of the apparatus for executing an arithmetic task provided by an embodiment of the present application. The apparatus 10 for executing an arithmetic task is applied to an execution system for an arithmetic task. The execution system for an arithmetic task may include a first master database node. The apparatus 10 for executing an arithmetic task may include an acquisition module 11, an execution module 12, and a re-execution module 13.
[0085] Among them, the acquisition module 11 is used for the first master database node to obtain an execution instruction, and the execution instruction carries an arithmetic task to be executed and a preset fault tolerance level.
[0086] The execution module 12 is used for the first master database node to execute the arithmetic task to be executed according to the preset fault tolerance level.
[0087] The re - execution module 13 is used to re - execute the to - be - executed operation task according to a preset fault - tolerance level if the first master database node fails to execute the to - be - executed operation task.
[0088] In some embodiments, if the preset fault - tolerance level is the first level, the execution module 12 can mainly be used for: the first master database node derives a child process;
[0089] The child process executes the to - be - executed operation task, and the first master database node monitors the execution exit code of the child process.
[0090] In this embodiment, the execution device can also be used for: before the first master database node fails to execute the to - be - executed operation task, if the first master database node monitors that the execution exit code is not the preset exit code, it is determined that the first master database node fails to execute the to - be - executed operation task; the re - execution module 13 can be used for: returning to execute the step of the first master database node deriving a child process until the execution exit code is the preset exit code.
[0091] In some embodiments, the execution system of the operation task further includes a slave database node corresponding to the first master database node. If the preset fault - tolerance level is the second level, the execution module 12 can mainly be used for: the first master database node executes the to - be - executed operation task, generates a key - point file, and sends the key - point file to the slave database node.
[0092] Furthermore, the re - execution module 13 can specifically be used for: if the first master database node fails to execute the to - be - executed operation task, restoring and executing the to - be - executed operation task according to the key - point file.
[0093] In some embodiments, the execution system of the operation task further includes a main control node for managing the first master database node and the slave database nodes. The slave database nodes corresponding to the first master database node include multiple slave database nodes. The execution device 10 can further include a migration module for: if the main control node determines that the first master database node fails, the main control node migrates the to - be - executed operation task to a target slave database node among the multiple slave database nodes, and upgrades the target slave database node to a second master database node. The target slave database node is the node corresponding to the target key value among the multiple slave database nodes; the second master database node restores and executes the to - be - executed operation task according to the key - point file.
[0094] In this embodiment, the migration module can also be used for: the first master database node periodically sends process status information to the main control node; if the process status information received by the main control node is the preset string, it is determined that the first master database node is valid; if the process status information received by the main control node is not the preset string, it is determined that the first master database node fails.
[0095] The execution device 10 for computing tasks provided by the embodiments of the present application. The first master library node obtains an execution instruction through the acquisition module 11. The execution instruction carries the computing task to be executed and a preset fault tolerance level. Then, the first master library node executes the computing task to be executed according to the preset fault tolerance level through the execution module 12. If the first master library node fails to execute the computing task to be executed, the re-execution module 13 will execute again if the first master library node fails to execute the computing task to be executed, thereby improving the success rate of executing MPI tasks, improving the efficiency of MPI task execution, and reducing the operation and maintenance costs.
[0096] In addition, the embodiments of the present application further provide a computer device, which can be a terminal, and the terminal can be a terminal device such as a laptop computer, a personal computer (PC), or a personal digital assistant (PDA). As Figure 8 shown, Figure 8 is a schematic structural diagram of the computer device provided by the embodiments of the present application. The computer device 2000 includes a processor 2001 with one or more processing cores, a memory 2002 with one or more computer-readable storage media, and a computer program stored in the memory 2002 and executable on the processor. Among them, the processor 2001 is electrically connected to the memory 2002. Those skilled in the art can understand that the structural diagram of the computer device shown in the figure does not constitute a limitation on the computer device, and it may include more or fewer components than shown in the figure, or combine certain components, or arrange different components.
[0097] The processor 2001 is the control center of the computer device 2000, connecting various parts of the entire computer device 2000 through various interfaces and lines. By running or loading software programs and / or modules stored in the memory 2002, and calling data stored in the memory 2002, it executes various functions of the computer device 2000 and processes data, thereby monitoring the entire computer device 2000.
[0098] In the embodiments of the present application, the processor 2001 in the computer device 2000 will load the instructions corresponding to the processes of one or more application programs into the memory 2002 according to the following steps, and the processor 2001 will run the application programs stored in the memory 2002 to implement various functions:
[0099] The first master library node obtains an execution instruction, and the execution instruction carries the computing task to be executed and a preset fault tolerance level;
[0100] The first master library node executes the computing task to be executed according to the preset fault tolerance level;
[0101] If the first primary database node fails to execute the operation task to be executed, the operation task to be executed is re-executed according to the preset fault tolerance level.
[0102] For the specific implementation of each of the above operations, reference may be made to the foregoing embodiments, which will not be elaborated herein.
[0103] Optionally, as Figure 8 shown, the computer device 2000 further includes: a touch display screen 2003, a radio frequency circuit 2004, an audio circuit 2005, an input unit 2006, and a power supply 2007. Among them, the processor 2001 is electrically connected to the touch display screen 2003, the radio frequency circuit 2004, the audio circuit 2005, the input unit 2006, and the power supply 2007 respectively. Those skilled in the art can understand that Figure 8 the computer device structure shown does not constitute a limitation on the computer device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0104] The touch display screen 2003 can be used to display a graphical user interface and receive operation instructions generated by a user's interaction with the graphical user interface. The touch display screen 2003 may include a display panel and a touch panel. Among them, the display panel can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the computer device. These graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. The touch panel can be used to collect touch operations of the user on or near it (such as operations of the user using any suitable object or accessory such as a finger or a stylus on or near the touch panel), and generate corresponding operation instructions, and the operation instructions execute the corresponding program. Optionally, the touch panel can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user and detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 2001, and can receive and execute the commands sent by the processor 2001. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it transmits it to the processor 2001 to determine the type of touch event. Subsequently, the processor 2001 provides a corresponding visual output on the display panel according to the type of touch event. In the embodiment of the present application, the touch panel and the display panel can be integrated into the touch display screen 2003 to implement input and output functions. However, in some embodiments, the touch panel and the touch panel can be implemented as two independent components to implement input and output functions. That is, the touch display screen 2003 can also be used as part of the input unit 2006 to implement the input function.
[0105] The radio frequency circuit 2004 can be used to transmit and receive radio frequency signals to establish wireless communication with a network device or other computer devices through wireless communication, and transmit and receive signals with the network device or other computer devices.
[0106] The audio circuit 2005 can be used to provide an audio interface between a user and a computer device through a speaker and a microphone. The audio circuit 2005 can convert the received audio data into an electrical signal and transmit it to the speaker, which converts it into a sound signal for output. On the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 2005 and then converted into audio data. After the audio data is output to the processor 2001 for processing, it is sent through the radio frequency circuit 2004 to, for example, another computer device, or the audio data is output to the memory 2002 for further processing. The audio circuit 2005 may also include an earphone jack to provide communication between a peripheral earphone and the computer device.
[0107] The input unit 2006 can be used to receive input digital, character information, or user feature information (such as fingerprint, iris, facial information, etc.), and generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function controls.
[0108] The power supply 2007 is used to supply power to each component of the computer device 2000. Optionally, the power supply 2007 can be logically connected to the processor 2001 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 2007 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0109] Although Figure 8 not shown in the figure, the computer device 2000 may also include a camera, a sensor, a Wi-Fi module, a Bluetooth module, etc., which will not be elaborated here.
[0110] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0111] As can be seen from the above, the computer device provided in this embodiment obtains an execution instruction through the first master library node, and the execution instruction carries a to-be-executed operation task and a preset fault tolerance level; the first master library node executes the to-be-executed operation task according to the preset fault tolerance level; if the first master library node fails to execute the to-be-executed operation task, the to-be-executed operation task is re-executed according to the preset fault tolerance level, thereby improving the success rate of executing the MPI task, improving the efficiency of executing the MPI task, and reducing the operation and maintenance cost.
[0112] Those of ordinary skill in the art can understand that all or part of the steps in the above various methods can be completed through instructions, or through instructions to control relevant hardware. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0113] To this end, an embodiment of the present application provides a computer-readable storage medium, which stores multiple computer programs that can be loaded by a processor to execute the steps in any execution method of an operation task provided by the embodiment of the present application. For example, the computer program can execute the following steps: The first master library node obtains an execution instruction, and the execution instruction carries the operation task to be executed and a preset fault tolerance level; the first master library node executes the operation task to be executed according to the preset fault tolerance level; if the first master library node fails to execute the operation task to be executed, the operation task to be executed is re-executed according to the preset fault tolerance level.
[0114] For the specific implementation of each of the above operations, reference can be made to the previous embodiments, which will not be elaborated here.
[0115] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.
[0116] Since the computer program stored in the storage medium can execute the steps in any execution method of an operation task provided by the embodiment of the present application, the beneficial effects that can be achieved by any execution method of an operation task provided by the embodiment of the present application can be realized. For details, refer to the previous embodiments, which will not be elaborated here.
[0117] The above has introduced in detail an execution method, device, storage medium, and computer device of an operation task provided by the embodiment of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for executing an operation task, characterized in that, An execution system for an operation task, the execution system of the operation task includes a first main library node, and the execution method of the operation task includes: The first main library node obtains an execution instruction, and the execution instruction carries the operation task to be executed and a preset fault tolerance level; The first main library node executes the operation task to be executed according to the preset fault tolerance level; If the first main library node fails to execute the operation task to be executed, then re-execute the operation task to be executed according to the preset fault tolerance level; if the preset fault tolerance level is the first level, then the first main library node executes the operation task to be executed according to the preset fault tolerance level, including: The first main library node derives a child process; The child process executes the operation task to be executed, and the first main library node monitors the execution exit code of the child process; Re-executing the operation task to be executed according to the preset fault tolerance level includes: returning to the step of the first main library node deriving a child process until the execution exit code is monitored to be a preset exit code.
2. The method for executing an arithmetic task according to claim 1, wherein Before the first main library node fails to execute the operation task to be executed, it includes: If the first main library node monitors that the execution exit code is not the preset exit code, it is determined that the first main library node fails to execute the operation task to be executed.
3. The method for executing an arithmetic task according to claim 1, wherein, The execution system of the operation task further includes a slave library node corresponding to the first main library node. If the preset fault tolerance level is the second level, then the first main library node executes the operation task to be executed according to the preset fault tolerance level, including: The first main library node executes the operation task to be executed, generates a key point file, and sends the key point file to the slave library node.
4. The method for executing an arithmetic task according to claim 3, wherein Re-executing the operation task to be executed according to the preset fault tolerance level includes: If the first main library node fails to execute the operation task to be executed, then resume executing the operation task to be executed according to the key point file.
5. The method for executing an arithmetic task according to claim 4, wherein The execution system of the operation task further includes a main control node, the main control node is used to manage the first main library node and the slave library node, the slave library nodes corresponding to the first main library node include multiple slave library nodes, and the execution method of the operation task further includes: If the main control node determines that the first main library node fails, the main control node migrates the operation task to be executed to a target slave library node among the multiple slave library nodes, and upgrades the target slave library node to a second main library node, and the target slave library node is the node corresponding to the target key value among the multiple slave library nodes; The second main library node resumes executing the operation task to be executed according to the key point file.
6. The method for executing an arithmetic task according to claim 5, wherein, Before the main control node determines that the first main library node fails, it further includes: The first main library node periodically sends process status information to the main control node; If the process status information received by the main control node is a preset string, it is determined that the first main library node is valid; If the process status information received by the main control node is not the preset string, it is determined that the first main library node fails.
7. An execution device for an operation task, characterized in that, An execution system applied to an operation task, the execution system of the operation task includes a first main library node, and the execution device of the operation task includes: An acquisition module, configured to acquire an execution instruction for the first main library node, where the execution instruction carries an operation task to be executed and a preset fault tolerance level; An execution module, configured to execute the operation task to be executed by the first main library node according to the preset fault tolerance level; A re-execution module, configured to re-execute the operation task to be executed according to the preset fault tolerance level if the execution of the operation task to be executed by the first main library node fails; If the preset fault tolerance level is the first level, the first main library node executes the operation task to be executed according to the preset fault tolerance level, including: The first main library node spawns a child process; The child process executes the operation task to be executed, and the first main library node monitors the execution exit code of the child process; Re-executing the operation task to be executed according to the preset fault tolerance level includes: returning to the step of the first main library node spawning a child process until the execution exit code is monitored to be a preset exit code.
8. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. Wherein, when the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A storage medium, characterized in that, A computer program is stored, wherein when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
System, method and device for assigning tasks and computer readable medium
CN112448977A
Real-time computing task processing method and device, equipment and storage medium
CN112835924A