Task execution result notification method and device

By configuring the first queue and the second queue in the head node of the computing cluster, the problem of being unable to actively notify the execution results after the task fails in a large-scale computing cluster, and timely notification of task execution results and effective utilization of resources are achieved.

CN120011109APending Publication Date: 2025-05-16SHANGHAI BILIBILI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510125377.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In large-scale computing clusters, failures in computing tasks occur from time to time due to thread, process or node failure, but the prior art cannot actively notify the execution results of tasks.

Method used

In the target head node of the cluster, the first queue and the second queue are configured. Through the task call of the first queue management work node, the task execution status in the work node is synchronized through the second queue. Based on the information in the second queue, the head node can actively obtain the execution result of the task and send it to the requesting party.

Benefits of technology

When a worker node process fails, the head node can promptly notify the task execution results to ensure timely feedback of the results and effective utilization of resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011109A_ABST
    Figure CN120011109A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a task execution result notification method, the task execution result notification method is used in a target head node of a cluster, the cluster further comprises a plurality of working nodes, and the target head node is configured with a first queue and a second queue; the method comprises the steps of receiving a request of a target task; extracting the target reference from the first queue; calling a target process of a target working node in the plurality of working nodes through the target reference to execute the target task; obtaining target execution state information returned by the target working node based on the target task; adding a target tuple comprising the target execution state information and the target reference into a second queue to update the second queue; and extracting the tuple from the updated second queue, and sending an execution result of the corresponding task to the requester based on the extracted tuple. According to the technical scheme provided by the embodiment of the invention, the execution result of the task in the process is timely notified under the condition that the process fails.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of Internet technology, and in particular, to a method, apparatus, computer device, computer-readable storage medium, and computer program product for notifying a task execution result. Background Art

[0002] With the development of computer technology, the complexity of computing tasks is increasing, and the scale of distributed computing resources is also expanding. However, in large-scale computing clusters, computing tasks often fail due to thread, process or node failures. However, the existing technology cannot actively notify the execution results of computing tasks in the process when a process fails.

[0003] It should be noted that the above content is not necessarily prior art, nor is it intended to limit the scope of patent protection of this application. Summary of the invention

[0004] Embodiments of the present application provide a method, apparatus, computer device, computer-readable storage medium, and computer program product for notifying a task execution result to solve or alleviate one or more of the technical problems raised above.

[0005] One aspect of an embodiment of the present application provides a method for notifying a task execution result, which is used in a target head node of a cluster, wherein the cluster further includes a plurality of working nodes, and the target head node is configured with a first queue and a second queue; the method includes: Receive a request for a target task; Retrieving a target reference from the first queue, the first queue being used to store a plurality of references; Calling a target process of a target working node among the multiple working nodes through the target reference to execute the target task; Obtaining target execution status information returned by the target working node based on the target task; adding a target tuple including the target execution state information and the target reference to the second queue to update the second queue; the second queue is used to store a plurality of tuples, one tuple corresponding to one task; A tuple is taken out from the updated second queue, and based on the taken out tuple, the execution result of the corresponding task is sent to the requesting party.

[0006] Optionally, taking out a tuple from the updated second queue, and sending the execution result of the corresponding task to the requester based on the taken out tuple, including: In case of receiving a new request, querying the updated second queue; When the updated second queue is not an empty queue, the tuple at the head of the queue is taken out from the updated second queue.

[0007] Optionally, taking out a tuple from the updated second queue, and sending the execution result of the corresponding task to the requester based on the taken out tuple, further comprising: When a new tuple is added to the updated second queue, the tuple at the head of the queue is taken out from the updated second queue.

[0008] Optionally, the method further comprises: When the execution result of the corresponding task is sent to the requester, the reference of the corresponding task is added back to the first queue.

[0009] Optionally, the method further comprises: In the case where the target execution status information includes error information, all tasks executed by the target work section are marked as failed.

[0010] Optionally, the method further comprises: In the case that both the target head node and the target working node fail, task information of multiple tasks executed by the target working node is obtained through a metadata service, and failure results of the multiple tasks are returned to the requesting party.

[0011] Another aspect of an embodiment of the present application provides a task execution result notification device, which is used in a target head node of a cluster, wherein the cluster further includes a plurality of working nodes, and the target head node is configured with a first queue and a second queue; the device includes: A receiving module, used for receiving a request for a target task; A first fetching module, used for fetching a target reference from a first queue, wherein the first queue is used for storing a plurality of references; A calling module, used for calling a target process of a target working node among the multiple working nodes through the target reference to execute the target task; An acquisition module is used to acquire target execution status information returned by the target working node based on the target task; An updating module, used for adding a target tuple including the target execution state information and the target reference to the second queue to update the second queue; the second queue is used for storing a plurality of tuples, one of the tuples corresponding to one task; The second fetching module is used to fetch tuples from the updated second queue, and send the execution result of the corresponding task to the requesting party based on the fetched tuples.

[0012] Another aspect of an embodiment of the present application provides a computer device, including: at least one processor; and a memory communicatively coupled to the at least one processor; Wherein: the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.

[0013] Another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method described above is implemented.

[0014] Another aspect of an embodiment of the present application provides a computer program product, including a computer program, which implements the method described above when executed by a processor.

[0015] The embodiment of the present application adopts the above technical solution and may include the following advantages: configuring the first queue and the second queue in the head node. The task call of each working node is managed by the first queue, and the task execution status of the tasks in each working node is synchronized by the second queue. Based on the second queue (taking tuples from the second queue), the execution result of the corresponding task can be actively obtained and sent to the requester. In this way, when a working node process fails, the head node can feedback the execution result of the task to the requester based on the information in the second queue, thereby ensuring timely notification of the task execution result. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings exemplarily illustrate the embodiments and constitute a part of the specification, and together with the text description of the specification, are used to explain the exemplary implementation of the embodiments. The embodiments shown are for illustrative purposes only and do not limit the scope of the claims. In all drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0017] Figure 1 A diagram schematically shows an operating environment of a method for notifying a task execution result according to Embodiment 1 of the present application; Figure 2 A flowchart schematically shows a method for notifying a task execution result according to the first embodiment of the present application; Figure 3 Schematically shows Figure 2 Flow chart of sub-steps of step S210; Figure 4 A diagram schematically shows an application example of a method for notifying a task execution result according to Embodiment 1 of the present application; Figure 5A diagram schematically shows an application example of a method for notifying a task execution result according to Embodiment 1 of the present application; Figure 6 A block diagram schematically shows an application example of the method for notifying the task execution result according to the first embodiment of the present application according to the second embodiment of the present application; and Figure 7 The hardware architecture diagram of the computer device according to the third embodiment of the present application is schematically shown. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0019] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of the present application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in the field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0020] In the description of the present application, it should be understood that the numerical labels before the steps do not indicate the order in which the steps are executed, but are only used to facilitate the description of the present application and to distinguish each step, and therefore should not be understood as a limitation on the present application.

[0021] First, the following terms are explained: Distributed computing framework for scaling AI and Python applications, providing a computational layer for parallel processing.

[0022] Actor: A concept in the distributed computing framework. The Python class is created through the API. It is essentially a stateful process that allows parallel and distributed execution of computing tasks.

[0023] ActorHandle: A concept in the distributed computing framework, which refers to a reference to a created Actor. Through this reference, users can interact with the remote Actor, call its methods and obtain results.

[0024] ObjectRef: A reference to a distributed object in a distributed computing framework that allows objects to be passed between different processes and nodes without actually transferring the contents of the object.

[0025] Liveness detection: Determine the health status of the service by sending requests to the specified path of the service and monitoring the returned status code and response time.

[0026] Secondly, in order to facilitate those skilled in the art to understand the technical solutions provided in the embodiments of the present application, the relevant technologies are described below: The applicant understands that in large-scale computing clusters, computing tasks often fail due to thread, process, or node failures. Online computing tasks with high real-time requirements usually require multiple tasks to be performed simultaneously, with one master and one backup or one master and multiple backups. When the main line fails, it is promptly switched to the backup line to ensure the stability of the real-time task.

[0027] To this end, the embodiment of the present application provides a technical solution for notifying the results of task execution. In this technical solution: (1) the main process failure or node failure can be actively notified; (2) the task execution results can be notified in a timely manner with low latency; (3) the misjudgment caused by the passive detection timeout is avoided, which leads to waste of computing resources. See below for details.

[0028] Finally, for ease of understanding, an exemplary operating environment is provided below.

[0029] like Figure 1 As shown, the operating environment diagram includes: requester 2 and resource pool. The resource pool may include multiple clusters 4. Requester 2 is the party that requests the computing task, which can be a local application, a server, a microservice in a cloud platform, a distributed data processing system, a scheduling system, etc.

[0030] Cluster 4 can be multiple service clusters, or it can be an automatically manageable cluster formed by distributing multiple containers to multiple physical or virtual computers for operation through a container orchestration platform. Cluster 4 is used to provide distributed computing services.

[0031] The resource pool may provide distributed computing services via a network. The network may include various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or the like. The network may include physical links, such as coaxial cable links, twisted pair cable links, optical fiber links, combinations thereof, and the like, or wireless links, such as cellular links, satellite links, Wi-Fi links, and the like.

[0032] It should be noted that the number of requesters 2 and clusters 4 in the figure is only for illustration and is not intended to limit the scope of patent protection of this application. Depending on actual conditions, there may be any number of requesters 2 and clusters 4.

[0033] The cluster 4 may be a distributed computing cluster (such as a Ray cluster) and includes a plurality of nodes, including a head node and a plurality of working nodes.

[0034] The head node is configured with multiple queues, such as the first queue and the second queue. Different queues are used to store different data.

[0035] The following uses one of the head nodes (hereinafter referred to as the target head node) in cluster 4 as the execution subject to introduce the technical solution of the present application through multiple embodiments. It should be noted that these embodiments can be implemented in a variety of different forms and should not be interpreted as being limited to the embodiments described here.

[0036] Embodiment 1 Figure 2 The flowchart schematically shows a method for notifying a task execution result according to the first embodiment of the present application.

[0037] like Figure 2 As shown, the method may include steps S200 to S210, wherein: Step S200: receiving a request for a target task.

[0038] Step S202: taking out a target reference from the first queue, where the first queue is used to store multiple references.

[0039] Step S204: calling a target process of a target working node among the multiple working nodes through the target reference to execute the target task.

[0040] Step S206: Obtain target execution status information returned by the target working node based on the target task.

[0041] Step S208, adding the target tuple including the target execution status information and the target reference to the second queue to update the second queue; the second queue is used to store multiple tuples, and one tuple corresponds to one task.

[0042] Step S210, taking out a tuple from the updated second queue, and sending the execution result of the corresponding task to the requester based on the taken out tuple.

[0043] The notification method of the task execution result provided in this embodiment configures the first queue and the second queue in the head node. The task call of each working node is managed through the first queue, and the task execution status of the tasks in each working node is synchronized through the second queue. Based on the second queue (taking tuples from the second queue), the execution result of the corresponding task can be actively obtained and sent to the requester. In this way, when the working node process fails, the head node can feedback the execution result of the task to the requester based on the information in the second queue, thereby ensuring timely notification of the task execution result.

[0044] The following combination Figure 2 , each step in steps S200~S210 and other optional steps are explained in detail.

[0045] Step S200 , receiving the request for the target task.

[0046] The request for the target task may be initiated by the requester through a scheduling service. The scheduling service may refer to a service for managing and scheduling tasks or workloads in a distributed computing, containerized environment, or computing cluster.

[0047] The request may be a video encoding request, a video decoding request, a computing request, a reading request, a writing request, a transcoding request, etc.

[0048] Step S202 , taking out a target reference from the first queue, where the first queue is used to store multiple references.

[0049] Each reference may correspond to a task to be executed. The reference may be used to point to a process in a working node, and the reference may identify and access the process of the working node. The underlying communication mechanism of the working node where the process is located may be ignored through the reference to facilitate the scheduling of the process. In some embodiments, the reference may be an ActorHandle of a distributed computing framework, and the process of the working node may be an Actor process of the distributed computing framework. When a request for a target task is received, the ActorHandle at the head of the queue may be taken out from the first queue, and the ActorHandle may be used to interact with the Actor process of the working node.

[0050] Step S204 , through the target reference, calling the target process of the target working node among the multiple working nodes to execute the target task.

[0051] There are multiple processes in a worker node, and each process is pre-allocated with resources, such as computing power, memory, etc. Multiple processes in a worker node can execute tasks in parallel. A worker node can be a computing node in a cluster. A worker node can be deployed in a container, or in a physical server or virtual machine, which is not limited here. A container is a lightweight virtualization technology that packages an application and its required dependencies, libraries, environments, etc. into an independent unit. In some embodiments, the target head node and multiple worker nodes can be deployed in a pod of the Kubernetes platform and run as containers. Kubernetes is an open source platform for automating the deployment, expansion, and management of containerized applications. It is used to package applications into containers and distribute these containers to multiple physical or virtual computers to run, forming an automatically manageable cluster. Pod is the smallest deployment unit in Kubernetes, and each Pod represents one or more running containers. The target process can be an Actor process, through which the target task can be executed.

[0052] Step S206 , obtain the target execution status information returned by the target working node based on the target task.

[0053] The target execution status information may be status information related to the task generated and returned by the target working node during or after the task is completed. The target execution status information may be used to indicate the execution status of the task, and may also include whether the task is successfully completed, the execution result, the current progress, the error status, etc., which are not limited here. In some embodiments, the target ActorHandle reference may be used to interact with the target Actor process of the target working node, and an ObjectRef response value with the target execution status information may be returned through the Remote method. Remote is the core mechanism for implementing distributed computing in distributed computing. Remote distributes tasks to the Actor process in the distributed computing cluster through the gRPC (high-performance remote procedure call framework) protocol for execution.

[0054] Step S208 , adding the target tuple including the target execution status information and the target reference to the second queue to update the second queue; the second queue is used to store multiple tuples, and one tuple corresponds to one task.

[0055] The tuple may include execution status information and references, and may also include other data, which is not limited here. The reference in the tuple corresponds to the execution status information. By storing the tuples in the second queue, the execution status information of the task corresponding to each tuple can be independently maintained in the second queue. In some embodiments, the target reference and target execution status information included in the target tuple may be a target ActorHandle reference and a corresponding ObjectRef response value.

[0056] Step S210 , take out a tuple from the updated second queue, and send the execution result of the corresponding task to the requester based on the taken out tuple.

[0057] Based on the execution status information of the retrieved tuple, the execution task result corresponding to the execution status information can be sent to the requesting party through the callback function. The callback function is a function automatically called by the program when a certain event occurs or an operation is completed. The callback function can be passed as a parameter to other functions and executed by another function under certain conditions. In some embodiments, taking the X framework as an example, the execution result of the task can be obtained through X.get(ObjectRef). X.get() is a function in the framework used to synchronously obtain the result of the remote task.

[0058] The specific process of taking out tuples from the updated second queue and sending the execution results of the corresponding tasks to the requesting party will be exemplarily introduced below.

[0059] In an optional embodiment, if Figure 3 As shown, step S210 may include: Step S300, when a new request is received, query the updated second queue; Step S302: when the updated second queue is not an empty queue, taking out the tuple at the head of the queue from the updated second queue.

[0060] If the updated second queue is not an empty queue, it means that the execution results of the corresponding tasks have not been notified to the requester, and the tuple at the head of the queue can be taken out from the updated second queue. Based on the execution status information of the tuple, the execution result of the task corresponding to the execution status information can be sent to the requester through the callback function. If the updated second queue is empty, it means that the execution results of the tasks have been notified to the requester.

[0061] In this embodiment, tuples can be obtained from the second queue in a non-blocking manner, and the corresponding task execution result can be notified, that is, the processing of other tasks will not be affected by waiting for the execution result of the task. In addition, timely querying the second queue can realize timely notification of the task execution result.

[0062] In an optional embodiment, step S210 may further include: When a new tuple is added to the updated second queue, the tuple at the head of the queue is taken out from the updated second queue.

[0063] When a new tuple is added to the updated second queue, it means that a new task execution result has been obtained. When a new task is completed, the tuple at the head of the queue can be taken out from the updated second queue, and based on the execution status information of the tuple, the execution task result corresponding to the execution status information can be sent to the requesting party through the callback function. In some embodiments, it can be set when to take out the tuple from the updated second queue and notify the corresponding task result according to the actual situation. For example, it can be set that when a task execution result is received, that is, when a new tuple is added to the updated second queue, the tuple is taken out from the updated second queue and the corresponding task result is notified. It can also be set to notify the task result after receiving any number of task execution results.

[0064] For example, taking the X framework as an example, the number of task execution results to be waited for can be set through the parameters of the X.wait() function. In order to ensure timeliness, the number of waiting can be set to 1, that is, the task result notification will be carried out when there is a task execution result. X.wait() is a function in the framework that is used to wait for the completion of the remote computing task and return the reference of the completed task.

[0065] In this embodiment, when a new tuple is added, it can also trigger the timely notification of the task execution result to the requesting party.

[0066] Regarding the processing of the corresponding reference after the task execution result notification, in an optional embodiment, the method further includes: when sending the execution result of the corresponding task to the requesting party, re-adding the reference of the corresponding task to the first queue.

[0067] When the execution result of the task is completed and sent to the requester, it means that the corresponding reference has ended the interaction with the corresponding process, and the reference can be re-added to the first queue to facilitate recycling the reference. In some embodiments, the first queue can also be supplemented with references by creating references, etc., which is not limited here.

[0068] In this embodiment, by re-adding the corresponding references of the execution results of the completed tasks into the first queue, frequent creation and destruction of references is avoided, thereby improving resource utilization.

[0069] Regarding the processing when the target working node fails, in an optional embodiment, the method also includes: when the target execution status information includes error information, marking all tasks executed by the target working node as failed.

[0070] The tuple containing the execution status information of the task failure result is updated to the second queue, and the requesting party is notified of the corresponding task execution failure by taking out the tuple from the second queue. In some embodiments, taking the X framework as an example, the error type of the target working node can be recorded by the GCS built into the target head node, such as the error type NodeDead, indicating a node failure or death. At the same time, based on the execution status information, that is, ObjectRef, the exception is captured through X.get(ObjectRef). For example, the captured exception type is ActorError, indicating an Actor process error in the cluster. GCS is a core component in the cluster, running on the head node of the cluster, for storing and managing the global state and metadata of the cluster, and coordinating various operations of the cluster, such as task scheduling, resource allocation, etc.

[0071] In this embodiment, the target working node failure is promptly reported to the target head node, and the task failure result executed on the target working node is promptly fed back to the requester, so that the requester can handle the task failure in a timely manner.

[0072] Regarding the processing when the target working node and the target head node fail, in an optional embodiment, the method also includes: when both the target head node and the target working node fail, obtaining task information of multiple tasks executed by the target working node through the metadata service, and returning the failure results of the multiple tasks to the requesting party.

[0073] The metadata service can supplement the deficiency of the single-point record of the GCS in the target head node. The metadata service is used to manage the data and task information in the system. The metadata service will save metadata about the task execution status, storage location, etc., to help other parts of the system coordinate and schedule tasks. In some embodiments, the requester can query the execution status of the task from the task information it maintains. If the task fails, the line switching or retry logic can be executed. The line switching can be switching different clusters, and the retry logic can be reissuing a new task.

[0074] In this embodiment, the metadata service avoids the loss of tasks caused by the failure of both the target node and the target working node, thereby increasing the robustness of the system.

[0075] In order to make this application easier to understand, the following Figure 4 , 5 An exemplary application is provided.

[0076] like Figure 4 As shown, the resource pool may include multiple distributed computing clusters, each distributed computing cluster may have a head node and multiple working nodes, and the target head node may be the head node in the distributed computing cluster selected by the scheduling service.

[0077] The workflow of the target head node is as follows: S1. Receive the request for the target task, that is, receive the task submitted by the scheduling service; S2. Pop the target reference from the first queue (idle queue), where the first queue is used to store multiple references (ActorHandle); S3. Call the target process (Actor) of the target working node among multiple working nodes through the target reference to execute the target task; S4. Obtain the target execution status information returned by the target working node based on the target task, where the target pointing status information is the response value returned by the Remote; S5, including the target execution status information and the target reference target tuple, join (push) the second queue (busy queue) to update the second queue (busy queue); S6. Take out a tuple from the second queue (busy queue), and send the execution result of the corresponding task to the requester based on the taken out tuple.

[0078] S7. When the execution result of the corresponding task is sent to the requesting party, the reference of the corresponding task is re-added (pushed) to the first queue (idle queue).

[0079] S8. When the target execution status information includes error information, all tasks executed by the target work section are marked as failed.

[0080] S9. When both the target head node and the target working node fail, Figure 5 As shown, based on the interactive communication between the metadata service and the GSC deployed in the target head node, the task information of multiple tasks executed by the target working node is obtained through the metadata service, and the failure results of multiple tasks are returned to the requester.

[0081] In this exemplary application, the target head node can promptly notify the task execution results, and can actively notify in the event of a process failure or node failure, while avoiding misjudgment caused by passive detection timeout and resulting in waste of computing resources.

[0082] Embodiment 2 Figure 6A block diagram of a notification device for task execution results according to the second embodiment of the present application is schematically shown. The device can be divided into one or more program modules, one or more program modules are stored in a storage medium, and are executed by one or more processors to complete the embodiment of the present application. The program module referred to in the embodiment of the present application refers to a series of computer program instruction segments that can perform specific functions. The following description will specifically introduce the functions of each program module in this embodiment. The device is used in a target head node of a cluster, and the cluster also includes multiple working nodes. The target head node is configured with a first queue and a second queue. Figure 6 As shown, the device 1600 may include: a receiving module 1610, a first fetching module 1620, a calling module 1630, an acquiring module 1640, an updating module 1650, and a second fetching module 1660, wherein: The receiving module 1610 is used to receive a request for a target task; A first fetching module 1620, configured to fetch a target reference from a first queue, wherein the first queue is configured to store a plurality of references; A calling module 1630 is used to call a target process of a target working node among the multiple working nodes through the target reference to execute the target task; An acquisition module 1640 is used to acquire target execution status information returned by the target working node based on the target task; An updating module 1650 is used to add a target tuple including the target execution state information and the target reference to the second queue to update the second queue; the second queue is used to store multiple tuples, and one tuple corresponds to one task; The second fetching module 1660 is used to fetch tuples from the updated second queue, and send the execution result of the corresponding task to the requesting party based on the fetched tuples.

[0083] As an optional embodiment, the second taking out module 1660 is further used for: In case of receiving a new request, querying the updated second queue; When the updated second queue is not an empty queue, the tuple at the head of the queue is taken out from the updated second queue.

[0084] As an optional embodiment, the second taking out module 1660 is further used for: When a new tuple is added to the updated second queue, the tuple at the head of the queue is taken out from the updated second queue.

[0085] As an optional embodiment, the apparatus 1600 further includes a joining module, configured to: When the execution result of the corresponding task is sent to the requester, the reference of the corresponding task is added back to the first queue.

[0086] As an optional embodiment, the apparatus 1600 further includes a marking module, which is used to: In the case where the target execution status information includes error information, all tasks executed by the target work section are marked as failed.

[0087] As an optional embodiment, the apparatus 1600 further includes a returning module, configured to: In the case that both the target head node and the target working node fail, task information of multiple tasks executed by the target working node is obtained through a metadata service, and failure results of the multiple tasks are returned to the requesting party.

[0088] Embodiment 3 Figure 7 The hardware architecture diagram of the computer device 10000 suitable for implementing the notification method of the task execution result according to the third embodiment of the present application is schematically shown. In some embodiments, the computer device 10000 can be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 can be a rack server, a blade server, a tower server or a cabinet server (including an independent server, or a server cluster composed of multiple servers), etc. Figure 7 As shown, the computer device 10000 includes but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can communicate with each other through a system bus. Among them: The memory 10010 includes at least one type of computer-readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 10010 can be an internal storage module of the computer device 10000, such as a hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 can also be an external storage device of the computer device 10000, such as a plug-in hard disk equipped on the computer device 10000, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Of course, the memory 10010 can also include both the internal storage module of the computer device 10000 and its external storage device. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed in the computer device 10000, such as the program code of the notification method of the task execution result, etc. In addition, the memory 10010 can also be used to temporarily store various data that have been output or will be output.

[0089] In some embodiments, the processor 10020 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other chips. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.

[0090] The network interface 10030 may include a wireless network interface or a wired network interface, and the network interface 10030 is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal through a network, and to establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an intranet, the Internet, the Global System of Mobile communication (GSM), Wideband Code Division Multiple Access (WCDMA), 4G network, 5G network, Bluetooth, Wi-Fi, etc.

[0091] It should be pointed out that Figure 7 Only a computer device having components 10010 - 10030 is shown, but it should be understood that implementation of all of the components shown is not a requirement, and more or fewer components may alternatively be implemented.

[0092] In this embodiment, the notification method of the task execution result stored in the memory 10010 can also be divided into one or more program modules and executed by one or more processors (such as processor 10020) to complete the embodiment of the present application.

[0093] Embodiment 4 An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the method for notifying the task execution result in the embodiment are implemented.

[0094] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (for example, SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the computer-readable storage medium can be an internal storage unit of a computer device, such as a hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium can also be an external storage device of a computer device, such as a plug-in hard disk equipped on the computer device, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Of course, the computer-readable storage medium can also include both the internal storage unit of the computer device and its external storage device. In this embodiment, the computer-readable storage medium is generally used to store the operating system and various application software installed on the computer device, such as the program code of the notification method of the task execution result in the embodiment. In addition, the computer-readable storage medium can also be used to temporarily store various types of data that have been output or are to be output.

[0095] Embodiment 5 An embodiment of the present application also provides a computer program product, including a computer program, which implements the method in the above embodiment when executed by a processor.

[0096] Obviously, those skilled in the art should understand that the modules or steps of the above-mentioned embodiments of the present application can be implemented by general-purpose computer devices, they can be concentrated on a single computer device, or distributed on a network composed of multiple computer devices, optionally, they can be implemented by executable program codes of computer devices, so that they can be stored in a storage device and executed by the computer device, and in some cases, the steps shown or described can be executed in a different order from that herein, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0097] It should be noted that the above are only preferred embodiments of the present application, and the patent protection scope of the present application is not limited thereto. Any equivalent structure or equivalent process transformation made using the contents of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A method for notifying a task execution result, characterized in that: In a target head node for a cluster, the cluster further includes a plurality of working nodes, and the target head node is configured with a first queue and a second queue; the method includes: Receive a request for a target task; Retrieving a target reference from the first queue, the first queue being used to store a plurality of references; Calling a target process of a target working node among the multiple working nodes through the target reference to execute the target task; Obtaining target execution status information returned by the target working node based on the target task; adding a target tuple including the target execution state information and the target reference to the second queue to update the second queue; the second queue is used to store multiple tuples, one tuple corresponding to one task; A tuple is taken out from the updated second queue, and based on the taken out tuple, the execution result of the corresponding task is sent to the requesting party.

2. The method according to claim 1, characterized in that: Taking out a tuple from the updated second queue, and sending the execution result of the corresponding task to the requester based on the taken out tuple, including: In case of receiving a new request, querying the updated second queue; When the updated second queue is not an empty queue, the tuple at the head of the queue is taken out from the updated second queue.

3. The method according to claim 1, characterized in that Taking out a tuple from the updated second queue, and sending the execution result of the corresponding task to the requester based on the taken out tuple, including: When a new tuple is added to the updated second queue, the tuple at the head of the queue is taken out from the updated second queue.

4. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: When the execution result of the corresponding task is sent to the requester, the reference of the corresponding task is added back to the first queue.

5. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: In the case where the target execution status information includes error information, all tasks executed by the target work section are marked as failed.

6. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: In the case that both the target head node and the target working node fail, task information of multiple tasks executed by the target working node is obtained through a metadata service, and failure results of the multiple tasks are returned to the requesting party.

7. A notification device for task execution results, characterized in that: In a target head node for a cluster, the cluster further includes a plurality of working nodes, and the target head node is configured with a first queue and a second queue; the device includes: A receiving module, used for receiving a request for a target task; A first fetching module, used for fetching a target reference from a first queue, wherein the first queue is used for storing a plurality of references; A calling module, used for calling a target process of a target working node among the multiple working nodes through the target reference to execute the target task; An acquisition module is used to acquire target execution status information returned by the target working node based on the target task; An updating module, used for adding a target tuple including the target execution state information and the target reference to the second queue to update the second queue; the second queue is used for storing a plurality of tuples, one of the tuples corresponding to one task; The second fetching module is used to fetch tuples from the updated second queue, and send the execution result of the corresponding task to the requesting party based on the fetched tuples.

8. A computer device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to claims 1 to 6 are implemented.