Task data processing method and device, computer, storage medium and program product
Through the coordinated work of scheduling service nodes and execution service nodes, task information acquisition and execution are decoupled, and task metadata and instance state management are used to solve the problem of low efficiency and security in task data processing, and efficient and secure task data processing is achieved.
Patent Information
- Application Number
- CN202510400700.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art is prone to task loss when processing large amounts of task data, resulting in low efficiency and security of task data processing.
Through the coordinated work of scheduling service nodes and execution service nodes, the acquisition and execution of task information is decoupled, and task metadata and instance state management are used to achieve efficient processing and secure storage of task data.
It improves the efficiency and security of task data processing, reduces the loss of task information, and can still effectively schedule and process task information when the amount of large data is large.
Smart Images

Figure CN120335959A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technologies, and in particular, to a method, apparatus, computer, storage medium, and program product for processing task data. Background Art
[0002] Executing tasks over the Internet has become a relatively common scenario. As the amount of task data increases, the processing of tasks becomes extremely important. Currently, generally, individual tasks are calculated and processed through a streaming framework, and it is impossible to wait for a certain task to complete for a long time. This is suitable for real-time processing of data itself. However, when the amount of task data is too large, task loss may occur, resulting in poor security for task data processing. Alternatively, a framework for processing scheduling tasks for operations initiated by users is used to process tasks. This framework provides a synchronous interface externally, making it impossible to process too many tasks simultaneously, which may lead to problems such as task loss, thus resulting in low efficiency and security for task data processing. Summary of the Invention
[0003] Embodiments of the present application provide a method, apparatus, computer, storage medium, and program product for processing task data, which can improve the efficiency and security of task data processing.
[0004] On the one hand, an embodiment of the present application provides a method for processing task data. This method is executed by a scheduling service node and includes:
[0005] Obtain first task information, convert the first task information into first task metadata belonging to a task data format, and add the first task metadata to a first task database; the task data format refers to the format of data provided by the scheduling service node for an execution service node;
[0006] Set the instance status of a first task instance corresponding to a first message task to an unexecuted status, and generate first instance data of the first task instance based on the instance status of the first task instance; the first message task refers to the message task indicated by the first task information.
[0007] When the scheduling service node issues tasks, based on the instance status of the instance data in the first task database, obtain the instance data to be processed from the instance data in the first task database, and send the instance data to be processed and the task metadata associated with the instance data to be processed to the execution service node, so that the execution service node processes the instance data to be processed and the task metadata associated with the instance data to be processed, and feeds back the processing status result for the instance data to be processed and the task metadata associated with the instance data to be processed to the scheduling service node; the instance data in the first task database includes first instance data, and the instance data to be processed includes instance data with an instance status of unexecuted status; the processing status result is used to update the instance status in the instance data to be processed.
[0008] On the one hand, an embodiment of the present application provides a task data processing method, which is executed by an execution service node. The method includes:
[0009] Receive the instance data to be processed and the task metadata associated with the instance data to be processed sent by the scheduling service node, process the instance data to be processed and the task metadata associated with the instance data to be processed, and obtain the processing status result for the instance data to be processed and the task metadata associated with the instance data to be processed; the instance data to be processed is generated based on the instance status of the to-be-processed task instances in the task metadata associated with the instance data to be processed, and the to-be-processed task instances refer to the task instances indicated by the instance data to be processed; the initial value of the instance status of the to-be-processed task instances is the unexecuted status, and the to-be-processed task instances include instance data with an instance status of unexecuted status; the task metadata associated with the instance data to be processed is data belonging to the task data format converted from the task information, and the task metadata associated with the instance data to be processed is added to the first task database by the scheduling service node; the task data format refers to the format of the data provided by the scheduling service node for the execution service node;
[0010] Feed back the processing status result to the scheduling service node, so that the scheduling service node updates the instance status in the instance data to be processed based on the processing status result.
[0011] On the one hand, an embodiment of the present application provides a task data processing device, which is applicable to a scheduling service node. The device includes:
[0012] A task processing module, configured to obtain first task information, convert the first task information into first task metadata belonging to the task data format, and add the first task metadata to the first task database; the task data format refers to the format of the data provided by the scheduling service node for the execution service node;
[0013] An instance processing module, configured to set the instance status of a first task instance corresponding to a first message task to an unexecuted status, and generate first instance data of the first task instance based on the instance status of the first task instance; the first message task refers to the message task indicated by the first task information.
[0014] A task distribution module, configured to, when the scheduling service node distributes tasks, obtain to-be-processed instance data from the instance data in the first task database based on the instance status of the instance data in the first task database, and send the to-be-processed instance data and the task metadata associated with the to-be-processed instance data to the execution service node, so that the execution service node processes the to-be-processed instance data and the task metadata associated with the to-be-processed instance data, and feeds back a processing status result for the to-be-processed instance data and the task metadata associated with the to-be-processed instance data to the scheduling service node; the instance data in the first task database includes the first instance data, and the to-be-processed instance data includes the instance data with an unexecuted status; the processing status result is used to update the instance status in the to-be-processed instance data.
[0015] Wherein, when obtaining the first task information, the task processing module can be used to:
[0016] Scan the first message queue or the second task database. If the first message queue or the second task database is not empty, obtain the first task information from the first message queue or the second task database; the first task information is added to the first message queue or the second task database by the application service node; the first task information is obtained by the application service node through converting an initial task message; the data format of the first task information is a scheduling data format, and the scheduling data format is the data format supported by the scheduling service node, which refers to the data format provided by the application service node for the scheduling service node.
[0017] Wherein, the scheduling service node belongs to a scheduling node cluster. When obtaining the first task information, the task processing module can be used to:
[0018] Obtain the to-be-scheduled task information and the task hash value corresponding to the to-be-scheduled task information, and convert the task hash value into a data shard value based on the number of scheduling nodes; the number of scheduling nodes is the number of nodes included in the scheduling node cluster.
[0019] Obtain the first shard processing range of the local scheduling service node, and determine the to-be-scheduled task information whose data shard value belongs to the first shard processing range as the first task information.
[0020] Obtain the first task information.
[0021] Wherein, the device further includes:
[0022] A node processing module, which is used to send a node offline message to other scheduling service nodes when a scheduling service node goes offline, so that other scheduling service nodes update the shard processing range of other scheduling service nodes based on the first shard processing range to obtain a second shard processing range; other scheduling service nodes refer to nodes other than the scheduling service node in the scheduling node cluster; the second shard processing range is used to represent the range of message tasks processed by other scheduling service nodes.
[0023] Wherein, when converting the first task information into the first task metadata belonging to the task data format, the task processing module can be used for:
[0024] Perform type recognition on the first task information to obtain the first task type indicated by the first task information;
[0025] Obtain the data parameters corresponding to the task data format, and obtain parameter-associated data from the first task information based on the data parameters;
[0026] Based on the data parameters, form the first task metadata by combining the first task type and the parameter-associated data.
[0027] Wherein, when converting the first task information into the first task metadata belonging to the task data format, the task processing module can be used for:
[0028] Convert the first task information into initial task metadata belonging to the task data format;
[0029] Obtain the priority determination parameter in the initial task metadata, and determine the task priority corresponding to the first task information based on the priority determination parameter; the priority determination parameter refers to the parameter used for priority determination;
[0030] Combine the task priority and the initial task metadata to form the first task metadata.
[0031] Wherein, when a scheduling service node issues a task and obtains the to-be-processed instance data from the instance data in the first task database based on the instance status of the instance data in the first task database, the task issuing module can be used for:
[0032] When a scheduling service node issues a task, obtain the instance status of the instance data included in the first task database, and determine the instance data with the instance status of unexecuted status or execution exception status as the initial instance data;
[0033] Obtain the number of valid execution nodes of the valid execution nodes with task execution conditions in the execution service node, and obtain the task execution data volume of the valid execution nodes;
[0034] Determine the number of instances to be dispatched based on the number of valid nodes and the amount of task execution data, and obtain the instance data to be processed from the initial instance data based on the number of instances to be dispatched.
[0035] Wherein, the device further includes:
[0036] A status update module, configured to update the instance status in the instance data to be processed to the in-execution status after sending the instance data to be processed and the task metadata associated with the instance data to be processed to the execution service node;
[0037] The status update module is further configured to update the instance status in the instance data to be processed based on the processing status result when receiving the processing status result feedback from the execution service node.
[0038] Wherein, the number of execution service nodes is M, and M is a positive integer; when sending the instance data to be processed and the task metadata associated with the instance data to be processed to the execution service node, the task dispatching module can be used for:
[0039] Determine a target execution service node from the M execution service nodes based on the node load and node status of the M execution service nodes; the node status of the target execution service node is the node running status;
[0040] Send the instance data to be processed and the task metadata associated with the instance data to be processed to the target execution service node.
[0041] An embodiment of the present application provides a task data processing device on the one hand. The device is applicable to an execution service node, and the device includes:
[0042] A data processing module, configured to receive the instance data to be processed and the task metadata associated with the instance data to be processed sent by the scheduling service node, process the instance data to be processed and the task metadata associated with the instance data to be processed, and obtain a processing status result for the instance data to be processed and the task metadata associated with the instance data to be processed; the instance data to be processed is generated based on the instance status of the to-be-processed task instances in the task metadata associated with the instance data to be processed, and the to-be-processed task instances refer to the task instances indicated by the instance data to be processed; the initial value of the instance status of the to-be-processed task instances is the unexecuted status, and the to-be-processed task instances include instance data with an instance status of unexecuted; the task metadata associated with the instance data to be processed is data belonging to the task data format converted from task information, and the task metadata associated with the instance data to be processed is added to the first task database by the scheduling service node; the task data format refers to the format of the data provided by the scheduling service node for the execution service node;
[0043] A result feedback module, configured to feedback the processing status result to the scheduling service node, so that the scheduling service node updates the instance status in the to-be-processed instance data based on the processing status result.
[0044] Wherein, when processing the to-be-processed instance data and the task metadata associated with the to-be-processed instance data to obtain the processing status result of the to-be-processed instance data and the task metadata associated with the to-be-processed instance data, the data processing module may be used for:
[0045] Obtain the task type, data to be updated, and task association path from the to-be-processed instance data and the task metadata associated with the to-be-processed instance data;
[0046] Determine the task logic corresponding to the to-be-processed instance data based on the task type, and generate an initial task instruction according to the task logic;
[0047] Update the initial task instruction with the data to be updated and the task association path to obtain a to-be-processed instruction, execute the to-be-processed instruction, and when the execution of the to-be-processed instruction is completed, obtain the processing status result of the to-be-processed instruction.
[0048] Wherein, the number of to-be-processed instance data is H, and H is a positive integer;
[0049] When processing the to-be-processed instance data and the task metadata associated with the to-be-processed instance data to obtain the processing status result of the to-be-processed instance data and the task metadata associated with the to-be-processed instance data, the data processing module may be used for:
[0050] Obtain the first priority corresponding to each of the H to-be-processed instance data from the task metadata respectively associated with the H to-be-processed instance data;
[0051] Obtain the second priority corresponding to each of the H to-be-processed instance data from the H to-be-processed instance data;
[0052] Determine the instance processing order of the H to-be-processed instance data based on the first priority and the second priority corresponding to each of the H to-be-processed instance data;
[0053] Process the H to-be-processed instance data in sequence according to the instance processing order, and when the processing of each to-be-processed instance data is completed, obtain the processing status result corresponding to the to-be-processed instance data.
[0054] An embodiment of the present application provides a computer device on the one hand, including a processor, a memory, and an input / output interface;
[0055] The processor is respectively connected to the memory and the input / output interface. Among them, the input / output interface is used to receive and output data, the memory is used to store computer programs, and the processor is used to call the computer programs so that the computer device including the processor executes the task data processing method in one aspect of the embodiments of the present application.
[0056] One aspect of the embodiments of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded and executed by a processor so that a computer device having the processor executes the task data processing method in one aspect of the embodiments of the present application.
[0057] One aspect of the embodiments of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions so that the computer device executes the methods provided in various alternative manners in one aspect of the embodiments of the present application. In other words, when the computer instructions are executed by the processor, the methods provided in various alternative manners in one aspect of the embodiments of the present application are implemented.
[0058] Implementing the embodiments of the present application will have the following beneficial effects:
[0059] In an embodiment of the present application, first task information is obtained, the first task information is converted into first task metadata belonging to a task data format, and the first task metadata is added to a first task database; the instance status of the first task instance corresponding to the first message task is set to an unexecuted status, and first instance data of the first task instance is generated based on the instance status of the first task instance; the first message task refers to the message task indicated by the first task information; when a scheduling service node issues a task, based on the instance status of the instance data in the first task database, the to-be-processed instance data is obtained from the instance data in the first task database, and the to-be-processed instance data and the task metadata associated with the to-be-processed instance data are sent to an execution service node, so that the execution service node processes the to-be-processed instance data and the task metadata associated with the to-be-processed instance data, and feeds back a processing status result for the to-be-processed instance data and the task metadata associated with the to-be-processed instance data to the scheduling service node; the instance data in the first task database includes the first instance data, and the to-be-processed instance data includes the instance data with an unexecuted status; the processing status result is used to update the instance status in the to-be-processed instance data. By combining the scheduling service node and the execution service node, the acquisition of task information and the execution of message tasks are decoupled, so as to implement the processing of tasks through different service nodes, enabling the scheduling service node to obtain and schedule task information, and the execution service node to process the message tasks indicated by the task information. The scheduling service node only obtains and schedules task information, which only consumes less resources and time. Even when the data volume of the message task is large, the scheduling service node can meet the requirements for resources such as the acquisition and scheduling of task information of the message task, thereby reducing the situation of task information loss and improving the security of task data processing. At the same time, multiple service nodes coordinate to work, and the instance status is introduced, enabling the scheduling service node to more conveniently obtain the instance data to be processed (i.e., the to-be-processed instance data), thereby improving the efficiency of task data processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0061] Figure 1 is a network interaction architecture diagram for task data processing provided by an embodiment of the present application;
[0062] Figure 2 is another network interaction architecture diagram for task data processing provided by an embodiment of the present application;
[0063] Figure 3 It is a schematic diagram of a task data processing scenario provided by an embodiment of the present application;
[0064] Figure 4 It is a flowchart of a method for task data processing provided by an embodiment of the present application;
[0065] Figure 5 It is another flowchart of a method for task data processing provided by an embodiment of the present application;
[0066] Figure 6 It is a schematic diagram of a task data processing interaction scenario provided by an embodiment of the present application;
[0067] Figure 7 It is a flowchart of a method interaction for task data processing provided by an embodiment of the present application;
[0068] Figure 8 It is a schematic diagram of a data display scenario provided by an embodiment of the present application;
[0069] Figure 9 It is a schematic diagram of a data synchronization event scenario provided by an embodiment of the present application;
[0070] Figure 10 It is a schematic diagram of a task data processing device provided by an embodiment of the present application;
[0071] Figure 11 It is another schematic diagram of a task data processing device provided by an embodiment of the present application;
[0072] Figure 12 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0073] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.
[0074] Among them, if it is necessary to collect data of an object (such as a user, etc.) in this application, a prompt interface or a pop-up window is displayed before and during the collection. The prompt interface or the pop-up window is used to prompt the user that some data is being collected currently. Only after obtaining the user's confirmation operation on the prompt interface or the pop-up window, the relevant steps for data acquisition are started; otherwise, the process ends. Moreover, the obtained user data will be used in reasonable, legal scenarios or for legal purposes, etc. Optionally, in some scenarios where user data needs to be used but the user's authorization has not been obtained, authorization can also be requested from the user, and the user data will be used only when the authorization is passed. That is to say, the use of user data in this application complies with the relevant regulations of laws and regulations.
[0075] In the embodiment of the present application, please refer to Figure 1 , Figure 1 which is a network interaction architecture diagram for task data processing provided by the embodiment of the present application, as Figure 1As shown in the figure, the task data processing architecture in this application includes, but is not limited to, a second message queue 101, an application service node 102, a scheduling service node 103, and an execution service node 104. Among them, the second message queue 101 is used to store the event information of the events generated by the business application. The business application can be any application program that can generate events, such as a communication application, a video application, or a social application, etc., which is not limited here. The application service node 102 refers to a node used to process the data generated by the business application, and is used to obtain the event information from the second message queue 101 and convert the event information into task information supported by the scheduling service node 103. Among them, scheduling refers to the process of arranging, directing, managing, or coordinating a certain work or activity to ensure its smooth progress or achieve the expected goal. As the name implies, the scheduling service node 103 refers to a node used to implement the scheduling of the data related to the message task, that is, used to store the data related to the message task and determine the execution service node that processes the data. Specifically, the scheduling service node 103 can be used to obtain the task information and perform scheduling processing on the message task indicated by the task information and the task instances included in the message task. That is, the message task and the task instances included in the message task are sent to the execution service node 104; the execution service node 104 refers to a node that processes the message task and is used to process the obtained task instances. For example, the number of execution service nodes 104 is multiple. When the scheduling service node 103 performs task distribution, it can determine the message task or task instance to be distributed, and which execution service node to send the determined message task or task instance to. Through the cooperation of each service node (such as the application service node 102, the scheduling service node 103, and the execution service node 104, etc.), the acquisition and processing of the message task are realized. During the acquisition and processing of the message task, through different service nodes, the decoupling between the acquisition of the message task and the processing of the message task is realized, so that the influence between the two is weak. When the scheduling service node maintains the task information, it does not need to wait for the message task corresponding to the task information to be executed, thereby improving the efficiency of task data processing. And even when the data volume of the task information is large, the scheduling service node can also realize the storage of the task information, thereby improving the security of task data processing.
[0076] Among them, the corresponding numbers of the above-mentioned second message queue 101, application service node 102, scheduling service node 103, and execution service node 104 can all be non-unique. For example, referring to Figure 2 , Figure 2 is another network interaction architecture diagram of task data processing provided by the embodiment of this application. As Figure 2 shown, the number of the second message queue 201 can be denoted as A, and A is a positive integer. For example, Figure 2The second message queues 201a, 201b, etc. shown in [the figure]; the number of application service nodes 202 can be denoted as B, where B is a positive integer, such as Figure 2 The application service nodes 202a, 202b, etc. shown in [the figure]; the number of scheduling service nodes can be denoted as N, where N is a positive integer, such as Figure 2 The scheduling service nodes 203a, 203b, etc. shown in [the figure]. It can be considered that N scheduling service nodes form a scheduling node cluster 203; the number of execution service nodes 204 can be denoted as M, where M is a positive integer, such as Figure 2 The execution service nodes 204a, 204b, etc. shown in [the figure].
[0077] Specifically, when a business application generates event information, the event information can be written into the second message queue 201. Among them, the number of the second message queue 201 can be A, where A is a positive integer, such as Figure 2The second message queues 201a, 201b, etc. shown in [the figure]. Optionally, the service devices associated with the service application can obtain the queue data parameters corresponding to A second message queues, and write the event information into the second message queue indicated by the queue data parameter that matches the event information. Here, the queue data parameter is used to represent the relevant information of the data stored in the corresponding second message queue. That is to say, the queue data parameter is used to represent what data the corresponding second message queue is used to store. Through this method, the classified storage of event information can be realized, and the organization of data management can be improved. Or, the service devices associated with the service application can randomly write the event information into one or more second message queues, and there is no limitation here. Further, the application service node 202 can obtain the event information from the second message queue 201, and convert the event information into task information that can be processed by the scheduling service node 203. Among them, any one of the application service nodes can scan the second message queue 201 to obtain the event information from the second message queue 201. Any one of the scheduling service nodes in the scheduling node cluster 203 can obtain the task information generated by the application service node, store and schedule the task information. Specifically, the obtained task information is converted into task metadata that can be processed by the execution service node and its associated instance data. Among them, the scheduling service node can be regarded as a distributed node, integrating a distributed scheduling engine for distributed processing of task information. Any one of the execution service nodes in the execution service node 204 can process the obtained task and obtain the processing status result of the task. Through multiple service nodes, and the number of each type of service node is multiple, the process of task data processing is split, so that when the scheduling service node obtains the task information, it does not need to wait for the task information processing to be completed, but stores the obtained task information, and then continuously sends the stored task information to the execution service node. This also enables the scheduling service node to store and schedule the task information when the data volume of the task information is large, thereby improving the efficiency and security of task data processing.
[0078] Specifically, reference can be made to Figure 3 , Figure 3 which is a schematic diagram of a task data processing scenario provided by an embodiment of the present application. As Figure 3As shown, the scheduling service node 301 can obtain the first task information 302, convert the first task information 302 into the first task metadata 303 belonging to the task data format, and add the first task metadata 303 to the first task database 304. Herein, the task data format refers to the data format supported by the scheduling service node 301, that is, the format of the data provided by the scheduling service node 301 for the execution service node, and the data belonging to this task data format can be processed by the execution service node. The scheduling service node 301 can set the instance state of the first task instance corresponding to the first message task to the unexecuted state, and this unexecuted state is used to indicate that the corresponding instance data has not been processed yet. The first message task refers to the message task indicated by the first task information. Further, the scheduling service node 301 can generate the first instance data of the first task instance based on the instance state of the first task instance. Among them, for the generation process of any task metadata added by the scheduling service node 301 to the first task database 304, the generation process of the first task metadata can be referred to. For the generation process of the instance data associated with any task metadata, the generation process of the first instance data can be referred to, and no further description will be given here. The scheduling service node 301 can, according to the instance state of the instance data in the first task database 304, obtain the instance data to be processed from the instance data in the first task database 304, and send the instance data to be processed and the task metadata associated with the instance data to be processed to the execution service node 305. The execution service node 305 can process the instance data to be processed and the task metadata associated with the instance data to be processed, and obtain the processing status result for the instance data to be processed and the task metadata associated with the instance data to be processed, and feedback the processing status result to the scheduling service node 301. The scheduling service node 301 can update the instance state in the instance data to be processed based on the processing status result. By this means, a task data processing architecture composed of multiple service nodes is constructed, and different service nodes implement different steps in the task data processing process, so that different task information will not have an impact during the task data processing process. Thus, even when the data volume of the task information is large, the task information can be processed, the situation of data loss is reduced, and the efficiency and security of task data processing are improved.
[0079] It can be understood that each service node mentioned in the embodiments of the present application (such as an application service node, a scheduling service node, or an execution service node) may be a type of computer device. The computer devices in the embodiments of the present application include, but are not limited to, terminal devices or servers. In other words, the computer device may be a server or a terminal device, or a system composed of a server and a terminal device. Among them, the above-mentioned terminal device may be an electronic device, including but not limited to mobile phones, tablet computers, desktop computers, laptop computers, handheld computers, in-vehicle devices, augmented reality / virtual reality (AR / VR) devices, head-mounted displays, smart TVs, wearable devices, smart speakers, digital cameras, cameras, and other mobile internet devices (MIDs) with network access capabilities, or terminal devices in scenarios such as trains, ships, and flights. Among them, the above-mentioned server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, vehicle-road collaboration, content delivery network (CDN), and big data and artificial intelligence platforms.
[0080] Optionally, the data involved in the embodiments of the present application may be stored in a computer device, or the data may be stored based on cloud storage technology or a blockchain network, which is not limited herein.
[0081] Further, please refer to Figure 4 , Figure 4 which is a flowchart of a method for task data processing provided by the embodiments of the present application. As Figure 4 shown, this task data processing process is implemented by a scheduling service node, and this task data processing process includes the following steps:
[0082] Step S401, obtain the first task information, convert the first task information into first task metadata belonging to the task data format, and add the first task metadata to the first task database.
[0083] In the embodiment of the present application, the scheduling service node (MasterServer) can obtain the first task information. Specifically, the scheduling service node can obtain the first task information from the first message queue or the second task database. Among them, the storage space of the first task information (that is, obtaining the first task information from the first message queue or the second task database) is determined jointly by the scheduling service node and the application service node. Specifically, when storing task information through the first message queue, the scheduling service node can scan the first message queue to obtain the first task information from the first message queue. Or, when storing task information through the second task database, the scheduling service node can scan the second task database to obtain the first task information from the second task database. Or, the scheduling service node can scan the first message queue or the second task database; if the first message queue or the second task database is not empty, obtain the first task information from the first message queue or the second task database; the first task information is added to the first message queue or the second task database by the application service node (Application Programming Interface Server, ApiServer); the first task information is obtained by the application service node converting the initial task message; the data format of the first task information is the scheduling data format, and the scheduling data format is the data format supported by the scheduling service node, that is, the data format required by the scheduling service node, which refers to the data format provided by the application service node for the scheduling service node. That is to say, the scheduling service node can obtain the first task information from the storage space determined by the scheduling service node and the application service node.
[0084] Optionally, the scheduling service node belongs to a scheduling node cluster. The number of nodes included in the scheduling node cluster is N, where N is a positive integer. The scheduling service node can be any one of the nodes in the scheduling node cluster. When describing with one node in the scheduling node cluster, this node can be denoted as the scheduling service node (or also referred to as the local scheduling service node), and the nodes in the scheduling node cluster other than this scheduling service node are denoted as other scheduling service nodes. Specifically, when obtaining the first task information, the scheduling service node can directly obtain the first task information. That is to say, each scheduling service node can obtain the task information generated by the application service node, and the task information obtained by the scheduling service node as the execution subject can be denoted as the first task information. Simply put, each scheduling service node can continuously scan the storage space of the task information and obtain the unprocessed task information from the storage space. By deploying multiple scheduling service nodes, a large amount of data can be processed simultaneously, achieving high-concurrency processing of task information, thereby improving the efficiency of task data processing. Among them, a single scheduling service node can easily schedule tens of millions of message tasks or task instances per day, and multiple scheduling service nodes can easily schedule hundreds of millions of message tasks or task instances per day. Thus, when this application is applied to scenarios of processing a large number of task data, it can also easily achieve the acquisition and scheduling of task information. Even with the performance optimization of the scheduling service nodes, the task magnitude that this application can handle will be higher, improving the efficiency and high concurrency of task data processing.
[0085] Specifically, the i-th scheduling service node obtains the first task information that is not associated with the read flag from the storage space of the task information. When the application service node detects that the first task information is obtained, it adds a read flag to the first task information; or, the i-th scheduling service node obtains the first task information from the storage space of the task information, and when the application service node detects that the first task information is obtained, it deletes the first task information from the storage space. Here, i is a positive integer less than or equal to N.
[0086] Alternatively, each scheduling service node can be responsible for a part of the data respectively. That is to say, the scheduling service node can obtain the first task information that meets the task acquisition conditions of the scheduling service node from the storage space of the task information. Specifically, the scheduling service node can obtain the to-be-scheduled task information and the task hash value corresponding to the to-be-scheduled task information, obtain the first shard processing range of the scheduling service node, and determine the first task information from the to-be-scheduled task information based on the task hash value and the first shard processing range; and obtain the first task information. Among them, the first shard processing range can be a hash range. At this time, the scheduling service node can determine the to-be-scheduled task information whose task hash value belongs to the first shard processing range as the first task information. Or, the first shard processing range can be a shard value range. At this time, the scheduling service node can obtain the to-be-scheduled task information and the task hash value corresponding to the to-be-scheduled task information, and convert the task hash value into a data shard value based on the number of scheduling nodes; the number of scheduling nodes is the number of nodes included in the scheduling node cluster, that is, N. For example, the remainder of the task hash value divided by the number of scheduling nodes can be used to obtain the data shard value of the to-be-scheduled task information, etc. The scheduling service node can obtain the first shard processing range of the scheduling service node, determine the to-be-scheduled task information whose data shard value belongs to the first shard processing range as the first task information; and obtain the first task information.
[0087] Optionally, when the scheduling service node goes offline, the scheduling service node may send a node offline message to other scheduling service nodes, so that the other scheduling service nodes update their shard processing ranges based on the first shard processing range to obtain a second shard processing range; the other scheduling service nodes refer to the nodes in the scheduling node cluster other than the scheduling service node; the second shard processing range is used to represent the range of message tasks processed by the other scheduling service nodes. For example, the first shard processing range may be divided among the other scheduling service nodes, that is, the second shard processing range corresponding to the other scheduling service nodes includes the shard processing range of the other scheduling service node before the update and the range allocated from the first shard processing range. For example, based on the number of other scheduling service nodes, the first shard processing range may be divided into K sub-shard ranges, and the K sub-shard ranges are allocated to the other scheduling service nodes. The j-th scheduling service node among the other scheduling service nodes combines the shard processing range of the j-th scheduling service node with the sub-shard range allocated to the j-th scheduling service node to form the second shard processing range of the j-th scheduling service node, where K is the number of nodes, K is a positive integer, and j is a positive integer less than or equal to K. Alternatively, based on the load of the other scheduling service nodes, the first shard processing range may be allocated to the other scheduling service nodes to obtain the second shard processing range of the other scheduling service nodes, that is, load balancing processing may be performed on the other scheduling service nodes, etc. That is to say, when deploying multiple scheduling service nodes, when a certain scheduling service node goes offline, the other scheduling service nodes can take over the task data processing range (i.e., the above-mentioned shard processing range) responsible by the offline scheduling service node, so that when some scheduling service nodes go offline (such as actively going offline or having an exception, etc.), it will not affect the acquisition and scheduling of task information, thereby improving the security of task data processing.
[0088] Further, the scheduling service node can convert the first task information into first task metadata belonging to the task data format. Among them, the scheduling service node can obtain the data parameters corresponding to the task data format, and convert the first task information into first task metadata based on the data parameters. Here, the data parameters refer to the parameters required to convert the data format of a certain data into the task data format, and are used to represent the parameter composition of the task metadata. Specifically, the scheduling service node can perform type recognition on the first task information to obtain the first task type indicated by the first task information. The first task type is used to represent the operation type of the first message task indicated by the first task information, such as the write task type, the delete task type, or the update task type, etc., which are not limited here. The scheduling service node can obtain the data parameters corresponding to the task data format, and obtain the parameter-associated data from the first task information based on the data parameters. The parameter-associated data is used to represent the value of the data parameters in the first task information; based on the data parameters, the first task type and the parameter-associated data are combined to form the first task metadata. The first task metadata can represent the parameters and the values of the parameters included in the first task information, and can be used to represent the task target and the involved data of the first message task indicated by the first task information. The task data format can be a data exchange format, such as the JSON format or the xml format, etc., or it can be a table format (in this case, one task metadata corresponds to one piece of data), which is not limited here. Optionally, the parameters included in different task information may be different. The scheduling service node can perform type recognition on the first task information to obtain the first task type indicated by the first task information; obtain the data parameters and the parameter-associated data corresponding to the data parameters from the first task information based on the task data format, and combine the first task type, the data parameters, and the parameter-associated data corresponding to the data parameters to form the first task metadata. For example, a possible first task metadata can be "Task ID: First Task ID; Task Type: First Task Type; Source Path: Data Table 1; Target Path: Data Table 2; Data Range: Data Range 1, etc.".
[0089] Optionally, the scheduling service node may convert the first task information into initial task metadata belonging to the task data format. This process may refer to the above process for obtaining the first task metadata. Further, a task priority may be added to the initial task metadata to obtain the first task metadata. Specifically, the scheduling service node may obtain a priority determination parameter in the initial task metadata, determine the task priority corresponding to the first task information based on the priority determination parameter. The priority determination parameter refers to a parameter used for priority determination, and may include, but is not limited to, the task generation time of the first task information and the task association path corresponding to the first task information, etc.; combine the task priority with the initial task metadata to form the first task metadata. Alternatively, the scheduling service node may obtain the task priority from the first task information, and combine the task priority with the initial task metadata to form the first task metadata. That is to say, when the business application uploads event information, it will directly write the task priority, etc. into the event information.
[0090] Step S402, set the instance status of the first task instance corresponding to the first message task to the unexecuted status, and generate first instance data of the first task instance based on the instance status of the first task instance.
[0091] In the embodiment of the present application, the first message task refers to the message task indicated by the first task information. Specifically, the scheduling service node may split the first task information into initial instance data corresponding to multiple first task instances respectively; set the instance status of the first task instance included in the first message task to the unexecuted status; combine the initial instance data corresponding to each first task instance with the instance status to form the first instance data of the first task instance. The scheduling service node may add the first instance data to the first task database, and the first instance data is associated with the first task metadata. Among them, the instance status of a task instance is used to represent the processing situation of the task instance, and the instance status may include, but is not limited to, the unexecuted status, the executing status, the execution completed status, and the execution abnormal status, etc. For example, the unexecuted status is used to represent that the corresponding task instance has not started to be processed; the executing status is used to represent that the corresponding task instance has been sent to the execution service node, but the feedback on the processing status result of the task instance has not been received; the execution completed status is used to represent that the corresponding task instance has been executed successfully; the execution abnormal status is used to represent that the corresponding task instance has been executed, but the execution has failed. By adding the instance status of each instance data to the first task database, the scheduling service node can conveniently and quickly implement the scheduling of each instance data, improving the convenience of task data processing.
[0092] Among them, the above first task metadata and first instance data can be used to refer to any task metadata and instance data in the first task database. That is, the generation and storage processes of the first task metadata and first instance data can be applied to the generation and storage processes of any task metadata and instance data in the first task database.
[0093] Step S403: When the scheduling service node issues a task, based on the instance status of the instance data in the first task database, obtain the instance data to be processed from the instance data in the first task database, and send the instance data to be processed and the task metadata associated with the instance data to be processed to the execution service node, so that the execution service node processes the instance data to be processed and the task metadata associated with the instance data to be processed, and feeds back the processing status result for the instance data to be processed and the task metadata associated with the instance data to be processed to the scheduling service node.
[0094] In the embodiment of the present application, when the scheduling service node issues a task, the scheduling service node can, based on the instance status of the instance data in the first task database, obtain the instance data to be processed from the instance data in the first task database, and send the instance data to be processed and the task metadata associated with the instance data to be processed to the execution service node (WorkerServer). The instance data to be processed may or may not include the first instance data. That is, the decoupling of the storage and scheduling of task information is realized, the convenient processing of task information is realized, and the convenience and efficiency of task data processing are improved. Among them, the instance data in the first task database includes the first instance data, and the instance data to be processed includes the instance data with the instance status of the unexecuted state; the processing status result is used to update the instance status in the instance data to be processed. Among them, the generation process of the instance data to be processed is the same as the generation process of the first instance data.
[0095] Among them, when obtaining the instance data to be processed, specifically, when the scheduling service node issues a task, obtain the instance status of the instance data included in the first task database, and determine the instance data with the instance status of the unexecuted state or the execution abnormal state as the initial instance data, so as to realize the processing of the unprocessed instance data and the reprocessing of the instance data with processing failures, and improve the security of task data processing. Further, the scheduling service node can determine the instance data to be processed from the initial instance data.
[0096] Specifically, the scheduling service node can determine the initial instance data as the instance data to be processed.
[0097] Alternatively, the scheduling service node may determine the number of instances to be dispatched based on the execution service nodes, and obtain the instance data to be processed from the initial instance data according to the number of instances to be dispatched, where the number of the instance data to be processed is less than or equal to the number of instances to be dispatched. For example, if the number of the initial instance data is greater than or equal to the number of instances to be dispatched, the instance data to be processed is obtained from the initial instance data. At this time, the number of the instance data to be processed is the number of instances to be dispatched. If the number of the initial instance data is less than the number of instances to be dispatched, the initial instance data is determined as the instance data to be processed. At this time, the number of the instance data to be processed is less than the number of instances to be dispatched.
[0098] Specifically, the scheduling service node may obtain the number of valid execution nodes of the execution service nodes that have the task execution conditions, and obtain the task execution data volume of the valid execution nodes. The valid execution node refers to an execution service node that has the instance processing condition, such as an execution service node in an idle state or with a node utilization rate less than the node operation threshold. That is to say, when receiving the instance data, the valid execution node can process the instance data. The node utilization rate is used to represent the node usage of the corresponding execution service node, and the node operation threshold is used to represent the upper limit of the node utilization rate restricted for the execution service node. The task execution data volume is used to represent the number of instance data that the corresponding valid execution node can currently process. Based on the number of valid nodes and the task execution data volume, the number of instances to be dispatched is determined, and the instance data to be processed is obtained from the initial instance data according to the number of instances to be dispatched. The number of the instance data to be processed is the number of instances to be dispatched. For example, the scheduling service node may determine the sum of the task execution data volumes of the valid execution nodes as the number of instances to be dispatched. Alternatively, the scheduling service node may obtain the number of valid execution nodes of the execution service nodes that have the task execution conditions, multiply the number of valid nodes by the unit processing data volume to determine the number of instances to be dispatched, and obtain the instance data to be processed from the initial instance data according to the number of instances to be dispatched. The unit processing data volume is used to represent the upper limit of the number of instance data that the scheduling service node sends to each execution service node.
[0099] By dispatching instances to multiple execution service nodes and having the multiple execution service nodes process the task instances, the efficiency of task instance processing is improved. When the number of instances to be dispatched is adopted, the node burden of the execution service nodes can be reduced, and the processing of the instance data by the scheduling service node does not involve execution, resulting in less resource consumption, so that the cooperation efficiency among various service nodes can be improved. Moreover, the scheduling service node can schedule the instance data in the first task database according to the load of the execution service nodes (such as the task execution data volume), so as to ensure the stability and efficiency of the task data processing architecture.
[0100] Among them, after sending the instance data to be processed and the task metadata associated with the instance data to be processed to the execution service node, the scheduling service node can update the instance status in the instance data to be processed to the in-execution status. When receiving the processing status result feedback from the execution service node, update the instance status in the instance data to be processed based on the processing status result. For example, if the processing status result is a processing completion result, update the instance status in the instance data to be processed to the execution completion status; if the processing status result is a processing exception result, update the instance status in the instance data to be processed to the execution exception status. Optionally, the number of times the task has been processed can also be added to the instance data to be processed. At this time, when obtaining the instance data to be processed, the instance data with the instance status being the execution exception status and the number of times the task has been processed being less than the exception handling times threshold, or the instance data with the instance status being the unexecuted status, can be determined as the instance data to be processed. Among them, the exception handling times threshold is used to represent the maximum number of times to process the same instance data. That is to say, it is possible to implement the processing of instance data and the retry processing when the processing of instance data fails, thereby improving the fault tolerance of task data processing.
[0101] Among them, the number of execution service nodes can be M, where M is a positive integer. When sending the instance data to be processed and the task metadata associated with the instance data to be processed to the execution service nodes, the scheduling service node can split the instance data to be processed into P sub-data to be processed, and send the P sub-data to be processed to P execution service nodes, with one sub-data to be processed corresponding to one execution service node, and P being a positive integer less than or equal to M. Specifically, the scheduling service node can determine the target execution service node from the M execution service nodes based on the node loads and node states of the M execution service nodes; the node state of the target execution service node is the node running state; send the instance data to be processed and the task metadata associated with the instance data to be processed to the target execution service node; at this time, P is the number of target execution service nodes. Or, P is M, that is, the scheduling service node can split the instance data to be processed into M sub-data to be processed, and send the M sub-data to be processed to the M execution service nodes, that is, one sub-data to be processed corresponds to one execution service node. Or, when obtaining the instance data to be processed based on the number of valid nodes and the amount of task execution data, P is the number of valid nodes, and the scheduling service node can split the instance data to be processed into P sub-data to be processed based on the amount of task execution data corresponding to each of the P valid execution nodes; based on the amount of task execution data corresponding to each of the P valid execution nodes, send the P sub-data to be processed to the P valid execution nodes. Or, when obtaining the instance data to be processed based on the number of valid nodes and the unit processing data volume, P is the number of valid nodes, and the scheduling service node can directly split the instance data to be processed into P sub-data to be processed, where the scheduling service node can evenly divide the instance data to be processed into P sub-data to be processed, or randomly divide the instance data to be processed into P sub-data to be processed, etc., and the scheduling service node can send the P sub-data to be processed to the P valid execution nodes.
[0102] In the embodiments of the present application, the scheduling service node can store and schedule the task information generated by the application service node, implement asynchronous scheduling of the task information, and can schedule based on the task priority, thereby ensuring the reasonable utilization of resources, improving the efficiency of task data processing and resource utilization rate, and improving the performance of the service node.
[0103] Further, please refer to Figure 5 , Figure 5 which is another flowchart of the task data processing method provided by the embodiments of the present application. As shown in Figure 5 , this task data processing process is implemented by the execution service node, and this task data processing process includes the following steps:
[0104] Step S501: Receive the instance data to be processed and the task metadata associated with the instance data to be processed sent by the scheduling service node, process the instance data to be processed and the task metadata associated with the instance data to be processed, and obtain the processing status result for the instance data to be processed and the task metadata associated with the instance data to be processed.
[0105] In the embodiment of the present application, receive the instance data to be processed and the task metadata associated with the instance data to be processed sent by the scheduling service node, process the instance data to be processed and the task metadata associated with the instance data to be processed, and obtain the processing status result for the instance data to be processed and the task metadata associated with the instance data to be processed; the instance data to be processed is generated based on the instance status of the to-be-processed task instance in the task metadata associated with the instance data to be processed, and the to-be-processed task instance refers to the task instance indicated by the instance data to be processed; the initial value of the instance status of the to-be-processed task instance is the unexecuted state, and the to-be-processed task instance includes instance data with an unexecuted state; the task metadata associated with the instance data to be processed is data belonging to the task data format converted from task information, and the task metadata associated with the instance data to be processed is added to the first task database by the scheduling service node, and the task data format refers to the format of the data provided by the scheduling service node for the execution service node. Among them, for the processing process of the instance data to be processed and the task metadata associated with the instance data to be processed, reference can be made to Figure 4 In steps S401 to S402 of, the processing process of the first instance data and the first task metadata will not be elaborated here.
[0106] Specifically, when processing the instance data to be processed and the task metadata associated with the instance data to be processed and obtaining the processing status result for the instance data to be processed and the task metadata associated with the instance data to be processed, the execution service node can obtain the task type, the data to be updated, and the task association path (such as the source path and the target path, etc.) from the instance data to be processed and the task metadata associated with the instance data to be processed. Among them, the task association path is used to represent the data table associated with the data to be updated. Determine the task logic corresponding to the instance data to be processed based on the task type, and generate an initial task instruction according to the task logic. Update the initial task instruction with the data to be updated and the task association path to obtain a to-be-processed instruction, execute the to-be-processed instruction, and when the execution of the to-be-processed instruction is completed, obtain the processing status result for the to-be-processed instruction. Among them, the completion of the execution of the to-be-processed instruction includes the following situations: the execution of the to-be-processed instruction is successful, or the execution of the to-be-processed instruction fails, etc.
[0107] For example, assume that the obtained task type is a write task type, the data to be updated is the data in Data Table 1 with a resource data volume greater than 10, and the task association path is between Data Table 1 and Data Table 2. Then, the task logic corresponding to the instance data to be processed can be determined based on the task type. It can be to read the data to be updated with a resource data volume greater than 10 from Data Table 1 and write the data to be updated into Data Table 2. At this time, the initial task instruction can include a read instruction and a write instruction. Further, update the initial task instruction with the data to be updated and the task association path to obtain the instruction to be processed. At this time, the instruction to be processed can include the instruction to be processed 1 for reading the data to be updated with a resource data volume greater than 10 from Data Table 1 and the instruction to be processed 2 for writing the data to be updated into Data Table 2. For example, assume that the data processing languages corresponding to Data Table 1 and Data Table 2 are both Structured Query Language (SQL). Then, the instruction to be processed 1 can be "to-be-updated-data = select * from Data Table 1 where resource data volume > 10", and the instruction to be processed 2 is "insert to-be-updated-data into Data Table 2". Or, the initial task instruction can include a combined instruction composed of a read instruction and a write instruction. At this time, the instruction to be processed can be "insert into Data Table 2 select * from Data Table 1 where resource data volume > 10", etc., which is not limited here.
[0108] Optionally, the number of instance data to be processed is H, where H is a positive integer. When processing the instance data to be processed and the task metadata associated with the instance data to be processed and obtaining the processing status result for the instance data to be processed and the task metadata associated with the instance data to be processed, the execution service node may obtain, from the task metadata respectively associated with the H instance data to be processed, the first priority corresponding to each of the H instance data to be processed, where the first priority is used to indicate the priority of the message task to which the to-be-processed task instance indicated by the corresponding instance data to be processed belongs; obtain, from the H instance data to be processed, the second priority corresponding to each of the H instance data to be processed, where the second priority is used to indicate the priority of the to-be-processed task instance indicated by the corresponding instance data to be processed within the message task to which the to-be-processed task instance belongs. Based on the first priority and the second priority corresponding to each of the H instance data to be processed, determine the instance processing order of the H instance data to be processed. Among them, in this instance processing order, for any two instance data to be processed, the first priority of the instance data to be processed in the front is higher than or equal to the first priority of the instance data to be processed in the back; for any two instance data to be processed corresponding to the same task metadata, the second priority of the instance data to be processed in the front is higher than the second priority of the instance data to be processed in the back. The execution service node may process the H instance data to be processed in sequence according to the instance processing order. When the processing of each instance data to be processed is completed, obtain the processing status result corresponding to the instance data to be processed. That is to say, when the processing of an instance data to be processed is completed, a processing status result for the instance data to be processed can be obtained. Among them, for the processing process of any instance data to be processed, reference may be made to the relevant process description in the above-mentioned process related to processing the instance data to be processed and the task metadata associated with the instance data to be processed.
[0109] Step S502: Feed back the processing status result to the scheduling service node, so that the scheduling service node updates the instance status in the instance data to be processed based on the processing status result.
[0110] In the embodiment of the present application, the execution service node may feed back the processing status result of the instance data to be processed to the scheduling service node that sends the instance data to be processed and the task metadata associated with the instance data to be processed. After obtaining the processing status result of the instance data to be processed, the scheduling service node may update the instance status in the instance data to be processed based on the processing status result. This process may refer to Figure 4 the relevant description in step S403 of
[0111] In the embodiments of the present application, the execution service node can process the received instance data to be processed and the task metadata associated with the instance data to be processed, implement task data processing, and feedback the processing status result of the instance data to be processed to the scheduling service node to achieve status synchronization of the instance data with the scheduling service node. Among them, the execution service node can be a distributed processing engine, which can be used to synchronously process multiple tasks, support synchronous tasks between dozens and hundreds, and even support more synchronous tasks with the optimization of the execution service node. At the same time, multiple execution service nodes are adopted to support the execution of larger-scale tasks, that is, support the simultaneous processing of more instance data, thereby improving the efficiency of task data processing, and enabling the present application to well achieve the synchronous processing of a large number of task information when applied to the processing scenario of a large number of task information, improving the concurrency of task data processing, and the increase in the task information processed simultaneously also enables the improvement of the real-time performance of task data processing in the processing scenario of a large number of task information.
[0112] Further, reference can be made to Figure 6 , Figure 6 which is a schematic diagram of an interaction scenario for task data processing provided by the embodiments of the present application. As Figure 6 shown, the task data processing architecture may include a second message queue 601, an application service node 602, a scheduling service node 603, an execution service node 604, etc. Among them, business events will be generated in the upstream business application of the task data processing architecture in the present application, and the event information of the business events can be sent to the second message queue 601. Among them, the business event can be a distributed file system (Hadoop Distributed File System, Hdfs) log, a database update event, a user operation, etc., which is not limited here. Among them, the second message queue 601 can be a distributed message system, such as a multi-layer architecture distributed system (Pulsar) or a single-layer architecture message system (kafka), etc. The second message queue 601 can include multiple second message queues, such as a second message queue 601a and a second message queue 601b, etc. Among them, the task data processing process includes the following parts:
[0113] ① The application service node 602 can include multiple application service nodes, such as an application service node 602a and an application service node 602b, etc. The application service node 602 can publish and subscribe to information for the second message queue 601, obtain event information from the second message queue 601 based on the subscription information, and filter the event information to obtain the filtered event information.
[0114] ② This step is an optional process. The application service node 602 can report the node status information of the application service node to the second task database.
[0115] ③ The application service node 602 can convert the filtered event information into task information (such as the first task information mentioned above), and add the task information to the storage space, which can be the first message queue or the second task database.
[0116] ④ The scheduling service node 603 can obtain task information (such as the first task information mentioned above) from the storage space, specifically by using an event acquisition command to obtain the task information from the storage space.
[0117] ⑤ The scheduling service node 603 can convert the obtained task information into task metadata in the task data format (such as the first task metadata corresponding to the first task information mentioned above), and can generate instance data (such as the first instance data associated with the first task metadata) based on the task metadata, and write the task metadata and the instance data into the first task database in an associated manner. Among them, the scheduling service node 603 can write the task metadata and the instance data into the first task database separately, or, after obtaining the instance data, it can also write the task metadata and the instance data into the first task database at the same time. Optionally, the first task database and the second task database can be the same database or different databases, which is not restricted here.
[0118] ⑥ The scheduling service node 603 can obtain the instance data to be processed and the task metadata associated with the instance data to be processed (which can be simply referred to as the task metadata to be processed) from the first task database, and send the instance data to be processed and the task metadata to be processed to the execution service node 604. Optionally, the scheduling service node 603 can perform data optimization on the task metadata in the first task database to obtain updated task metadata, or, in step ⑤, it can perform data optimization on the generated task metadata and write the optimized task metadata into the first task database; at this time, the task metadata in the first task database is the task metadata after data optimization. Optionally, the data optimization can include but is not limited to task deduplication, consistency processing (such as adding instance status for instance data), and task data processing optimization (such as determining task priorities for task metadata).
[0119] ⑦ The execution service node 604 can perform instance dispatch, that is, process the obtained instance data to be processed and the task metadata to be processed.
[0120] ⑧The execution service node 604 can feedback the processing status result for the to-be-processed instance data and the to-be-processed task metadata to the scheduling service node 603.
[0121] ⑨The scheduling service node 603 can update the instance status in the to-be-processed instance data based on the processing status result.
[0122] Specifically, please refer to Figure 7 , Figure 7 which is a method interaction flowchart for task data processing provided by an embodiment of the present application. As Figure 7 shown, the task data processing process includes the following steps:
[0123] Step S701, in response to a processing request for a service event, add the event information of the service event to a second message queue.
[0124] In an embodiment of the present application, a service application can, in response to a processing request for a service event, add the event information of the service event to a second message queue. Among them, the number of second message queues is A, and the service application can add the event information of the service event to A second message queues. Alternatively, each of the A second message queues corresponds to a data type, and the service application can identify the event type of the service event and, based on the event type of the service event and the data types corresponding to the A second message queues respectively, add the event information of the service event to the A second message queues. Among them, the second message queue can be regarded as a distributed message system, which can support subscription / reception of data, can support a transmission bandwidth of terabytes (TB) per second, and may even support a larger transmission bandwidth after optimization, thereby improving the efficiency of big data processing.
[0125] Optionally, the data types corresponding to the respective second message queues may be determined based on the event source. For example, if the second message queue 1 corresponds to the event source 1 and the second message queue 2 corresponds to the event source 2, then the business application can identify the event type of the business event. At this time, the event type of the business event is used to represent the event source of the business event, and the event information of the business event is written into the second message queue corresponding to the event source of the business event; or, the data types corresponding to the respective second message queues may be determined based on the event topic, which may include, but is not limited to, data warehouse partition update events, Hdfs directory update events, table update events, etc. For example, if the data type corresponding to the second message queue 1 is "data warehouse partition update event", the data type corresponding to the second message queue 2 is "Hdfs directory update event", and the data type corresponding to the second message queue 3 is "table update event", then the event type of the business event is used to represent the topic of the business event. For example, when the event type of the business event is "table update event", the event information of the business event can be written into the second message queue 3, etc.
[0126] Step S702, obtain the event information to be processed from the second message queue.
[0127] In the embodiment of the present application, the application service node may publish subscription information for the second message queue, and obtain the event information to be processed from the second message queue based on the subscription information. The application service node may be stateless and can be horizontally scaled to hundreds or thousands of service nodes, and can support processing event information at the TB level per second, so as to achieve real-time performance in scenarios where a large number of event information is processed. Optionally, the number of application service nodes may be B. Different application service nodes may integrate different message processing components or the same message processing components, which is not limited here. Among them, the message processing components may include, but are not limited to, a multi-layer architecture distributed system (pulsar client), a large-scale message middleware system (Tube client), and an application programming interface (Application Programming Interface, API), etc. And these message processing components can all support the processing of large amounts of data. Through the subscription information, the application service node can actively obtain event information from the second message queue in a streaming manner, and by using multiple application service nodes and message processing components, the application service node can support high-concurrency processing of event information and can also ensure that the event information is not lost, thereby improving the security of task data processing.
[0128] Optionally, the application service node may obtain the initial event information from the second message queue and determine the initial event information as the event information to be processed. Or, the application service node may filter the initial event information to obtain the event information to be processed (i.e.,Figure 6 (the filtered event information in step ①). Specifically, the application service node can obtain an event filtering condition and filter the initial event information that meets the event filtering condition in the initial event information to obtain the event information to be processed. Among them, the event filtering condition can include, but is not limited to, read-only filtering, dirty data filtering, and event structure filtering, etc. The event filtering condition can be determined by a business object (such as a manager of a task data processing architecture), or can be determined according to historical event processing situations. The event filtering condition can be updated based on needs, and is not limited here. Among them, dirty data filtering refers to filtering various incorrect, incomplete, inaccurate, or inconsistent data, that is, filtering out abnormal event information; read-only filtering refers to filtering event information that only indicates reading data; event structure filtering refers to filtering event information whose event structure does not conform to the standard event structure processed in the task data processing architecture, and the event structure is used to represent the data composition of the corresponding event information.
[0129] Step S703, convert the event information to be processed into task information belonging to the scheduling data format.
[0130] In the embodiment of the present application, the application service node can convert the event information to be processed into task information belonging to the scheduling data format, and the scheduling data format refers to the format of data that the scheduling service node can process. Specifically, the application service node can parse the event information to be processed to obtain event parameters, obtain the parameter values corresponding to the event parameters from the event information to be processed, and based on the scheduling data format, form the parameter values corresponding to the event information to be processed into task information. Among them, the scheduling data format can be a data exchange format, and the event parameters are used to represent the parameters included in the event information to be processed. For example, a possible example of task information can be "Task type: data synchronization; Source path: data table 1; Target path: data table 2; Data range: data range 1 in data table 1", indicating that the message task indicated by this task information is to synchronize and write the data within data range 1 in data table 1 into data table 2.
[0131] Further, the application service node can write the task information into the storage space, which can be the first message queue or the second task database. Of course, the storage space can also include both the first message queue and the second task database at the same time. At this time, the application service node can randomly select one of the first message queue and the second task database as the storage space to be used, or can determine the storage space to be used from the first message queue and the second task database based on the event parameters of the event information to be processed, convert the event information to be processed into task information belonging to the scheduling data format, and write the task information into the storage space to be used. Of course, when the storage space is the first message queue, the scheduling data format is the format of the data maintained by the first message queue; when the storage space is the second task database, the scheduling data format is the format of the data maintained by the second task database. By this means, various modes of data interaction can be achieved between the application service node and the scheduling service node, supporting message queues or databases, etc., and the event information can be more efficiently converted into task information, that is, the business event can be more efficiently converted into message tasks and task instances, thereby improving the efficiency of task data processing.
[0132] Optionally, the second task database can also include the node status of each service node (including the application service node, the scheduling service node, and the execution service node), and the node status is used to indicate whether the corresponding service node is in the online state or the offline state. For example, in Figure 6 step ② shown, the application service node 602 can write the node status information of the application service node into the second task database. For example, a possible storage method for the node status of each service node can be seen in Table 1:
[0133] Table 1
[0134] Service Node Node Status Application Service Node 1 Go Online Scheduling Service Node 1 Go Online Execution Service Node 1 Go Offline … …
[0135] As shown in Table 1, the second task database can also include the node description information of each service node, and the node description information is used to represent other information of the corresponding service node, such as the node address, the components integrated in the service node, and the node functions included in the service node, etc., which are not limited here.
[0136] Step S704, obtain the first task information.
[0137] In the embodiments of the present application, the scheduling service node can include multiple business functions, such as Figure 6The task splitting function, concurrency control function, priority sorting function, load processing function, etc. shown in the figure. Among them, the task splitting function is used to split the message task indicated by the task information into multiple task instances; the concurrency control function is used to implement the concurrent processing of multiple task information; the priority sorting function is used to generate the task priorities corresponding to each task information; the load processing function is used to manage the node loads of each scheduling service node, obtain the node loads of each execution service node, and can be used to implement the distribution of instance data to the execution service node based on the node loads of each execution service node.
[0138] Specifically, the scheduling service node can obtain the first task information from the storage space, and this process can refer to Figure 4 the relevant description in step S401 of
[0139] Step S705, convert the first task information into the first task metadata belonging to the task data format, and add the first task metadata to the first task database.
[0140] In the embodiment of the present application, this process can refer to Figure 4 the relevant description in step S401 of
[0141] Table 2
[0142]
[0143] As shown in Table 2, each row of data corresponds to a task metadata, and the task identifier is used to uniquely indicate a task metadata. After adding the first task metadata to the first task database, the first task database includes the first task metadata. Optionally, the data parameter may further include a task name, a data start time, a task period, and a task associated object, etc., and a task executable operation component may also be associated with the task metadata. Among them, the task name is used to represent the name of the message task indicated by the corresponding task metadata. The data start time is used to represent the start time of the message task indicated by the corresponding task metadata. The task period is used to represent the processing period of the corresponding task metadata. For example, it can be non-periodic (that is, the task metadata is not scheduled periodically, or the task metadata is scheduled once), or a task scheduling period (that is, the corresponding task metadata is scheduled periodically using the task scheduling period). The task associated object is used to represent the object (i.e., the person in charge) responsible for managing the corresponding task metadata, and the object (i.e., the data responsible user) responsible for managing the data involved in the corresponding task metadata. The task executable operation component is used to represent the component that can perform operations on the corresponding task metadata, such as an instance display operation component, a data supplement operation component, and a view display operation component, etc.; among them, when the instance display operation component is triggered, it is used to display the instance data associated with the corresponding task metadata; the data supplement operation component is used to obtain the data submitted based on the data supplement operation component and add the data to the corresponding task metadata when triggered; the view display operation component is used to display the view of the corresponding task metadata.
[0144] Step S706, set the instance status of the first task instance corresponding to the first message task to the unexecuted status, and generate the first instance data of the first task instance based on the instance status of the first task instance.
[0145] In the embodiment of the present application, this process can refer to Figure 4 the relevant description in step S402 of
[0146] Table 3
[0147]
[0148] As shown in Table 3, one task metadata is associated with multiple instance data. After the first instance data is written into the first task database, the first instance data is included in Table 3 above.
[0149] Alternatively, the scheduling service node can write the first instance data into a data table for processing instance data. For example, a possible storage method for instance data can be seen in Table 4:
[0150] Table 4
[0151] Task ID Instance Data Task ID 1 Instance Data 1 … …
[0152] As shown in Table 4, through the task identifier, the instance data is associated with the task metadata associated with the instance data in the first task database.
[0153] Among them, the instance parameters corresponding to the instance data may include but are not limited to the data start time, instance execution time, execution duration, instance status, and the number of times the task has been processed, etc. It is also possible to associate an instance executable operation component with the instance data. Among them, the data start time is used to represent the initiation time of the message task indicated by the task metadata associated with the instance data. The instance execution time is used to represent the processing time period of the instance data. The execution duration is used to represent the duration consumed for processing the instance data. The instance status is used to represent the execution situation of the instance data, such as it can be an unexecuted state, an executing state, an executed state, or an execution exception state, etc. The number of times the task has been processed is used to represent the number of times the instance data has been processed. The instance executable operation component is used to represent the component that can perform operations on the instance data, such as an instance log viewing operation component, an instance detection operation component, and an instance retry operation component, etc. Among them, when the instance log viewing operation component is triggered, it is used to display the running log of the instance data; when the instance detection operation component is triggered, it is used to detect the instance data and the running log of the instance data; when the instance retry operation component is triggered, it is used to reprocess the instance data.
[0154] Optionally, in this application, the task metadata and instance data, etc. can be displayed for the user through the front-end page. For example, it can be seen in Figure 8 , Figure 8 is a schematic diagram of a data display scenario provided by an embodiment of this application. As Figure 8As shown, the service device can respond to a task viewing request and display a task management page 801, which displays task metadata and task executable operation components associated with the task metadata, such as task metadata 8021 and task executable operation component 8022 associated with task metadata 8021. In response to a trigger operation for displaying an operation component for an instance of the second task metadata, assuming the second task metadata is task metadata 8021, that is, in response to a trigger operation for displaying operation component 802a for an instance of task metadata 8021, display an instance management page 803 for the instance of the second task metadata, which displays instance data associated with the second task metadata and instance executable operation components associated with the instance data, such as instance data 8041 and instance operation component 8042 associated with instance data 8041. Optionally, the service device can forward the trigger operation for the instance retry operation component for the second instance data to the scheduling service node in response to the trigger operation for the instance retry operation component for the second instance data. When detecting the trigger operation for the instance retry operation component for the second instance data, for example, assuming the second instance data is instance data 8041, that is, detecting the trigger operation for instance retry operation component 804a for instance data 8041, the scheduling service node can send the second instance data to the execution service node to enable the execution service node to reprocess the second instance data. Among them, Figure 8 In "1 / 3" of the number of times the task shown in has been executed, "1" is used to represent the number of times instance data 8041 has been processed, and "3" is used to represent the exception handling times threshold, that is, instance data 8041 is processed at most 3 times.
[0155] Step S707, when the scheduling service node issues a task, based on the instance status of the instance data in the first task database, obtain the instance data to be processed from the instance data in the first task database.
[0156] In the embodiment of this application, this process can refer to Figure 4 the relevant description in step S403 of, which will not be elaborated here.
[0157] Step S708, send the instance data to be processed and the task metadata associated with the instance data to be processed.
[0158] In the embodiment of this application, the scheduling service node can send the instance data to be processed and the task metadata associated with the instance data to be processed to the execution service node. This process can refer to Figure 4 the relevant description in step S403 of, which will not be elaborated here.
[0159] Step S709: Process the instance data to be processed and the task metadata associated with the instance data to be processed, and obtain the processing status result.
[0160] In the embodiments of the present application, the execution service node can process the obtained instance data to be processed and the task metadata associated with the instance data to be processed, and obtain the processing status result for the instance data to be processed and the task metadata associated with the instance data to be processed. This process can refer to the relevant description in Figure 5 Step S501. Among them, the execution service node also supports horizontal expansion. A single execution service node can concurrently process dozens or even more instance data. By horizontal elastic expansion, that is, increasing the number of execution service nodes, the task data processing architecture of the present application can even support the processing of billions of instance data per day, improving the efficiency of task data processing. As Figure 6 shown, the execution service node can also include an alarm component, a detection component, a log component, etc. Among them, the alarm component is used to give an alarm prompt when an exception occurs during the processing of instance data; the detection component is used to detect the processing process of instance data; the log component is used to generate log information for the processing process of instance data, etc.
[0161] Optionally, an execution service node can generate one or more instance execution threads. Specifically, any execution service node can generate an instance execution thread (TaskExecuteThread) according to the data volume of the obtained instance data to be processed; call the instance execution thread to process the instance data to be processed and the task metadata associated with the instance data to be processed, and obtain the processing status result corresponding to the instance data to be processed. Among them, the execution service node can generate an instance execution thread based on the data volume of the obtained instance data to be processed, so as to achieve multi-threaded concurrent processing of the instance data to be processed, thereby improving the efficiency of task data processing.
[0162] Step S710: Feedback the processing status result.
[0163] In the embodiments of the present application, the execution service node can feedback the processing status result to the scheduling service node. This process can refer to the relevant description in Figure 5 Step S502. Optionally, the execution service node can feedback the processing status result to the scheduling service node through Remote Procedure Call (RPC) for state machine maintenance, that is, update the instance status in the instance data.
[0164] Step S711: Update the instance status of the instance data to be processed based on the processing status result.
[0165] In an embodiment of the present application, the scheduling service node may update the instance status of the instance data to be processed based on the processing status result. For the relevant description of this process, reference may be made to Figure 4 the relevant description in step S403 of
[0166] In the above process, the business device may display the execution status of each step in the task data processing on the front-end page, enabling the user to intuitively view the processing status of any message task and realizing the visualization of task data processing. In the present application, the scheduling service node and the execution service node can perceive each other's node status to achieve automatic disaster tolerance. For example, the scheduling service node may select a suitable execution service node to process the instance data based on the task execution data volume of the execution service node, etc.
[0167] For example, refer to Figure 9 Figure 9 which is a schematic diagram of a data synchronization event scenario provided by an embodiment of the present application. As Figure 9 shown, the update (Schema Change) event for a data table may include, but is not limited to, a data addition event (AddColumn), a data deletion event (Drop Column), a data modification event (Change Column), a table renaming event (Rename Table), etc. When performing data synchronization, for example, if the message task corresponding to the metadata of the task to be processed is to synchronize the data in an offline registration (Hive) table to a data warehouse (OLAP) engine table, the business application may send the event information of the business event generated during data synchronization to the second message queue, which can avoid the situation of event loss caused by directly sending to a certain synchronization service, such as the situation of event loss caused by abnormal synchronization interfaces, etc.; after obtaining the event information through the second message queue, the application service node may convert the event information into task information and write the task information into the storage space. Optionally, the business event may be converted into a message task, and based on the message task, the event information may be converted into the task information corresponding to the message task. For example, in data synchronization, the business event may be converted into a message task of synchronizing from a Hive table (i.e., data table 1) to an OLAP engine table (i.e., data table 2), or the business event may be converted into a message task of data optimization task for the OLAP engine table, etc. Further, the data synchronization may further include the following steps:
[0168] 1. The scheduling service node may detect the task information through a data detector, trigger task generation through the data detector, and obtain the task information.
[0169] 2. The scheduling service node can convert the obtained task information into task metadata in the task data format, add the task metadata to the first task database, and generate instance data of the task instances included in the message task indicated by the task metadata, and add the instance data to the first task database. This process can be referred to in Figure 4 the relevant processing procedures of the first task metadata and the first instance data in steps S401 and S402 of
[0170] 3. The scheduling service node can obtain the instance data to be processed and the task metadata associated with the instance data to be processed (which can be denoted as the task metadata to be processed) from the first task database, and send the instance data to be processed and the task metadata to be processed to the execution service node. Alternatively, the scheduling service node can determine the task instance to be processed from the first task database, send an instance processing request for the task instance to be processed to the execution service node, and the execution service node can, based on the instance processing request, obtain the instance data to be processed and the task metadata associated with the instance data to be processed from the first task database. This process can be referred to in Figure 4 the relevant description in step S403 of
[0171] For example, a possible task metadata can be referred to in Table 5:
[0172] Table 5
[0173] Task ID Data Parameter Parameter Value Parameter ID Creation Time Update Time Task ID 1 Database Database 1 Parameter ID 1 Creation Time 1 Update Time 1 Task ID 1 Data Table Data Table 1 Parameter ID 2 Creation Time 1 Update Time 1 … … … … … …
[0174] Or, a possible task metadata can be referred to in Table 6:
[0175] Table 6
[0176] Data Parameter Parameter Value Task ID Task ID 1 Task Type Task Type 1 Task Name Task Name 1 … …
[0177] 4. The execution service node can request task execution and call the instance execution thread.
[0178] 5. The execution service node can perform data synchronization based on the instance execution thread. For example, by 5.1, synchronizing all the data in Data Table 1 to Data Table 2 in full; or by 5.2, synchronizing data to Data Table 2 according to the variables in Data Table 1, etc.
[0179] For example, a possible instruction to be processed is used to read data in partition 1 of a Hive table through a big data computing engine (such as Spark), and then call the database connection interface (Java Database Connectivity, Jdbc) of an OLAP engine table to write the read data into partition 1 of the OLAP engine table. Or, a possible instruction to be processed is used to synchronize data in a Hive table to an OLAP engine table using SQL through the capabilities of the OLAP engine table itself, such as "insert into OLAP engine table select * from Hive table".
[0180] 6. The execution service node can, through an instance execution thread, respond to a task execution request and obtain a processing status result for the instance data to be processed.
[0181] 7. The processing status result is fed back to the scheduling service node, and the scheduling service node updates the instance status in the instance data to be processed in the first task database based on the processing status result. Or, the execution service node can directly update the instance status in the instance data to be processed in the first task database based on the processing status result.
[0182] In an embodiment of the present application, first task information is obtained, the first task information is converted into first task metadata belonging to a task data format, and the first task metadata is added to a first task database; the instance status of a first task instance included in the first message task is set to an unexecuted status, and first instance data of the first task instance is generated based on the instance status of the first task instance; the first message task refers to the message task indicated by the first task information; when the scheduling service node issues a task, based on the instance status of the instance data in the first task database, the to-be-processed instance data is obtained from the instance data in the first task database, and the to-be-processed instance data and the task metadata associated with the to-be-processed instance data are sent to the execution service node, so that the execution service node processes the to-be-processed instance data and the task metadata associated with the to-be-processed instance data, and feeds back a processing status result for the to-be-processed instance data and the task metadata associated with the to-be-processed instance data to the scheduling service node; the instance data in the first task database includes the first instance data, and the to-be-processed instance data includes the instance data with an instance status of unexecuted; the processing status result is used to update the instance status in the to-be-processed instance data. By combining the scheduling service node and the execution service node, the acquisition of task information and the execution of message tasks are decoupled, so as to process tasks through different service nodes, enabling the scheduling service node to obtain and schedule task information, and the execution service node to process the message tasks indicated by the task information. The scheduling service node only obtains and schedules task information, consuming less resources and time. Even when the data volume of the message task is large, the scheduling service node can meet the requirements for resources such as the acquisition and scheduling of task information of the message task, thereby reducing the situation of task information loss and improving the security of task data processing. At the same time, multiple service nodes coordinate to work, and the instance status is introduced, enabling the scheduling service node to more conveniently obtain the instance data to be processed (i.e., the to-be-processed instance data), thus improving the efficiency of task data processing.
[0183] Further, please refer to Figure 10 , Figure 10 which is a schematic diagram of a task data processing device provided in an embodiment of the present application. The task data processing device 1000 may be a computer program (including program codes, etc.) running in a computer device. For example, the task data processing device may be an application software; the device may be used to execute the corresponding steps in the method provided in the embodiment of the present application. As Figure 10 shown, the task data processing device 1000 may be used for Figure 4 the computer device corresponding to the corresponding embodiment. Specifically, the device may include: a task processing module 11, an instance processing module 12, and a task issuing module 13.
[0184] The task processing module 11 is used to obtain the first task information, convert the first task information into first task metadata belonging to the task data format, and add the first task metadata to the first task database; the task data format refers to the data format provided by the scheduling service node for the execution service node.
[0185] The instance processing module 12 is used to set the instance status of the first task instance corresponding to the first message task to the unexecuted status, and generate first instance data of the first task instance based on the instance status of the first task instance; the first message task refers to the message task indicated by the first task information.
[0186] The task distribution module 13 is used to, when the scheduling service node distributes tasks, obtain the to-be-processed instance data from the instance data in the first task database based on the instance status of the instance data in the first task database, and send the to-be-processed instance data and the task metadata associated with the to-be-processed instance data to the execution service node, so that the execution service node processes the to-be-processed instance data and the task metadata associated with the to-be-processed instance data, and feeds back the processing status result for the to-be-processed instance data and the task metadata associated with the to-be-processed instance data to the scheduling service node; the instance data in the first task database includes the first instance data, and the to-be-processed instance data includes the instance data with the instance status being the unexecuted status; the processing status result is used to update the instance status in the to-be-processed instance data.
[0187] Among them, when obtaining the first task information, the task processing module 11 can be used to:
[0188] Scan the first message queue or the second task database. If the first message queue or the second task database is not empty, obtain the first task information from the first message queue or the second task database; the first task information is added to the first message queue or the second task database by the application service node; the first task information is obtained by the application service node through converting the initial task message; the data format of the first task information is the scheduling data format, and the scheduling data format is the data format supported by the scheduling service node, which refers to the data format provided by the application service node for the scheduling service node.
[0189] Among them, the scheduling service node belongs to the scheduling node cluster. When obtaining the first task information, the task processing module 11 can be used to:
[0190] Obtain the to-be-scheduled task information and the task hash value corresponding to the to-be-scheduled task information, and convert the task hash value into a data shard value based on the number of scheduling nodes; the number of scheduling nodes is the number of nodes included in the scheduling node cluster.
[0191] Obtain the first shard processing range of the local scheduling service node, and determine the scheduling task information whose data shard value belongs to the first shard processing range as the first task information;
[0192] Obtain the first task information.
[0193] Among them, the device 1000 further includes:
[0194] The node processing module 14 is used to send a node offline message to other scheduling service nodes when the local scheduling service node goes offline, so that other scheduling service nodes update the shard processing range of other scheduling service nodes based on the first shard processing range to obtain the second shard processing range; other scheduling service nodes refer to scheduling service nodes other than the local scheduling service node in the scheduling node cluster; the second shard processing range is used to represent the range of message tasks processed by other scheduling service nodes.
[0195] Among them, when converting the first task information into the first task metadata belonging to the task data format, the task processing module 11 can be used for:
[0196] Perform type recognition on the first task information to obtain the first task type indicated by the first task information;
[0197] Obtain the data parameters corresponding to the task data format, and obtain parameter-associated data from the first task information based on the data parameters;
[0198] Based on the data parameters, compose the first task type and the parameter-associated data into the first task metadata.
[0199] Among them, when converting the first task information into the first task metadata belonging to the task data format, the task processing module 11 can be used for:
[0200] Convert the first task information into initial task metadata belonging to the task data format;
[0201] Obtain the priority determination parameter in the initial task metadata, and determine the task priority corresponding to the first task information based on the priority determination parameter; the priority determination parameter refers to the parameter used for priority determination;
[0202] Compose the task priority and the initial task metadata into the first task metadata.
[0203] Among them, when the scheduling service node issues tasks, when obtaining the instance data to be processed from the instance data in the first task database based on the instance status of the instance data in the first task database, the task issuing module 13 can be used for:
[0204] When the scheduling service node issues a task, obtain the instance status of the instance data included in the first task database, and determine the instance data with the instance status being the unexecuted status or the execution exception status as the initial instance data;
[0205] Obtain the number of valid execution nodes of the execution service nodes that have the task execution conditions, and obtain the task execution data volume of the valid execution nodes;
[0206] Based on the number of valid nodes and the task execution data volume, determine the instance distribution quantity, and obtain the to-be-processed instance data from the initial instance data based on the instance distribution quantity.
[0207] Wherein, the apparatus 1000 further includes:
[0208] A status update module 15, configured to update the instance status in the to-be-processed instance data to the in-execution status after sending the to-be-processed instance data and the task metadata associated with the to-be-processed instance data to the execution service node;
[0209] The status update module 15 is further configured to update the instance status in the to-be-processed instance data based on the processing status result when receiving the processing status result fed back by the execution service node.
[0210] Wherein, the number of execution service nodes is M, and M is a positive integer; when sending the to-be-processed instance data and the task metadata associated with the to-be-processed instance data to the execution service node, the task distribution module 13 may be configured to:
[0211] Determine a target execution service node from the M execution service nodes based on the node load and node status of the M execution service nodes; the node status of the target execution service node is the node running status;
[0212] Send the to-be-processed instance data and the task metadata associated with the to-be-processed instance data to the target execution service node.
[0213] In the embodiment of the present application, the apparatus is applied to the scheduling service node, and can store and schedule the task information generated by the application service node, realize the asynchronous scheduling of the task information, and can perform scheduling based on the task priority, so as to ensure the reasonable utilization of resources, improve the efficiency of task data processing and resource utilization rate, and improve the performance of the service node.
[0214] Further, please refer to Figure 11 , Figure 11It is a schematic diagram of another task data processing device provided by an embodiment of the present application. The task data processing device may be a computer program (including program code, etc.) running in a computer device. For example, the task data processing device may be an application software. The device may be used to execute the corresponding steps in the method provided by the embodiment of the present application. As Figure 11 shown, the task data processing device 1100 may be used for Figure 5 the computer device in the corresponding embodiment. Specifically, the device 1100 may include: a data processing module 21 and a result feedback module 22.
[0215] The data processing module 21 is configured to receive the to-be-processed instance data sent by the scheduling service node and the task metadata associated with the to-be-processed instance data, process the to-be-processed instance data and the task metadata associated with the to-be-processed instance data, and obtain a processing status result for the to-be-processed instance data and the task metadata associated with the to-be-processed instance data. The to-be-processed instance data is generated based on the instance status of the to-be-processed task instance in the task metadata associated with the to-be-processed instance data. The to-be-processed task instance refers to the task instance indicated by the to-be-processed instance data. The initial value of the instance status of the to-be-processed task instance is the unexecuted state. The to-be-processed task instance includes instance data with an unexecuted state. The task metadata associated with the to-be-processed instance data is data belonging to the task data format converted from task information. The task metadata associated with the to-be-processed instance data is added to the first task database by the scheduling service node. The task data format refers to the format of the data provided by the scheduling service node for the execution service node.
[0216] The result feedback module 22 is configured to feedback the processing status result to the scheduling service node, so that the scheduling service node updates the instance status in the to-be-processed instance data based on the processing status result.
[0217] Wherein, when processing the to-be-processed instance data and the task metadata associated with the to-be-processed instance data to obtain a processing status result for the to-be-processed instance data and the task metadata associated with the to-be-processed instance data, the data processing module 21 may be used to:
[0218] Obtain the task type, the data to be updated, and the task association path from the to-be-processed instance data and the task metadata associated with the to-be-processed instance data;
[0219] Determine the task logic corresponding to the to-be-processed instance data based on the task type, and generate an initial task instruction according to the task logic;
[0220] Update the initial task instruction with the data to be updated and the task association path to obtain a to-be-processed instruction, execute the to-be-processed instruction, and when the execution of the to-be-processed instruction is completed, obtain a processing status result for the to-be-processed instruction.
[0221] Among them, the number of instance data to be processed is H, and H is a positive integer;
[0222] When processing the instance data to be processed and the task metadata associated with the instance data to be processed, and obtaining the processing status result for the instance data to be processed and the task metadata associated with the instance data to be processed, the data processing module 21 can be used for:
[0223] Obtain the first priorities corresponding to the H instance data to be processed respectively from the task metadata associated with the H instance data to be processed;
[0224] Obtain the second priorities corresponding to the H instance data to be processed respectively from the H instance data to be processed;
[0225] Determine the instance processing order of the H instance data to be processed based on the first priorities and the second priorities corresponding to the H instance data to be processed respectively;
[0226] Process the H instance data to be processed in sequence according to the instance processing order, and obtain the processing status result corresponding to the instance data to be processed when the processing of each instance data to be processed is completed.
[0227] An embodiment of the present application provides a task data processing device. The device can run in an execution service node, obtain first task information, convert the first task information into first task metadata belonging to a task data format, and add the first task metadata to a first task database; set the instance state of a first task instance corresponding to a first message task to an unexecuted state, and generate first instance data of the first task instance based on the instance state of the first task instance; the first message task refers to the message task indicated by the first task information; when the scheduling service node issues a task, based on the instance state of the instance data in the first task database, obtain the instance data to be processed from the instance data in the first task database, and send the instance data to be processed and the task metadata associated with the instance data to be processed to the execution service node, so that the execution service node processes the instance data to be processed and the task metadata associated with the instance data to be processed, and feeds back a processing status result for the instance data to be processed and the task metadata associated with the instance data to be processed to the scheduling service node; the instance data in the first task database includes the first instance data, and the instance data to be processed includes the instance data with an instance state of unexecuted state; the processing status result is used to update the instance state in the instance data to be processed. By combining the scheduling service node and the execution service node, the acquisition of task information and the execution of message tasks are decoupled, so as to implement the processing of tasks through different service nodes, enabling the scheduling service node to obtain and schedule task information, and the execution service node to process the message tasks indicated by the task information. The scheduling service node only obtains and schedules task information, which only consumes less resources and time. Even when the data volume of the message task is large, the scheduling service node can meet the requirements for resources such as the acquisition and scheduling of task information of the message task, thereby reducing the situation of task information loss and improving the security of task data processing. At the same time, multiple service nodes coordinate to work, and the instance state is introduced, enabling the scheduling service node to more conveniently obtain the instance data to be processed (i.e., the instance data to be processed), thereby improving the efficiency of task data processing.
[0228] See Figure 12 , Figure 12 is a schematic structural diagram of a computer device provided by an embodiment of the present application. As Figure 12As shown in the figure, the computer device in the embodiment of the present application may include: one or more processors 1201, a memory 1202, and an input / output interface 1203. The processor 1201, the memory 1202, and the input / output interface 1203 are connected through a bus 1204. The memory 1202 is used to store a computer program, and the computer program includes program instructions. The input / output interface 1203 is used to receive and output data, such as for data interaction between the application service node and the second message queue, or for data interaction between the application service node and the scheduling service node, or for data interaction between the scheduling service node and the execution service node, etc.; the processor 1201 is used to execute the program instructions stored in the memory 1202.
[0229] Among them, the processor 1201 is located in the scheduling service node and can perform the following operations:
[0230] Obtain the first task information, convert the first task information into the first task metadata belonging to the task data format, and add the first task metadata to the first task database; the task data format refers to the format of the data provided by the scheduling service node for the execution service node;
[0231] Set the instance state of the first task instance corresponding to the first message task to the unexecuted state, and generate the first instance data of the first task instance based on the instance state of the first task instance; the first message task refers to the message task indicated by the first task information;
[0232] When the scheduling service node issues a task, based on the instance state of the instance data in the first task database, obtain the instance data to be processed from the instance data in the first task database, and send the instance data to be processed and the task metadata associated with the instance data to be processed to the execution service node, so that the execution service node processes the instance data to be processed and the task metadata associated with the instance data to be processed, and feeds back the processing status result for the instance data to be processed and the task metadata associated with the instance data to be processed to the scheduling service node; the instance data in the first task database includes the first instance data, and the instance data to be processed includes the instance data with the instance state of the unexecuted state; the processing status result is used to update the instance state in the instance data to be processed.
[0233] Among them, the processor 1201 is located in the execution service node and can perform the following operations:
[0234] Receive the to-be-processed instance data sent by the scheduling service node and the task metadata associated with the to-be-processed instance data, process the to-be-processed instance data and the task metadata associated with the to-be-processed instance data, and obtain the processing status result for the to-be-processed instance data and the task metadata associated with the to-be-processed instance data; the to-be-processed instance data is generated based on the instance status of the to-be-processed task instance in the task metadata associated with the to-be-processed instance data, and the to-be-processed task instance refers to the task instance indicated by the to-be-processed instance data; the initial value of the instance status of the to-be-processed task instance is the unexecuted state, and the to-be-processed task instance includes instance data with an instance status of unexecuted state; the task metadata associated with the to-be-processed instance data is data belonging to the task data format converted from the task information, and the task metadata associated with the to-be-processed instance data is added by the scheduling service node to the first task database; the task data format refers to the format of the data provided by the scheduling service node for the execution service node.
[0235] Feedback the processing status result to the scheduling service node so that the scheduling service node can update the instance status in the to-be-processed instance data based on the processing status result.
[0236] In some possible implementation manners, the processor 1201 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0237] The memory 1202 may include a read-only memory and a random access memory, and provide instructions and data to the processor 1201 and the input / output interface 1203. A part of the memory 1202 may also include a non-volatile random access memory. For example, the memory 1202 may also store information about the device type.
[0238] In specific implementation, the computer device can execute through its built-in various functional modules as provided by the Figure 4 or Figure 5 in each step, and specifically refer to the implementation manners provided by each step in the Figure 4 or Figure 5 , which will not be elaborated here.
[0239] An embodiment of the present application provides a computer device, including: a processor, an input / output interface, and a memory. The processor obtains a computer program in the memory and executes the Figure 4 or Figure 5 steps of the method shown in, and performs task data processing operations. The embodiment of the present application realizes obtaining first task information, converting the first task information into first task metadata belonging to the task data format, and adding the first task metadata to the first task database; setting the instance status of the first task instance included in the first message task to the unexecuted status, and generating first instance data of the first task instance based on the instance status of the first task instance; the first message task refers to the message task indicated by the first task information; when the scheduling service node issues a task, based on the instance status of the instance data in the first task database, obtaining the instance data to be processed from the instance data in the first task database, and sending the instance data to be processed and the task metadata associated with the instance data to be processed to the execution service node, so that the execution service node processes the instance data to be processed and the task metadata associated with the instance data to be processed, and feeds back the processing status result for the instance data to be processed and the task metadata associated with the instance data to be processed to the scheduling service node; the instance data in the first task database includes the first instance data, and the instance data to be processed includes the instance data with the instance status of the unexecuted status; the processing status result is used to update the instance status in the instance data to be processed. By combining the scheduling service node and the execution service node, the acquisition of task information and the execution of message tasks are decoupled, so that the task can be processed through different service nodes, enabling the scheduling service node to obtain and schedule task information, and the execution service node to process the message task indicated by the task information. The scheduling service node only obtains and schedules task information, which only consumes less resources and time. Even when the data volume of the message task is large, the scheduling service node can meet the requirements of resources such as obtaining and scheduling the task information of the message task, thereby reducing the situation of task information loss and improving the security of task data processing. At the same time, multiple service nodes coordinate their work, and the instance status is introduced, enabling the scheduling service node to more conveniently obtain the instance data to be processed (i.e., the instance data to be processed), thereby improving the efficiency of task data processing.
[0240] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, and the computer program is suitable for being loaded and executed by the processor Figure 4 or Figure 5 the task data processing method provided by each step in, and specifically, reference can be made to the Figure 4 or Figure 5The implementation manners provided in each step are not described herein again. In addition, the beneficial effects of adopting the same method are not described again. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in this application, please refer to the description of the method embodiments of this application. As an example, the computer program can be deployed to be executed on a computer device, or on multiple computer devices located at one place, or on multiple computer devices distributed at multiple places and interconnected through a communication network.
[0241] The computer-readable storage medium may be the task data processing device provided in any of the foregoing embodiments or an internal storage unit of the computer device, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store the data that has been output or will be output.
[0242] The embodiments of this application also provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 4 or Figure 5 the methods provided in the various alternative manners in, realizing the decoupling of the acquisition of task information and the execution of message tasks by combining the scheduling service node and the execution service node, so as to process tasks through different service nodes, enabling the scheduling service node to acquire and schedule task information, and the execution service node to process the message tasks indicated by the task information. The scheduling service node only acquires and schedules task information, and only needs to consume less resources and time. Even when the data volume of the message tasks is large, the scheduling service node can also meet the requirements for resources such as the acquisition and scheduling of task information of the message tasks, thereby reducing the situation of task information loss and improving the security of task data processing. At the same time, multiple service nodes work in coordination, and the instance state is introduced, enabling the scheduling service node to more conveniently acquire the instance data to be processed (i.e., the instance data to be processed), thereby improving the efficiency of task data processing.
[0243] In the description, claims, and drawings of the embodiments of this application, terms such as "first" and "second" are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product, or equipment that includes a series of steps or units is not limited to the listed steps or modules, but may optionally further include unlisted steps or modules, or may optionally further include other step units inherent to these processes, methods, devices, products, or equipment.
[0244] In the embodiments of this application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other relevant parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.
[0245] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in this description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0246] The methods and related devices provided in the embodiments of this application are described with reference to the method flowcharts and / or structural schematic diagrams provided in the embodiments of this application. Specifically, each process and / or block of the method flowchart and / or structural schematic diagram, and the combination of the processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable task data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable task data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable task data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions in the processFigure 1 One process or multiple processes and / or structure schematic Figure 1 The functions specified in one box or multiple boxes. These computer program instructions can also be loaded onto a computer or other programmable task data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 One process or multiple processes and / or structure schematic steps for the functions specified in one box or multiple boxes.
[0247] The steps in the method embodiments of this application can be adjusted, combined, and deleted according to actual needs.
[0248] The modules in the device embodiments of this application can be combined, divided, and deleted according to actual needs.
[0249] The above-disclosed are only the preferred embodiments of this application. Of course, the scope of rights of this application cannot be limited thereby. Therefore, equivalent changes made according to the claims of this application still fall within the scope covered by this application.
Claims
1. A task data processing method, characterized in that, The method is executed by a scheduling service node, and the method includes: Obtain first task information, convert the first task information into first task metadata belonging to a task data format, and add the first task metadata to a first task database; the task data format refers to the format of the data provided by the scheduling service node for an execution service node; Set the instance state of a first task instance corresponding to a first message task to an unexecuted state, and generate first instance data of the first task instance based on the instance state of the first task instance; the first message task refers to the message task indicated by the first task information; When the scheduling service node issues a task, based on the instance state of the instance data in the first task database, obtain to-be-processed instance data from the instance data in the first task database, and send the to-be-processed instance data and the task metadata associated with the to-be-processed instance data to the execution service node, so that the execution service node processes the to-be-processed instance data and the task metadata associated with the to-be-processed instance data, and feeds back a processing status result for the to-be-processed instance data and the task metadata associated with the to-be-processed instance data to the scheduling service node; the instance data in the first task database includes the first instance data, and the to-be-processed instance data includes instance data with an instance state of the unexecuted state; the processing status result is used to update the instance state in the to-be-processed instance data.
2. The method according to claim 1, wherein The obtaining of the first task information includes: Scan a first message queue or a second task database. If the first message queue or the second task database is not empty, obtain first task information from the first message queue or the second task database; the first task information is added to the first message queue or the second task database by an application service node; the first task information is obtained by the application service node through conversion of an initial task message; the data format of the first task information is a scheduling data format, and the scheduling data format is a data format supported by the scheduling service node, and refers to the format of the data provided by the application service node for the scheduling service node.
3. The method according to claim 1, wherein The scheduling service node belongs to a scheduling node cluster; the obtaining of the first task information includes: Obtain to-be-scheduled task information and a task hash value corresponding to the to-be-scheduled task information, and convert the task hash value into a data shard value based on the number of scheduling nodes; the number of scheduling nodes is the number of nodes included in the scheduling node cluster; Obtain a first shard processing range, and determine the to-be-scheduled task information whose data shard value belongs to the first shard processing range as the first task information; Obtain the first task information.
4. The method according to claim 3, wherein The method further includes: When the scheduling service node goes offline, it sends a node offline message to other scheduling service nodes, so that the other scheduling service nodes update their shard processing ranges based on the first shard processing range to obtain a second shard processing range; the other scheduling service nodes refer to the nodes in the scheduling node cluster except the scheduling service node; the second shard processing range is used to represent the range of message tasks processed by the other scheduling service nodes.
5. The method according to claim 1, characterized in that The conversion of the first task information into first task metadata belonging to the task data format includes: Performing type recognition on the first task information to obtain a first task type indicated by the first task information; Obtaining data parameters corresponding to the task data format, and obtaining parameter-associated data from the first task information based on the data parameters; Based on the data parameters, forming the first task metadata from the first task type and the parameter-associated data.
6. The method according to claim 1, wherein The conversion of the first task information into first task metadata belonging to the task data format includes: Converting the first task information into initial task metadata belonging to the task data format; Obtaining a priority determination parameter in the initial task metadata, and determining a task priority corresponding to the first task information based on the priority determination parameter; the priority determination parameter refers to a parameter used for priority determination; Combining the task priority and the initial task metadata to form the first task metadata.
7. The method according to claim 1, characterized in that, When the scheduling service node issues a task, based on the instance status of the instance data in the first task database, obtaining the instance data to be processed from the instance data in the first task database includes: When the scheduling service node issues a task, obtaining the instance status of the instance data included in the first task database, and determining the instance data with the instance status being the unexecuted status or the execution exception status as the initial instance data; Obtaining the number of valid execution nodes of the execution service nodes that have task execution conditions, and obtaining the task execution data volume of the valid execution nodes; Based on the number of valid nodes and the task execution data volume, determining the instance issuance quantity, and obtaining the instance data to be processed from the initial instance data based on the instance issuance quantity.
8. The method according to claim 1, characterized in that, The method further includes: After sending the instance data to be processed and the task metadata associated with the instance data to be processed to the execution service node, updating the instance status in the instance data to be processed to the in-execution status; When receiving the processing status result fed back by the execution service node, updating the instance status in the instance data to be processed based on the processing status result.
9. The method according to claim 1, characterized in that The number of the execution service nodes is M, and M is a positive integer; sending the instance data to be processed and the task metadata associated with the instance data to be processed to the execution service node includes: Based on the node loads and node statuses of the M execution service nodes, determining a target execution service node from the M execution service nodes; the node status of the target execution service node is the node running status. Send the to-be-processed instance data and the task metadata associated with the to-be-processed instance data to the target execution service node.
10. A task data processing method, characterized in that, The method is executed by an execution service node, and the method includes: Receiving the to-be-processed instance data and the task metadata associated with the to-be-processed instance data sent by a scheduling service node, processing the to-be-processed instance data and the task metadata associated with the to-be-processed instance data, and obtaining a processing status result for the to-be-processed instance data and the task metadata associated with the to-be-processed instance data; the to-be-processed instance data is generated based on the instance status of a to-be-processed task instance in the task metadata associated with the to-be-processed instance data, and the to-be-processed task instance refers to the task instance indicated by the to-be-processed instance data; the initial value of the instance status of the to-be-processed task instance is the unexecuted state, and the to-be-processed task instance includes instance data with the instance status of the unexecuted state; the task metadata associated with the to-be-processed instance data is data belonging to a task data format converted from task information, and the task metadata associated with the to-be-processed instance data is added by the scheduling service node to a first task database; the task data format refers to the format of the data provided by the scheduling service node for the execution service node; Feedback the processing status result to the scheduling service node, so that the scheduling service node updates the instance status in the to-be-processed instance data based on the processing status result.
11. The method according to claim 10, wherein The processing the to-be-processed instance data and the task metadata associated with the to-be-processed instance data, and obtaining a processing status result for the to-be-processed instance data and the task metadata associated with the to-be-processed instance data includes: Obtain a task type, data to be updated, and a task association path from the to-be-processed instance data and the task metadata associated with the to-be-processed instance data; Determine the task logic corresponding to the to-be-processed instance data based on the task type, and generate an initial task instruction according to the task logic; Update the initial task instruction with the data to be updated and the task association path to obtain a to-be-processed instruction, execute the to-be-processed instruction, and when the execution of the to-be-processed instruction is completed, obtain a processing status result for the to-be-processed instruction.
12. The method according to claim 10, wherein The number of the to-be-processed instance data is H, and H is a positive integer; The processing the to-be-processed instance data and the task metadata associated with the to-be-processed instance data, and obtaining a processing status result for the to-be-processed instance data and the task metadata associated with the to-be-processed instance data includes: Obtain the first priorities respectively corresponding to the H to-be-processed instance data from the task metadata respectively associated with the H to-be-processed instance data; Obtain the second priorities respectively corresponding to the H to-be-processed instance data from the H to-be-processed instance data; Determine the instance processing order of the H to-be-processed instance data based on the first priorities and the second priorities respectively corresponding to the H to-be-processed instance data; Process the H to-be-processed instance data in sequence according to the described instance processing order. When the processing of each to-be-processed instance data is completed, obtain the processing status result corresponding to the to-be-processed instance data.
13. A task data processing device, characterized in that, The device is applicable to a scheduling service node, and the device includes: A task processing module, configured to obtain first task information, convert the first task information into first task metadata belonging to a task data format, and add the first task metadata to a first task database; the task data format refers to the format of data provided by the scheduling service node for the execution service node; An instance processing module, configured to set the instance status of a first task instance corresponding to a first message task to an unexecuted status, and generate first instance data of the first task instance based on the instance status of the first task instance; the first message task refers to the message task indicated by the first task information; A task distribution module, configured to, when the scheduling service node distributes tasks, obtain to-be-processed instance data from the instance data in the first task database based on the instance status of the instance data in the first task database, and send the to-be-processed instance data and the task metadata associated with the to-be-processed instance data to the execution service node, so that the execution service node processes the to-be-processed instance data and the task metadata associated with the to-be-processed instance data, and feeds back a processing status result for the to-be-processed instance data and the task metadata associated with the to-be-processed instance data to the scheduling service node; the instance data in the first task database includes the first instance data, and the to-be-processed instance data includes instance data with an instance status of the unexecuted status; the processing status result is used to update the instance status in the to-be-processed instance data.
14. A task data processing device, characterized in that, The device is applicable to an execution service node, and the device includes: A data processing module, configured to receive the to-be-processed instance data and the task metadata associated with the to-be-processed instance data sent by the scheduling service node, process the to-be-processed instance data and the task metadata associated with the to-be-processed instance data, and obtain a processing status result for the to-be-processed instance data and the task metadata associated with the to-be-processed instance data; the to-be-processed instance data is generated based on the instance status of a to-be-processed task instance in the task metadata associated with the to-be-processed instance data, and the to-be-processed task instance refers to the task instance indicated by the to-be-processed instance data; the initial value of the instance status of the to-be-processed task instance is the unexecuted status, and the to-be-processed task instance includes instance data with an instance status of the unexecuted status; the task metadata associated with the to-be-processed instance data is data belonging to the task data format converted from task information, and the task metadata associated with the to-be-processed instance data is added to the first task database by the scheduling service node; the task data format refers to the format of data provided by the scheduling service node for the execution service node; A result feedback module, configured to feedback the processing status result to the scheduling service node, so that the scheduling service node updates the instance status in the to-be-processed instance data based on the processing status result.
15. A computer device, characterized in that, It includes a processor, a memory, and an input / output interface; The processor is respectively connected to the memory and the input / output interface. Among them, the input / output interface is used to receive and output data, the memory is used to store computer programs, and the processor is used to call the computer programs so that the computer device executes the method described in any one of claims 1-9, or executes the method described in any one of claims 10-12.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which is adapted to be loaded and executed by a processor, so that a computer device having the processor executes the method described in any one of claims 1-9, or executes the method described in any one of claims 10-12.
17. A computer program product, comprising computer instructions, characterized in that, When the computer instructions are executed by a processor, the method described in any one of claims 1-9 is implemented, or the method described in any one of claims 10-12 is executed.