Data task arrangement method and device

By generating and analyzing data task nodes and assigning them to task queues, the problems of cumbersome task orchestration operations and lack of flexibility in dependency management in the existing technology are solved, and efficient and stable distributed data processing is achieved.

CN119938252APending Publication Date: 2025-05-06FENGTAI SCI & TECH (BEIJING) CO LTD
View PDF -1 Cites 2 Cited by

Patent Information

Application Number
CN202411813107.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, the task orchestration method of distributed data tasks is cumbersome to operate, it is difficult to adapt to task changes in dynamic environments, and task execution dependency management lacks flexibility, making it difficult to deal with complex task links and conditional branches.

Method used

By responding to the user's task node creation operation of the distributed processing system, multiple data task nodes are generated, and the node type of each data task node is parsed, and it is allocated to the corresponding task queue. The execution order of the data task nodes in the distributed processing system is determined based on the task queue.

Benefits of technology

It improves the flexibility of distributed data task orchestration, simplifies operational processes, improves user experience, and realizes efficient and stable operation of distributed processing systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938252A_ABST
    Figure CN119938252A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of distributed data processing, and provides a data task arrangement method and device.According to the method, multiple data task nodes are generated by responding to task node creation operation of a user on a distributed processing system, and the data format of the data task nodes is in a JavaScript object numbered musical notation JSON form; analyzing each data task node to obtain the node type of each data task node; according to the node type of each data task node, each data task node is allocated to a corresponding task queue, and each task queue corresponds to a different node type; and determining an execution sequence of the plurality of data task nodes in the distributed processing system according to the task queue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of distributed data processing technology, and more specifically, to a data task scheduling method and device. Background Art

[0002] In the existing technology, there are still many shortcomings when using distributed processing systems for data governance and data modeling. For example, the task orchestration method of distributed data tasks such as data governance and data modeling is not flexible enough, most of which require manual configuration and are difficult to adapt to task changes in dynamic environments; task execution dependency management is often relatively static and lacks flexibility, making it difficult to handle complex task links and conditional branches; and cumbersome operations lead to poor user experience.

[0003] The problem of cumbersome operation of the task arrangement method of distributed data tasks in the above-mentioned prior art directly affects the application effect of the distributed processing system in complex business scenarios, making it difficult to achieve efficient and stable operation. Summary of the invention

[0004] The purpose of the embodiments of the present application is to provide a data task scheduling method and device, aiming to solve the technical problem of cumbersome operation of the task scheduling method of distributed data tasks in the prior art.

[0005] To achieve the above object, according to a first aspect of the present application, a data task scheduling method is provided, the method comprising:

[0006] In response to the user's task node creation operation on the distributed processing system, a plurality of data task nodes are generated, wherein the data format of the data task nodes is in the JavaScript object notation JSON format;

[0007] Parsing each of the data task nodes to obtain the node type of each of the data task nodes;

[0008] According to the node type of each data task node, each data task node is assigned to a corresponding task queue, wherein each task queue corresponds to a different node type;

[0009] The execution order of the plurality of data task nodes in the distributed processing system is determined according to the task queue.

[0010] Optionally, in a possible implementation manner of the first aspect, in response to a user creating an operation on a task node of the distributed processing system, generating a plurality of the data task nodes includes:

[0011] In response to the task node creation operation performed by the user through the business configuration interface of the distributed processing system, a plurality of the data task nodes are generated, wherein different business configuration interfaces correspond to different business scenario requirements.

[0012] Optionally, in a possible implementation manner of the first aspect, parsing each of the data task nodes to obtain the node type of each of the data task nodes includes:

[0013] A task scheduling rule engine is used to parse each of the data task nodes to obtain the node type of each of the data task nodes, wherein the task scheduling rule engine is used to define the node type and node priority corresponding to each data task node, and the node type includes: source type, handler type, and aggregation type.

[0014] Optionally, in a possible implementation of the first aspect, allocating each of the data task nodes to a corresponding task queue according to the node type of each of the data task nodes includes:

[0015] Determine, according to the task queue corresponding to the node type of each data task node, a mapping relationship between each data task node and a predecessor node list in the corresponding task queue;

[0016] Based on the mapping relationship, each data task node is stored in the corresponding task queue with the node identifier of each data task node as the key and the predecessor node list corresponding to each data task node as the value.

[0017] Optionally, in a possible implementation manner of the first aspect, determining the execution order of the plurality of data task nodes in the distributed processing system according to the task queue includes:

[0018] If at least two of the plurality of data task nodes are in different task queues, respectively, determining the node priority of each data task node in each task queue;

[0019] If at least two of the plurality of data task nodes are in the same task queue, determining a dependency relationship between the at least two data task nodes in the same task queue;

[0020] The execution order of the multiple data task nodes in the distributed processing system is determined according to the node priorities and / or dependencies of the multiple data task nodes.

[0021] Optionally, in a possible implementation manner of the first aspect, if at least two of the plurality of data task nodes are in different task queues, respectively, determining the node priority of each data task node in each task queue includes:

[0022] If at least two of the plurality of data task nodes are in different task queues, respectively, determining the node types corresponding to the different task queues, wherein the node type includes at least one of the following: a source type, a processing program type, and a convergence type;

[0023] If the node type corresponding to the first task queue in the different task queues is a source type, the node priority of the data task node in the first task queue is the first priority;

[0024] If the node type corresponding to the second task queue in the different task queues is a processing program type, the node priority of the data task node in the second task queue is a second priority, wherein the second priority is less than the first priority;

[0025] If the node type corresponding to the third task queue in the different task queues is the aggregation type, the node priority of the data task node in the third task queue is the third priority, wherein the third priority is lower than the second priority.

[0026] Optionally, in a possible implementation manner of the first aspect, if at least two of the plurality of data task nodes are in the same task queue, determining the dependency relationship between the at least two data task nodes in the same task queue includes:

[0027] If at least two of the plurality of data task nodes are in the same task queue, determining whether there is a predecessor node in the at least two data task nodes in the same task queue;

[0028] If there is one of the predecessor nodes among at least two of the data task nodes, determining that a non-predecessor node among at least two of the data task nodes depends on the predecessor node;

[0029] If there are two predecessor nodes in at least two of the data task nodes, determining a dependency relationship between the two predecessor nodes based on the creation order of the two predecessor nodes;

[0030] If the predecessor node does not exist in at least two of the data task nodes, the dependency relationship between the two predecessor nodes is determined based on the creation order of the two predecessor nodes.

[0031] Optionally, in a possible implementation manner of the first aspect, determining the dependency relationship between the two predecessor nodes based on the creation order of the two predecessor nodes includes:

[0032] If the creation order of the two predecessor nodes is different, determining a predecessor node that is created first among the two predecessor nodes as the dependent predecessor node, wherein the creation order is determined based on an initialization timestamp or a node identifier size;

[0033] If the creation orders of the two predecessor nodes are different, then according to the size of the node identifier of each predecessor node, a predecessor node with a smaller node identifier will be determined as the dependent predecessor node.

[0034] Optionally, in a possible implementation manner of the first aspect, after determining the execution order of the plurality of data task nodes in the distributed processing system according to the task queue, the method further includes:

[0035] Using a predefined dependency enqueue method, obtain a predecessor node list corresponding to each data task node in the task queue, wherein each predecessor node list is used to store at least one predecessor node of the data task node;

[0036] Detecting the status of the identification bit of each predecessor node in the predecessor node list;

[0037] If the identification bit status of each of the predecessor nodes in the predecessor node list is a true value, it is determined that each of the predecessor nodes has been processed, and the data task node is added to the priority queue;

[0038] According to the execution order of the multiple data task nodes, the multiple data task nodes are taken out from the priority queue in sequence and executed.

[0039] Optionally, in a possible implementation manner of the first aspect, the method further includes:

[0040] In determining the execution order of the plurality of data task nodes, or in the process of executing the plurality of data task nodes, corresponding log records are generated.

[0041] According to a second aspect of the present application, a data task scheduling device is provided, the device comprising:

[0042] A generating unit, for generating a plurality of data task nodes in response to a user's task node creation operation on a distributed processing system, wherein the data format of the data task nodes is a JavaScript object notation JSON format;

[0043] A parsing unit, used for parsing each of the data task nodes to obtain the node type of each of the data task nodes;

[0044] An allocating unit, configured to allocate each of the data task nodes to a corresponding task queue according to the node type of each of the data task nodes, wherein each of the task queues corresponds to a different node type;

[0045] A sorting unit is used to determine the execution order of multiple data task nodes in the distributed processing system according to the task queue.

[0046] The second aspect and any implementation of the second aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the second aspect and any implementation of the second aspect can refer to the technical effects corresponding to the first aspect and any implementation of the first aspect, which will not be repeated here.

[0047] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the electronic device implements any of the methods described in one of the embodiments.

[0048] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements any of the methods described above.

[0049] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product is run on an electronic device, the electronic device executes any one of the methods described in the first aspect.

[0050] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.

[0051] An embodiment of the present application provides a data task orchestration method and device, which generates multiple data task nodes in response to a user's task node creation operation on a distributed processing system, wherein the data format of the data task node is in the JavaScript object notation JSON format; parses each data task node to obtain the node type of each data task node; assigns each data task node to a corresponding task queue according to the node type of each data task node, wherein each task queue corresponds to a different node type; and determines the execution order of multiple data task nodes in the distributed processing system according to the task queue.

[0052] Through this data task orchestration method, the problems of insufficient flexibility in data task orchestration, lack of flexibility in task execution dependency management, and cumbersome operation in the prior art are solved, thereby improving the application effect of the distributed processing system in complex business scenarios, and achieving efficient and stable operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0054] Figure 1 It is a flowchart of a data task scheduling method provided in an embodiment of the present application;

[0055] Figure 2 It is a flowchart of an optional data task scheduling method provided in an embodiment of the present application;

[0056] Figure 3 It is a flowchart of an optional data task scheduling method provided in an embodiment of the present application;

[0057] Figure 4 It is a flowchart of an optional data task scheduling method provided in an embodiment of the present application;

[0058] Figure 5 is a flowchart of an optional data task scheduling method provided by another embodiment of the present application;

[0059] Figure 6 It is a structural diagram of a data task scheduling device provided in an embodiment of the present application;

[0060] Figure 7 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0061] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.

[0062] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.

[0063] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0064] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.

[0065] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0066] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0067] This application example provides an example of a data task orchestration method, please refer to Figure 1 As shown, Figure 1 A schematic flow chart of a data task scheduling method provided by the present application is shown. As an example but not a limitation, the method can be applied to or run in an electronic device. The method includes:

[0068] S101, in response to a user's operation of creating a task node of a distributed processing system, a plurality of data task nodes are generated, wherein the data format of the data task nodes is in the form of JavaScript object notation JSON.

[0069] S102, parsing each data task node to obtain the node type of each data task node.

[0070] S103, allocating each data task node to a corresponding task queue according to the node type of each data task node, wherein each task queue corresponds to a different node type.

[0071] S104: determining the execution order of the multiple data task nodes in the distributed processing system according to the task queue.

[0072] In the present application example, when the user performs a task node creation operation on the distributed processing system, the distributed processing system responds to this operation and generates multiple data task nodes, and the data format of the data task node is the JavaScript object notation JSON format. The JSON data format is concise, easy to understand and process, and can clearly describe the relevant information of each data task node. In this way, the user's operation of creating a task node is converted into a data task node presented in JSON format, so that subsequent automated processing can be performed based on these standardized formats, and no longer completely relying on manual configuration, which improves the flexibility of task scheduling, and can dynamically generate different data task nodes according to actual needs, and can better adapt to task changes in a dynamic environment.

[0073] Next, by parsing multiple data task nodes, the node type of each data task node is obtained. By parsing the data task nodes in JSON format, the key attributes of each data task node, namely the node type, can be accurately extracted, so that the distributed processing system can clearly understand the essential characteristics of each data task node.

[0074] According to the node types of the determined multiple data task nodes, these task nodes are assigned to the corresponding task queues, and each task queue corresponds to different node types. This allocation method allows data task nodes with the same or similar characteristics (divided by node type) to be classified into the same queue, thereby realizing the classified management of data task nodes. When processing complex task links and conditional branches, the execution order and processing method of tasks can be flexibly arranged according to the conditions of different task queues, no longer based on the more static dependency management mode in traditional technology, effectively improving the flexibility of task execution dependency management.

[0075] Finally, the execution order of multiple data task nodes in the distributed processing system is determined according to the task queue. Users no longer need to perform tedious manual operations such as setting the execution order, which greatly simplifies the operation process and improves the user experience.

[0076] Through the data task scheduling method, multiple data task nodes are generated in response to the user's task node creation operation on the distributed processing system, and each data task node is parsed to obtain the node type of each data task node. Based on the node type, each data task node is assigned to a corresponding task queue, and finally the execution order of multiple data task nodes in the distributed processing system is determined according to the task queue. The problems of inflexible data task scheduling, lack of flexibility in task execution dependency management, and cumbersome operation in the prior art are effectively solved, thereby improving the application effect of the distributed processing system in complex business scenarios, and further achieving efficient and stable operation.

[0077] In a possible implementation, in response to a user's operation of creating a task node of a distributed processing system, multiple data task nodes are generated, including:

[0078] In response to a task node creation operation performed by a user through a business configuration interface of the distributed processing system, a plurality of data task nodes are generated, wherein different business configuration interfaces correspond to different business scenario requirements.

[0079] In the above implementation, multiple data task nodes corresponding to the task node creation operation performed by the user through the specific business configuration interface of the distributed processing system are generated. It can better adapt to the needs of different business scenarios, so that the subsequent data task arrangement and processing can be more in line with the actual business situation, thereby improving the flexibility and effectiveness of the entire distributed processing system in data governance and data modeling.

[0080] It should be understood that distributed processing systems usually provide users with a series of interfaces to facilitate users to configure and customize related data processing tasks according to their own business needs. The business configuration interface is specially designed for the needs of different business scenarios, and different interfaces often correspond to different business functions or processing flows. Specifically, users can operate through the business configuration interface provided by the distributed processing system, and the specific operation content is to create data task nodes. For example, in an e-commerce business scenario, there may be a business configuration interface for product management, which is used to create task nodes related to product data processing, such as task nodes for product information updates, inventory management, etc.; in the logistics business scenario, there will be business configuration interfaces involving order tracking, cargo delivery, etc., which are used to create corresponding task nodes.

[0081] When the user performs a task node creation operation through a specific business configuration interface, the system will create a data task node in JSON format. As a lightweight data exchange format, JSON is concise, easy to understand and process. By converting each data task node into JSON format, the various properties and related information of the data task node can be clearly described, such as the node type of the data task node (input node of SOURCE type, processing node of HANDLER type, output node of SINK type, etc.), the data fields involved, processing logic, and the relationship with other task nodes. These data task nodes presented in JSON format can be easily parsed and processed by related components such as the rule engine.

[0082] Based on the above implementation method, by closely combining the user's task node creation operation on a specific business configuration interface with the generation of corresponding JSON format data task nodes, and taking into account the differences in requirements of different business scenarios, a flexible and effective method is provided for subsequent data task orchestration and processing, which helps to solve some problems faced in the existing technology when using distributed processing systems for data governance and data modeling.

[0083] It should be noted that the present application method can be applied to but is not limited to any of the following fields: Financial field: real-time analysis and processing of large amounts of transaction data, automatic orchestration of complex risk control tasks, such as automatic adjustment of real-time risk control rules, dynamic deployment of risk control models, etc.; Internet industry: In large-scale Internet platforms, real-time monitoring systems can quickly integrate data streams through flexible task orchestration, and provide real-time alarms when the system is abnormal to ensure service stability; Internet of Things field: processing large amounts of data streams from sensors, cameras and other devices, real-time orchestration of traffic signals, emergency management and other tasks, and efficient management of smart cities; and any other industry fields that require the orchestration of task flows.

[0084] In a possible implementation, each data task node is parsed to obtain the node type of each data task node, including:

[0085] A task scheduling rule engine is used to parse each data task node to obtain the node type of each data task node. The task scheduling rule engine is used to define the node type and node priority corresponding to each data task node. The node types include: source type, handler type, and aggregation type.

[0086] In the above possible implementation, the task scheduling rule engine is used to parse the JSON data corresponding to the data task node, so as to accurately obtain the node type of each data task node. In addition, the task scheduling rule engine also bears the important responsibility of defining the relevant attributes of each data task node, such as node type and node priority, so as to facilitate the subsequent data task scheduling and processing flow.

[0087] It should be understood that in the example of this application, the task scheduling rule engine is a software component or module specially designed to handle matters related to data task scheduling, which has the ability to parse JSON data and can read and understand information about each data task node stored in JSON format. Through in-depth analysis of these JSON data, various attribute information required for subsequent processing, such as node type, can be extracted.

[0088] For example, in a complex distributed data processing system, there may be a large number of data task nodes with different functions and features, and these data task nodes describe their various situations in JSON format. The task scheduling rule engine is like an intelligent "parser" that can analyze each data task node in an orderly manner, thereby providing an accurate basis for subsequent processing steps.

[0089] When the task scheduling rule engine parses each data task node, it will identify and extract the node type information of each task node from the data task node in JSON format according to the established rules and logic. In an optional embodiment, the node types include: source type SOURCE, handler type HANDLER, and aggregation type SINK. For example, the SOURCE node type usually serves as the source of data and is responsible for receiving data sources from the outside, just like the source of water flow, providing raw data for subsequent data processing; the HANDLER node type acts as a handler, mainly responsible for various processing or conversion operations on the data, such as cleaning, processing, and analyzing the data; the SINK node type can be regarded as a convergence point, and its main responsibility is to output the processed good data to the target system to complete the last step of the data processing process.

[0090] In addition to being able to parse the data task nodes in JSON format to obtain node type information, the task scheduling rule engine can also, but is not limited to, define the node type and node priority corresponding to each data task node. That is, during the design and operation of the entire data processing system, the engine can also actively specify the type that each task node should have and its priority in the processing flow. For example, it is clearly stipulated that the SOURCE node type has a higher priority in the entire task process, because as the data source, if it is not executed first, the subsequent nodes may not be able to obtain the required original data; the HANDLER node type has a lower priority than the SOURCE node type, and the HANDLER node type can only be effectively processed after the SOURCE node provides the original data; the SINK node type has the lowest priority, because the SINK node type will output the processed data only after the previous SOURCE and HANDLER nodes have completed the corresponding work.

[0091] Through this implementation method, the task scheduling rule engine establishes a clear rule system for the entire data processing process, so that each data task node can be arranged and processed in a reasonable order and priority, thereby ensuring the efficiency, accuracy and stability of data processing.

[0092] Based on this possible implementation method, by using the task scheduling rule engine to parse the data task nodes, the node type information can be accurately obtained. With the help of the engine's definition function of node type and node priority, a solid foundation is provided for subsequent data task scheduling and processing, which helps to realize efficient and stable data processing processes under distributed processing systems.

[0093] In a possible implementation, each data task node is assigned to a corresponding task queue according to its node type, including:

[0094] According to the task queue corresponding to the node type of each data task node, determine the mapping relationship between each data task node and the predecessor node list in the corresponding task queue;

[0095] Based on the mapping relationship, each data task node is stored in the corresponding task queue with the node identifier of each data task node as the key and the predecessor node list corresponding to each data task node as the value.

[0096] In this possible implementation, data task nodes are accurately assigned to corresponding task queues based on their respective node types. This process involves first determining the mapping relationship between each data task node and the list of predecessor nodes in the corresponding task queue, and then using the mapping relationship to store each data task node in the corresponding task queue in the form of a specific key-value pair, so as to standardize the order of data task nodes in the queue and facilitate subsequent task processing.

[0097] First, different node types (such as source type SOURCE, handler type HANDLER, sink type SINK) have their corresponding task queues. This is determined based on the previous classification of node types and the design of the entire data processing architecture. For example, the task queue corresponding to the SOURCE node type is the sourceQueue queue, the task queue corresponding to the HANDLER node type is the handlerQueue queue, and the task queue corresponding to the SINK node type is the sinkQueue queue.

[0098] Then, clarify the relationship between each data task node and the list of predecessor nodes in the task queue where it is located. A predecessor node refers to a node that precedes the data task node in the task execution order and has an impact on the execution of the data task node. By determining this mapping relationship, it is clear how each data task node depends on other nodes in its corresponding task queue, which facilitates the accurate arrangement of the execution order of multiple data task nodes. For example, in a specific business scenario, for a node of a HANDLER type, by analyzing its business logic and data dependencies, it is determined that the list of predecessor nodes in the handlerQueue queue may contain certain SOURCE type nodes and other specific HANDLER node types.

[0099] Furthermore, once the above mapping relationship is determined, each data task node is stored in the corresponding task queue in the form of a specific key-value pair. Specifically, the node identifier of each data task node can be selected as the key. The node identifier can be information such as the node ID number that can uniquely identify the node; the value is the predecessor node list corresponding to the data task node determined previously. For example, suppose there is a data task node with a node identifier of "Node101", which is a HANDLER node type. After the previous steps, it is determined that the predecessor node list of the data task node in the handlerQueue queue is ["SourceNode001", "HandlerNode020"]. Then, with "Node101" as the key and ["SourceNode001", "HandlerNode020"] as the value, this data task node is stored in the handlerQueue queue.

[0100] Through this key-value pair storage method, each data task node and its dependent predecessor node can be clearly seen in the corresponding task queue, so that when processing multiple data task nodes later, tasks can be executed in order according to the dependency relationship of the data task nodes. For example, when a data task node is to be executed, it can be checked first whether the predecessor nodes of the data task node have been executed, so as to ensure the accuracy and consistency of task execution.

[0101] By adopting this possible implementation method, by first determining the mapping relationship between the node and the predecessor node list, and then storing each data task node in the corresponding task queue in the form of a key-value pair, the data task nodes are effectively arranged and managed, which helps to achieve the orderly progress of the entire data processing process and improve the efficiency and reliability of data task processing in a distributed processing system.

[0102] For a possible implementation, please refer to Figure 2 As shown, Figure 2 It is a flowchart of an optional data task scheduling method provided in an embodiment of the present application, which determines the execution order of multiple data task nodes in a distributed processing system according to a task queue, including:

[0103] S201, if at least two of the multiple data task nodes are in different task queues, respectively, determining the node priority of each data task node in each task queue;

[0104] S202, if at least two of the multiple data task nodes are in the same task queue, determining a dependency relationship between the at least two data task nodes in the same task queue;

[0105] S203, determining the execution order of the multiple data task nodes in the distributed processing system according to the node priorities and / or dependencies of the multiple data task nodes.

[0106] In this possible implementation, determining the execution order of multiple data task nodes in a distributed processing system needs to be considered on a case-by-case basis. When data task nodes are distributed in different task queues, the execution order is determined based on the priority of the nodes in each task queue; when there are at least two data task nodes in the same task queue, the execution order is determined by analyzing the dependencies between these data task nodes. Finally, the node priorities and / or dependencies are combined to derive the complete execution order of all data task nodes to ensure that the data processing process can be carried out in an orderly and efficient manner in a distributed processing system.

[0107] Different task queues usually correspond to different types of data task nodes. For example, the data task nodes of the source type SOURCE mentioned above are placed in the sourceQueue queue, the data task nodes of the handler type HANDLER are placed in the handlerQueue queue, and the data task nodes of the aggregation type SINK are placed in the sinkQueue queue, etc.

[0108] Each data task node in the task queue has a corresponding node priority. For example, the corresponding priority is set based on the logic of the data processing flow and the role of each node type in the entire process. Specifically, the SOURCE node type, as the source of data, generally has a higher priority, because the processing of subsequent data task nodes often depends on the original data provided by the SOURCE type node; the HANDLER node type has the second highest priority and is mainly responsible for processing and converting data, and needs to start work after the SOURCE node provides data; the SINK node type has a relatively low priority and is to output the processed data task, and needs to wait for the aforementioned SOURCE and HANDLER nodes to complete the corresponding operations.

[0109] When there are at least two data task nodes in different task queues, the node priorities of these data task nodes in each task queue must be clarified first. For example, on a factory production line, different workshops (corresponding to different task queues) produce different parts (corresponding to different node types), and the production of parts in each workshop has its sequence (corresponding to node priority). By determining these sequences, it can provide a basis for subsequent cross-workshop coordination of production (determining the execution order of nodes in different queues).

[0110] If at least two of the multiple data task nodes are in the same task queue, the order of execution cannot be determined simply by node priority at this time, because they are at the same priority level (the same task queue), therefore, it is necessary to deeply analyze the specific dependencies between these data task nodes. It should be understood that in the example of the present application, dependency refers to whether the execution of a data task node depends on the execution result of another data task node. For example, in a data processing task, a HANDLER node type may need to rely on another HANDLER node type to perform specific preprocessing on the data before it can perform subsequent processing operations. This dependency can be determined by analyzing factors such as the business logic corresponding to the node and the data flow.

[0111] Finally, the execution order of all data task nodes in the distributed processing system is finally determined by comprehensively considering the information determined in the previous two situations, namely, the priority of nodes in different task queues and the dependency relationship between nodes in the same task queue. For example, the execution order between different queues is first arranged according to the priority order of the task queues (determined by the priority of the node type corresponding to each queue), that is, the tasks of the queue where the SOURCE node type is located are executed first, then the tasks of the queue where the HANDLER node type is located, and finally the tasks of the queue where the SINK node type is located. Then, within each queue, the execution order is further refined according to the dependency relationship between the nodes to ensure that each data task node is executed at the right time, avoiding problems such as processing before the data is ready, thereby ensuring the smooth progress of the entire data processing process.

[0112] The above possible implementation method considers the situations of data task nodes in different or the same task queues respectively, determines the node priority and dependency respectively, and then combines these factors to accurately determine the execution order of multiple data task nodes in the distributed processing system, which helps to achieve efficient and orderly data processing operations.

[0113] For a possible implementation, please refer to Figure 3 As shown, Figure 3 It is a flow chart of an optional data task scheduling method provided by an embodiment of the present application. If at least two data task nodes among multiple data task nodes are in different task queues, the node priority of each data task node in each task queue is determined, including:

[0114] S301, if at least two of the multiple data task nodes are in different task queues, determine the node types corresponding to the different task queues, wherein the node type includes at least one of the following: source type, processing program type, and aggregation type.

[0115] S302: If the node type corresponding to the first task queue among the different task queues is a source type, the node priority of the data task node in the first task queue is a first priority.

[0116] S303: If the node type corresponding to the second task queue in the different task queues is a processing program type, the node priority of the data task node in the second task queue is a second priority, wherein the second priority is lower than the first priority.

[0117] S304: If the node type corresponding to the third task queue among the different task queues is the aggregation type, the node priority of the data task node in the third task queue is the third priority, wherein the third priority is lower than the second priority.

[0118] In this possible implementation, when there are multiple data task nodes located in different task queues, the node type corresponding to each task queue is determined, and then the node priority of the data task node in each task queue is clarified according to the different node types. Specifically, according to the different roles and order of several common node types such as source type, processing program type, and aggregation type in the data processing flow, the corresponding priorities are set for the nodes in different task queues to ensure that data processing can be carried out efficiently in a reasonable order.

[0119] First, when it is found that there are at least two data task nodes in different task queues, it is necessary to analyze these different task queues to determine the node types they correspond to. In the data processing architecture, nodes are usually divided into different types according to their functions in the entire process. For example, the source type (SOURCE) node is mainly responsible for receiving external data sources and is the starting point for data to enter the system; the handler type (HANDLER) node is responsible for various processing and conversion operations on the data; the sink type (SINK) node outputs the processed data to the target system and is the end point of the data processing process.

[0120] Since each task queue is generally used to store a specific type of node, it is possible to clearly identify the node types corresponding to different task queues, whether they are source types, handler types, or sink types. That is, there is one task queue dedicated to storing source node types, another task queue dedicated to storing handler node types, and another task queue dedicated to storing sink node types.

[0121] Further, once the node type corresponding to the task queue is determined, if the node type corresponding to one of the task queues (the first task queue) is a source type (SOURCE), the node priority of all data task nodes in the first task queue is set to the first priority. This is because the source node type plays a vital role in the entire data processing process. It is the source of the data, and all subsequent processing and operations depend on the original data it provides. If the source node type cannot execute and provide data first, then other types of nodes will not be able to work properly, so it should have the highest priority, which is called the first priority here.

[0122] Next, if the node type corresponding to another task queue (second task queue) is a handler type (HANDLER), then the node priority of the data task node in the second task queue is set to the second priority. The handler node type is mainly responsible for processing and converting data. It can only work on the basis that the source node type has provided the original data, so its priority is lower than the priority of the source node type, that is, the second priority is lower than the first priority.

[0123] Finally, if there is another task queue (the third task queue) whose corresponding node type is a SINK, the node priority of the data task node in the third task queue is set to the third priority. The main responsibility of the sink node type is to output the processed data to the target system. It needs to wait for the previous source node type to provide the original data and the processing program node type to complete the data processing and conversion operations before it can start working, so its priority is the lowest, that is, the third priority is lower than the second priority.

[0124] By setting the node priority of data task nodes in different task queues according to the node type, it can be ensured that in the data processing flow, the data first flows in from the source node type, is processed and converted by the processor node type, and is finally output by the aggregation node type. The whole process is carried out in an orderly manner according to a reasonable order, thereby improving the efficiency and accuracy of data processing.

[0125] For a possible implementation, please refer to Figure 4 As shown, Figure 4 It is a flow chart of an optional data task scheduling method provided by an embodiment of the present application. If at least two data task nodes among multiple data task nodes are in the same task queue, the dependency relationship between at least two data task nodes in the same task queue is determined, including:

[0126] S401, if at least two data task nodes among the multiple data task nodes are in the same task queue, determine whether there is a predecessor node among the at least two data task nodes in the same task queue.

[0127] S402: If there is a predecessor node among the at least two data task nodes, determine whether a non-predecessor node among the at least two data task nodes depends on the predecessor node.

[0128] S403: If there are two predecessor nodes in at least two data task nodes, determine the dependency relationship between the two predecessor nodes based on the creation order of the two predecessor nodes.

[0129] S404: If there is no predecessor node in at least two data task nodes, determine the dependency relationship between the two predecessor nodes based on the creation order of the two predecessor nodes.

[0130] In the above possible implementations, when multiple data task nodes are in the same task queue, in order to determine the execution order between these data task nodes, it is necessary to deeply analyze the dependencies between them. By judging whether there is a predecessor node and the different situations when there are multiple predecessor nodes, the dependencies between the data task nodes can be clarified, so as to facilitate the subsequent reasonable arrangement of the execution order of the data task nodes in the same task queue.

[0131] When it is found that at least two data task nodes are in the same task queue, first check whether there is a predecessor node in at least two data task nodes. Specifically, the predecessor node here refers to a node that is executed before other nodes in the task execution order and whose execution result will affect the subsequent nodes. For example, in a data processing task, node A needs to perform some preprocessing on the data before node B can continue the subsequent processing based on the result of node A's processing, then node A is the predecessor node of node B.

[0132] By analyzing each data task node in the same task queue and checking their business logic, data flow and other aspects, it is determined whether there is such a predecessor node relationship. If it is found through detection that there is a predecessor node in at least two data task nodes, it is determined that the non-predecessor node is dependent on this predecessor node. In other words, in order for the non-predecessor node to perform its task normally, it must wait for the predecessor node to complete the execution first and provide the corresponding processing results. For example, in a task queue for data cleaning and analysis, there is a node C responsible for performing preliminary cleaning operations on the original data (predecessor node), and then node D will perform data analysis operations based on the clean data cleaned by node C (non-predecessor node). Then node D depends on node C. Only when node C completes the cleaning task first can node D smoothly carry out data analysis work.

[0133] When it is determined that there are two predecessor nodes in at least two data task nodes, the situation is relatively complicated. At this time, it is necessary to determine the dependency between the two predecessor nodes based on the order in which they are created. Generally speaking, the creation order can be determined by the timestamp when the node is created or the identification information such as the node ID. If the creation time of a predecessor node is earlier than that of another predecessor node (or its ID is earlier under a certain sorting rule), it can be considered that the predecessor node created first may have a certain priority in the execution order. For example, node E and node F are both predecessor nodes in a task queue, and node E is created earlier, then node E needs to be executed first, and then node F, because node E created first may have prepared the corresponding data or completed certain pre-operations, which can ensure that the subsequent nodes can smoothly obtain the required data.

[0134] Even if there is no obvious predecessor node relationship in the same task queue (that is, all nodes do not seem to be directly dependent on other nodes to be executed first to some extent), it is still necessary to determine the dependency relationship between the nodes based on the order in which they are created. The timestamp when the node was created or the node ID and other identification information are also used to make the judgment. According to the order of creation, the node created first may be given a certain priority in the execution order. For example, in a task queue, there are nodes G and H, and there is no obvious predecessor node relationship between them. However, if the creation time of node G is earlier than that of node H, then when arranging the execution order, node G may be considered for execution first, and then node H, so as to ensure the orderliness and consistency of task execution.

[0135] Through the above analysis and processing of different situations in the same task queue, the dependency relationship between at least two data task nodes can be accurately determined, thereby providing a basis for reasonably arranging the execution order of these nodes in the same task queue, and helping to achieve efficient and orderly data processing flow in the same task queue.

[0136] In a possible implementation, determining the dependency relationship between two predecessor nodes based on the creation order of the two predecessor nodes includes:

[0137] If the creation order of the two predecessor nodes is different, a predecessor node that is created first among the two predecessor nodes is determined as the dependent predecessor node, wherein the creation order is determined based on the initialization timestamp or the node identifier size.

[0138] If the creation order of two predecessor nodes is different, then according to the node identifier size of each predecessor node, the predecessor node with a smaller node identifier will be determined as the dependent predecessor node.

[0139] In this implementation, when the dependency relationship between two predecessor nodes needs to be determined based on the creation order of the two predecessor nodes, the two factors are mainly based on the initialization timestamp or the size of the node identifier. By comparing the order of the creation time of the two predecessor nodes or the size relationship of the node identifiers, it is clear which predecessor node is dependent, which is convenient for the subsequent reasonable arrangement of the task execution order.

[0140] If the creation order of two predecessor nodes is different, the initialization timestamp is first considered to determine their creation order. The initialization timestamp is a time mark assigned when the node is created, which can accurately record the moment when the node is created.

[0141] If the initialization timestamp of a predecessor node is earlier than that of another predecessor node, it can be determined that the earlier created predecessor node is the dependent predecessor node. This is because in many data processing scenarios, the first created node often performs some initialization operations or preparatory work first, and the subsequent nodes may need to rely on these preliminary work completed by it to successfully carry out their own tasks. For example, in a data processing flow, node A is created at 9 am (as can be seen from its initialization timestamp), and node B is created at 10 am, then node B may depend on node A, because node A is created before node B, and may have completed some preliminary processing of data or preparation of resources. Node B needs to continue to perform tasks based on the results of node A.

[0142] In addition to determining the creation order based on the initialization timestamp, if the creation order of two predecessor nodes is different, the predecessor node with a smaller node ID will be determined as the dependent predecessor node based on the node ID size of each predecessor node. The node ID is a unique identification information of each data task node, which can be in the form of a number, a string, etc., and usually has certain sorting rules.

[0143] When the creation order of two predecessor nodes is different (the creation order here can also be determined by other methods, such as the initialization timestamp mentioned above), the predecessor node that is dependent is further determined based on the node identifier size. If the node identifier of a predecessor node is smaller than the node identifier of another predecessor node under the established sorting rules, the predecessor node with the smaller node identifier of this data task is determined as the dependent predecessor node. For example, the node identifier of node C is "001" and the node identifier of node D is "002". According to the sorting rules of digital size, the node identifier of node C is smaller, so it can be considered that node D may be dependent on node C. Because in some cases, the size of the node identifier may be related to factors such as the creation order of the node or its priority in the system. Determining the dependency in this way helps to arrange the task execution order more reasonably.

[0144] Through the above implementation method, that is, determining the dependency relationship between two predecessor nodes based on the initialization timestamp or the node identifier size, it is possible to more accurately determine which predecessor node is dependent, thereby providing a basis for reasonably arranging the execution order of data task nodes in the same task queue in the future, ensuring that the data processing process can be carried out efficiently and orderly.

[0145] For a possible implementation, please refer to Figure 5 As shown, Figure 5 : is a flow chart of an optional data task scheduling method provided by an embodiment of the present application. After determining the execution order of multiple data task nodes in the distributed processing system according to the task queue, the method further includes:

[0146] S501, using a predefined dependency enqueue method to obtain a predecessor node list corresponding to each data task node in each task queue, wherein each predecessor node list is used to store at least one predecessor node of the data task node.

[0147] S502: Detect the status of the identification bit of each predecessor node in the predecessor node list.

[0148] S503, if the identification bit status of each predecessor node in the predecessor node list is a true value, it is determined that each predecessor node has been processed and the data task node is added to the priority queue.

[0149] S504, taking out and executing the multiple data task nodes from the priority queue in sequence according to the execution order of the multiple data task nodes.

[0150] After determining the execution order of multiple data task nodes in the distributed processing system according to the task queue, in order to further ensure that the task nodes can be executed in sequence in the correct order and when the dependency conditions are met, the list of predecessor nodes can be obtained by adopting a pre-defined dependency enqueue method, and the identification bit status of the predecessor node can be detected to determine whether the execution conditions are met. Only when all predecessor nodes are processed, the current data task node is added to the priority queue to wait for execution. Finally, each data task node is executed in sequence from the priority queue according to the established execution order, so as to ensure the accuracy and consistency of the entire data processing process.

[0151] First, a predefined dependency enqueue method is used. It should be understood that this method is specially designed to handle the dependencies between task nodes and to correctly queue nodes that meet the conditions. For each data task node in each task queue, a corresponding predecessor node list can be obtained. The function of the predecessor node list is to store those nodes that are ranked before the current data task node in the execution order, and whose execution results will affect the current data task node, that is, the predecessor nodes. For example, in a data processing task, data task node A may depend on node B and node C that have been executed before, then node B and node C will be recorded in the predecessor node list of data task node A.

[0152] Once the list of predecessor nodes is obtained, the next step is to check the flag status of each predecessor node in the list. The flag status is a way to mark whether the predecessor node has completed processing, which can usually be represented by a Boolean value (true or false). For example, when a predecessor node completes the part of the data processing task it is responsible for, its flag status may be set to true, indicating that the processing has been completed; if the processing has not been completed, the flag status is false. By checking this flag status, you can clearly know the current processing progress of each predecessor node.

[0153] When the detection finds that the identification bit status of each predecessor node in the predecessor node list is true, that is, it indicates that the processing has been completed, this means that all the predecessor nodes that the current data task node depends on have successfully completed their respective tasks. In this case, it can be determined that the current data task node meets the dependency conditions for execution, so the data task node is added to the priority queue. Joining the priority queue is to be able to execute these nodes that meet the conditions in a certain order (usually according to the previously determined execution order) later.

[0154] Finally, according to the execution order of multiple data task nodes that have been determined before, these data task nodes are taken out from the priority queue and executed in sequence. For example, if the execution order determined before is to execute data task node A first, and then execute data task node B, when each data task node meets its own dependency conditions and is added to the priority queue, node A and node B will be taken out from the priority queue in this order for execution, thereby ensuring that the entire data processing process is carried out in an orderly manner according to the predetermined plan and logic, avoiding errors or data processing problems caused by unsatisfied dependencies, and ensuring that data processing tasks can be completed efficiently, accurately, and orderly in a distributed processing system.

[0155] In a possible implementation manner, the method further includes:

[0156] In determining the execution order of multiple data task nodes, or in the process of executing multiple data task nodes, corresponding log records are generated.

[0157] In this application example, in the entire data task node processing process, whether it is the stage of determining the execution order of these nodes or the actual execution of multiple data task nodes, corresponding log records must be generated. The purpose of this is to be able to track and record the entire data processing process in detail, so as to facilitate subsequent monitoring, troubleshooting, performance analysis, and process optimization operations, thereby ensuring that data processing tasks can run more stably and efficiently under the distributed processing system.

[0158] When determining the execution order of multiple data task nodes, many complex operations and judgments are involved, such as analyzing the priority of nodes in different task queues, the dependencies between nodes in the same task queue, etc., in order to finally determine the order in which each data task node is executed. By generating corresponding log records at each step, each step taken in the process of determining the execution order, the rules based on, and each intermediate result obtained can be recorded in detail. For example, record the process of determining that the node type of a task queue is a source type (SOURCE), and then determine that the node in the queue has a higher priority; or record the process of analyzing the dependency between two data task nodes in the same task queue, and finally determine the execution order by comparing their creation order or other factors.

[0159] In addition, the above log records can provide a clear basis for subsequent review and understanding of why the data task nodes are arranged in such an execution order. It also helps to quickly locate which step in determining the execution order has an error when a problem occurs.

[0160] When each data task node is actually executed, the execution status of each data task node, the data content processed, the execution time node and other information can be recorded through the log. For example, record the specific time when a data task node starts to execute, whether there are any abnormalities during the execution process (such as data reading errors, processing logic errors, etc.), if there are abnormalities, record the type and specific manifestations of the abnormality in detail; you can also record the output results of the node after processing the data and other information.

[0161] In this way, during the execution of data processing tasks, by viewing log records, you can understand the working status of each data task node in real time, discover potential problems in time and take corresponding measures to deal with them. Moreover, after the task is completed, these log records can also provide important reference for evaluating the performance of the entire data processing task, such as processing speed, data accuracy, etc.

[0162] By generating corresponding log records at the key stages of determining the execution order and executing data task nodes, the data processing flow can be monitored and recorded in all aspects, providing a strong guarantee for the smooth operation and subsequent optimization of data processing tasks in a distributed processing system.

[0163] The data task scheduling method provided in the example of this application is implemented based on the distributed real-time processing architecture of the distributed processing system to solve the problems of inflexibility and cumbersome operation of data governance and data modeling in the prior art. Through automated task dependency analysis and fault-tolerant processing mechanism, efficient scheduling of tasks can be achieved to ensure the stable operation of the system in high-concurrency scenarios. At the same time, the example of this application also provides a business-oriented task configuration interface to simplify the operation process of task scheduling and improve the usability and scalability of the distributed real-time processing framework in practical applications. Through the above-mentioned data task scheduling method, it can be ensured that the order of logical dependencies is met during the actual task execution process. Complex task workflows can be automatically generated and handed over to the distributed computing framework for execution, which improves the user experience, reduces the cumbersomeness and inflexibility of data modeling and data governance, and reduces cost investment.

[0164] Specifically, the task orchestration method has the following significant advantages: 1) Flexible orchestration: Users can directly orchestrate the corresponding task flow through the business configuration interface according to the needs of different business scenarios. 2) Improve user experience: Users only need to drag and drop to perform complex data governance and data modeling, which greatly reduces the requirements for users' professional literacy, thereby improving the overall user experience. 3) Improve development efficiency: Users only need to focus on specific businesses, and do not need to pay too much attention to the complex data modeling and data governance process, which greatly improves development efficiency. 4) Reliability of data dependency: According to the above sorting rules, it can be ensured that at each stage of the task flow, the data has been processed and prepared at the previous node, which can avoid the situation of "data not ready" or "unable to obtain data" when the subsequent nodes rely on the previous data, thereby improving the reliability of the task flow. 5) Avoid concurrency conflicts: Through strict sorting, it is ensured that the execution order of the nodes conforms to the logical dependency of the data, avoiding data read and write conflicts or resource preemption during the concurrent execution of tasks. Especially in large-scale distributed systems, this sorting mechanism can greatly reduce concurrency conflicts. 6) Improved execution efficiency of task flows: Since the SOURCE and HANDLER nodes are executed first and generate temporary views, downstream nodes can quickly obtain processed data and avoid repeated processing. This orderly execution mechanism improves the overall processing efficiency of task flows, especially in complex multi-node dependency scenarios, which can significantly reduce execution time. 7) Reduced error risk: This sorting strategy avoids the risk of SINK or HANDLER nodes processing when data is not ready, thereby reducing the possibility of data loss or execution failure and enhancing system stability. 8) Adapt to complex business logic: In actual applications, many business processes may involve multiple data sources, multi-layer data processing, and multi-terminal data output. Through this reasonable sorting rule, the system can automatically perform optimal sorting according to business logic, reduce human intervention, and ensure the correctness of business processes and data consistency.

[0165] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0166] Corresponding to the data task arrangement method in the above embodiment, Figure 6 is a schematic diagram of the structure of a data task scheduling device provided in an embodiment of the present application. The device can be implemented as part or all of a computer device by software, hardware, or a combination of both. The computer device can be Figure 7 Electronic equipment shown.

[0167] Reference Figure 6, the data task arrangement device comprises:

[0168] A generating unit 601 is used to generate a plurality of data task nodes in response to a user's task node creation operation on a distributed processing system, wherein the data format of the data task nodes is a JavaScript object notation JSON format;

[0169] A parsing unit 602 is used to parse each data task node to obtain the node type of each data task node;

[0170] The allocation unit 603 is used to allocate each data task node to a corresponding task queue according to the node type of each data task node, wherein each task queue corresponds to a different node type;

[0171] The sorting unit 604 is used to determine the execution order of multiple data task nodes in the distributed processing system according to the task queue.

[0172] It should be noted that the data task orchestration device provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0173] The functional units and modules in the above embodiments may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit, and the above integrated units may be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present application.

[0174] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0175] An embodiment of the present application further provides an electronic device, the electronic device comprising one or more processors and a memory;

[0176] The memory is coupled to one or more processors, and the memory is used to store computer program codes, the computer program codes include computer instructions, and one or more processors call the computer instructions to enable the electronic device to execute the data task arrangement method shown above.

[0177] Figure 7The schematic diagram of the structure of an electronic device provided in the embodiment of the present application is that the electronic device 700 can be a mobile phone, a smart screen, a tablet computer, a wearable electronic device, an in-vehicle electronic device, an augmented reality (AR) device, a virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a projector, or a communication device such as a server, a storage device, a base station, or a smart car, etc. The embodiment of the present application does not impose any restrictions on the specific type of the electronic device.

[0178] The memory 701 can be used to store computer software programs 702 and modules, and the processor 703 executes various functional applications and data processing of the electronic device by running the software programs and modules stored in the memory 701. The memory 701 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the electronic device (such as audio data, a phone book, etc.), etc. In addition, the memory 701 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0179] Among them, the processor 703 may include one or more processors such as a central processing unit, an application processor (AP), a baseband processor, etc. The processor may be the nerve center and command center of the wireless router. The processor 703 may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions. The memory 701 may be used to store computer executable program codes, and the executable program codes include instructions. The processor 703 executes various functional applications and data processing of the network device by running the instructions stored in the memory. The memory 701 may include a program storage area and a data storage area, such as storing data of a sound signal to be played. For example, the memory may be a double rate synchronous dynamic random access memory DDR or a flash memory Flash.

[0180] An embodiment of the present application also provides a computer-readable storage medium, in which computer instructions are stored; when the computer-readable storage medium runs on an electronic device, the electronic device executes the data task scheduling method shown above.

[0181] Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. Computer-readable storage media can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. that includes one or more available media integrated therein. Available media can be magnetic media (e.g., floppy disks, hard disks, tapes), optical media, or semiconductor media (e.g., solid state disks (SSDs)), etc.

[0182] The embodiment of the present application also provides a computer program product including computer instructions. When the computer program product is executed on an electronic device, the electronic device can execute the data task scheduling method shown above.

[0183] The computer storage medium and computer program product provided in the above-mentioned embodiments of the present application are used to execute the method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects corresponding to the method provided above, and will not be repeated here.

[0184] In the above embodiments, it can also be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When loading and executing computer instructions on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from a website site, computer, server or data center by wired (such as: coaxial cable, optical fiber, data subscriber line (Digital Subscriber Line, DSL)) or wireless (such as: infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server, data center, etc. that includes one or more available media integration. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a digital versatile disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)).

[0185] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0186] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments applied for herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0187] In the embodiments provided in the present application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0188] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0189] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A data task scheduling method, characterized in that: include: In response to the user's task node creation operation on the distributed processing system, a plurality of data task nodes are generated, wherein the data format of the data task nodes is in the JavaScript object notation JSON format; Parsing each of the data task nodes to obtain the node type of each of the data task nodes; According to the node type of each data task node, each data task node is assigned to a corresponding task queue, wherein each task queue corresponds to a different node type; The execution order of the plurality of data task nodes in the distributed processing system is determined according to the task queue.

2. The method according to claim 1, characterized in that The step of generating a plurality of data task nodes in response to a user creating a task node of the distributed processing system comprises: In response to the task node creation operation performed by the user through the business configuration interface of the distributed processing system, a plurality of the data task nodes are generated, wherein different business configuration interfaces correspond to different business scenario requirements.

3. The method according to claim 1, characterized in that The parsing of each of the data task nodes to obtain the node type of each of the data task nodes includes: A task scheduling rule engine is used to parse each of the data task nodes to obtain the node type of each of the data task nodes, wherein the task scheduling rule engine is used to define the node type and node priority corresponding to each data task node, and the node type includes: source type, handler type, and aggregation type.

4. The method according to claim 1, characterized in that: According to the node type of each of the data task nodes, each of the data task nodes is assigned to a corresponding task queue, including: Determine, according to the task queue corresponding to the node type of each data task node, a mapping relationship between each data task node and a predecessor node list in the corresponding task queue; Based on the mapping relationship, each data task node is stored in the corresponding task queue with the node identifier of each data task node as the key and the predecessor node list corresponding to each data task node as the value.

5. The method according to claim 1, characterized in that Determining the execution order of the plurality of data task nodes in the distributed processing system according to the task queue includes: If at least two of the plurality of data task nodes are in different task queues, respectively, determining the node priority of each data task node in each task queue; If at least two of the plurality of data task nodes are in the same task queue, determining a dependency relationship between the at least two data task nodes in the same task queue; The execution order of the multiple data task nodes in the distributed processing system is determined according to the node priorities and / or dependencies of the multiple data task nodes.

6. The method according to claim 5, characterized in that If at least two of the plurality of data task nodes are in different task queues, determining the node priority of each data task node in each task queue includes: If at least two of the plurality of data task nodes are in different task queues, respectively, determining the node types corresponding to the different task queues, wherein the node type includes at least one of the following: a source type, a processing program type, and a convergence type; If the node type corresponding to the first task queue in the different task queues is a source type, the node priority of the data task node in the first task queue is the first priority; If the node type corresponding to the second task queue in the different task queues is a processing program type, the node priority of the data task node in the second task queue is a second priority, wherein the second priority is less than the first priority; If the node type corresponding to the third task queue in the different task queues is the aggregation type, the node priority of the data task node in the third task queue is the third priority, wherein the third priority is lower than the second priority.

7. The method according to claim 5, characterized in that If at least two of the plurality of data task nodes are in the same task queue, determining the dependency relationship between the at least two data task nodes in the same task queue includes: If at least two of the plurality of data task nodes are in the same task queue, determining whether there is a predecessor node in the at least two data task nodes in the same task queue; If there is one of the predecessor nodes among at least two of the data task nodes, determining that a non-predecessor node among at least two of the data task nodes depends on the predecessor node; If there are two predecessor nodes in at least two of the data task nodes, determining a dependency relationship between the two predecessor nodes based on the creation order of the two predecessor nodes; If the predecessor node does not exist in at least two of the data task nodes, the dependency relationship between the two predecessor nodes is determined based on the creation order of the two predecessor nodes.

8. The method according to claim 7, characterized in that Determining a dependency relationship between the two predecessor nodes based on a creation order of the two predecessor nodes includes: If the creation order of the two predecessor nodes is different, determining a predecessor node that is created first among the two predecessor nodes as the dependent predecessor node, wherein the creation order is determined based on an initialization timestamp or a node identifier size; If the creation orders of the two predecessor nodes are different, then according to the size of the node identifier of each predecessor node, a predecessor node with a smaller node identifier will be determined as the dependent predecessor node.

9. The method according to any one of claims 1 to 8, characterized in that After determining the execution order of the plurality of data task nodes in the distributed processing system according to the task queue, the method further includes: Using a predefined dependency enqueue method, obtain a predecessor node list corresponding to each data task node in the task queue, wherein each predecessor node list is used to store at least one predecessor node of the data task node; Detecting the status of the identification bit of each predecessor node in the predecessor node list; If the identification bit status of each of the predecessor nodes in the predecessor node list is a true value, it is determined that each of the predecessor nodes has been processed, and the data task node is added to the priority queue; According to the execution order of the multiple data task nodes, the multiple data task nodes are taken out from the priority queue in sequence and executed.

10. A data task scheduling device, characterized in that: include: A generating unit, for generating a plurality of data task nodes in response to a user's task node creation operation on a distributed processing system, wherein the data format of the data task nodes is a JavaScript object notation JSON format; A parsing unit, used for parsing each of the data task nodes to obtain the node type of each of the data task nodes; An allocating unit, configured to allocate each of the data task nodes to a corresponding task queue according to the node type of each of the data task nodes, wherein each of the task queues corresponds to a different node type; A sorting unit is used to determine the execution order of multiple data task nodes in the distributed processing system according to the task queue.

Citation Information

Cited By

  • Type definition processing method of EDA simulator, electronic equipment and medium

    CN121902727A

  • Type definition processing method of EDA simulator, electronic device and medium

    CN121902727B