Task processing method and corresponding device
By dividing nodes in directed acyclic graph into task groups, the problem of cloud systems affecting performance due to memory overflow when processing big data is solved, and more efficient memory usage and system performance improvement is achieved.
Patent Information
- Application Number
- PCT/CN2024/139019
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-15
- Filing Date
- 2024-12-13
- Publication Date
- 2025-06-19
AI Technical Summary
When cloud systems process big data, they are prone to memory overflow due to a large amount of intermediate data occupies memory, affecting system performance.
By dividing nodes in the directed acyclic graph into at least two task groups, the number of tasks in each task group is not greater than N, and the incoming degree of tasks in the task group is zero, so as to reduce the use of machine memory and improve task processing efficiency.
It reduces memory usage during task processing, improves system performance, reduces end-to-end delay, and improves resource utilization.
Smart Images

Figure CN2024139019_19062025_PF_FP_ABST
Abstract
Description
A task processing method and corresponding device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on December 15, 2023, with application number 202311735767.4 and application name “A method for task processing and corresponding device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of cloud computing technology, and in particular to a task processing method and corresponding device. Background Art
[0003] When cloud systems process big data, they typically convert it into processing different tasks. Processing a large amount of data involves multiple processing tasks. Currently, most data processing tasks can be abstracted into a directed acyclic graph (DAG), with tasks as nodes and task dependencies as edges. Task orchestration and scheduling can then be performed based on this DAG.
[0004] Current task scheduling and dispatching solutions typically employ best-effort algorithms or QoS-constrained algorithms. Best-effort algorithms disregard resource costs and prioritize the earliest possible task completion or minimize the overall workflow completion time. QoS-constrained algorithms prioritize not only the earliest possible task completion but also the resource costs of using different resources to meet the service quality requirements of different users.
[0005] Best-effort algorithms incur significant material costs and waste resources. Currently, there has been some progress in algorithms based on quality of service control, which can shorten completion times by executing tasks in parallel. However, the limited memory space available on cloud systems limits the parallelism of in-memory computing. In in-memory computing, tasks generate intermediate data, which cannot be consumed and released in a timely manner. This leads to memory accumulation and can easily cause memory overflow, impacting system performance. Summary of the Invention
[0006] This application provides a task processing method for reducing memory usage and improving system performance during big data task processing. This application also provides corresponding devices, systems, computer-readable storage media, and computer program products.
[0007] In a first aspect, the present application provides a method for task processing, comprising: obtaining a first directed acyclic graph for describing a data processing task; the first directed acyclic graph comprises a plurality of nodes and a plurality of edges for connecting the nodes, wherein each node is used to represent a task and the edge is used to represent a dependency relationship between two connected nodes; dividing the plurality of nodes into at least two task groups; wherein the number of tasks in each task group is not greater than N, and the in-degree of the nodes associated with the tasks in each task group is zero, N being determined by the size of the machine memory, the size of the memory occupied by the instance, and the size of the memory occupied by the output data of the task, and N being a positive integer; and processing each task group.
[0008] In the present application, each node in the first directed acyclic graph may represent a task or a task set, and there is no dependency between the tasks in the task set.
[0009] In the present application, the process of dividing the first directed acyclic graph into at least two task groups may be performed in multiple rounds, with each round only dividing the tasks associated with nodes with zero in-degree.
[0010] In this application, machine memory can be the memory of the physical machine that executes the task group in the cloud system, or it can refer to the memory of the virtual machine (VM) that executes the task group; instance refers to software resources such as processes or threads that execute the task group, or hardware resources such as virtual machines or containers; the output data of a task refers to the data generated after the task is executed.
[0011] In this application, tasks in the same task group can be processed in parallel, and task groups that have no dependencies between different task groups can also be processed in parallel.
[0012] In this first aspect, the in-degree of the node associated with the task in each task group is 0, indicating that the tasks in the same task group are independent of each other and do not need to generate intermediate data, which can reduce the occupancy of machine memory; in addition, the number of tasks in the task group is no more than N, which can limit the amount of output data of the task, thereby reducing the memory occupancy of the task output data; moreover, the tasks in the same task group are fully decoupled and independent of each other, and can be processed in parallel, which can improve the efficiency of task processing and reduce the end-to-end (edge to edge, E2E) delay of task processing, thereby improving the system performance of task processing.
[0013] In one possible implementation, the above step of dividing multiple nodes into at least two task groups includes: dividing the nodes with zero in-degree in the first directed acyclic graph into at least one task group to obtain a second directed acyclic graph, and updating the in-degree of the nodes in the second directed acyclic graph; dividing the nodes with zero in-degree in the second directed acyclic graph into at least one task group.
[0014] In this possible implementation, when dividing the task groups, only one task group can be divided in each round. Even if the number of nodes with in-degree 0 is greater than N, only one task group is divided. This can avoid rapid memory occupation. Task groups can also be divided in batches when the number of nodes with in-degree 0 is greater than N. When there is sufficient memory, task groups can be divided in batches to increase the speed of directed acyclic graph division.
[0015] In one possible implementation, the above step of: dividing the nodes with zero in-degree in the first directed acyclic graph into at least one task group includes: when the number of nodes with zero in-degree in the first directed acyclic graph is not greater than N, dividing the nodes with zero in-degree into one task group; when the number of nodes with zero in-degree in the first directed acyclic graph is greater than N, dividing no more than N nodes that meet the affinity requirements into the same task group.
[0016] In this possible implementation, tasks that meet affinity requirements can be understood as tasks that have dependencies on the same subsequent task. These tasks facilitate the processing of subsequent tasks, and there are no dependencies between these tasks. Grouping tasks by affinity improves the processing efficiency of related subsequent tasks.
[0017] In one possible implementation, the above steps: processing each task group includes: establishing a correspondence between the P tasks in the first task group and the M instances according to the task scheduling goal, P≤N, and P is a positive integer, M is a positive integer, M instances are used to execute the P tasks, the first task group is any one of the at least two task groups, and the task scheduling goal is the shortest sum of the time for the M instances to execute the P tasks.
[0018] In this possible implementation, the sum of the time it takes for M instances to execute P tasks can be derived from statistics from previous runs. This can be achieved by pre-calculating the time it takes for each instance to execute different tasks, and then calculating the sum of the time it would take for the P tasks to be executed by M instances. For each task group, when assigning instances to tasks, the P tasks are scheduled based on the time it takes for M instances to execute them, minimizing the time it takes for the P tasks to be executed. This minimizes the risk of a single task taking too long to process and slowing down the processing of the entire task group, thereby reducing end-to-end latency and improving resource utilization.
[0019] In one possible implementation, when P>M, the above steps: establishing a correspondence between the P tasks in the first task group and the M instances according to the task scheduling strategy, including: dividing the P tasks in the first task group into M subgroups according to the task scheduling goals; wherein the absolute value of the difference between the execution time of the tasks in at least one subgroup and the time reference value is less than a first threshold, and the time reference value includes the time for executing the P tasks and the ratio to M; establishing a correspondence between the M subgroups and the M instances.
[0020] In this possible implementation, when P ≤ M, each task can be associated with an instance. When P > M, the P tasks can be grouped so that the execution time of the tasks in each subgroup is as close as possible to the time reference value. That is, the absolute value of the difference between the execution time of the tasks in at least one subgroup and the time reference value is as close to 0 as possible. The first threshold can be a very small positive number close to 0, such as 0.1, 0.01, and other possible values. In this application, the P tasks are arranged and divided into different subgroups according to the requirement that the absolute value of the difference between the execution time of the tasks in at least one subgroup and the time reference value is less than the first threshold, which is conducive to improving the efficiency of task processing.
[0021] In a possible implementation, the M subgroups include at least one first subgroup, the first subgroup includes a task, and the execution time of the task is greater than the time reference value.
[0022] In this possible implementation, when the execution time of a task is greater than the time reference value, the task can be divided into a subgroup and executed exclusively by an instance, thereby preventing the task from slowing down other tasks.
[0023] In a possible implementation, the M subgroups include at least one second subgroup, the second subgroup includes at least two tasks, and the absolute value of the difference between the execution time of the at least two tasks and the time reference value is less than a first threshold.
[0024] In this possible implementation, when the execution time of some tasks is short, at least two such tasks can be divided into a subgroup, and at least two tasks in the subgroup can be executed by the same instance. In this way, the execution time of each instance can be roughly the same, there is no need to enter the waiting state too early, and resource utilization is improved.
[0025] In one possible implementation, when P>M, the P tasks in the first task group are established in correspondence with the M instances according to the task scheduling strategy, including: establishing an allocation relationship between each of the P tasks and the M instances; wherein the absolute value of the difference between the execution time of at least two tasks pointing to the same instance and the time reference value is less than a first threshold, and the time reference value includes the time for executing the P tasks and the ratio to M; configuring priorities for the at least two tasks pointing to the same instance, and the priorities are used to indicate the order in which the at least two tasks pointing to the same instance are allocated to the corresponding instances.
[0026] In this possible implementation, there's no need to divide the P tasks into subgroups. Instead, the instances for executing the tasks can be directly determined. When two or more tasks are configured with the same instance, it's necessary to determine the execution priority for these tasks. This way, when subsequently assigning tasks to instances, the tasks can be assigned to the corresponding instances based on the instance-task assignment relationship and priority. When determining the instance for each task, the absolute value of the difference between the execution time of at least two tasks pointing to the same instance and the time reference value is less than a first threshold. This ensures that each instance executes tasks at roughly the same time, eliminating the need to enter a waiting state prematurely and improving resource utilization.
[0027] In one possible implementation, the method further includes: obtaining a maximum number of instances, a memory size occupied by each instance, a memory size occupied by output data of a single task, and a machine memory size; and determining N using the following relationship;
[0028] Among them, X is the maximum number of instances, MemAllocate insi ResSize is the memory size occupied by each instance taskj The size of the memory occupied by the output data of a single task, Limit mem The size of the machine's memory.
[0029] In this possible implementation, the value of N can be dynamically determined using the above relationship, which is beneficial for improving the accuracy of task group division.
[0030] In one possible implementation, the above steps of obtaining a first directed acyclic graph for describing data processing tasks include: receiving requirements for processing data, the requirements being used to determine the tasks to be performed; parsing the dependencies between the tasks to be performed, and constructing a first directed acyclic graph.
[0031] In this possible implementation, a data processing requirement may be proposed by a user or tenant. For example, the requirement may be one or more indicators in statistical big data. Based on the requirement, multiple tasks to be performed are determined, and then a first directed acyclic graph is constructed based on the dependencies between the tasks. This directed acyclic graph can improve the efficiency of subsequent task processing.
[0032] A second aspect of the present application provides a task processing apparatus for executing the method of the first aspect or any possible implementation of the first aspect. Specifically, the task processing apparatus includes modules or units for executing the method of the first aspect or any possible implementation of the first aspect, such as an acquisition unit, a first processing unit, and a second processing unit.
[0033] The third aspect of the present application provides a task processing device, including a transceiver, a processor and a memory, wherein the transceiver and the processor are coupled to the memory, and the memory is used to store programs or instructions. When the program or instructions are executed by the processor, the task processing device executes the method in the aforementioned first aspect or any possible implementation of the first aspect.
[0034] The fourth aspect of the present application provides a chip system, which includes one or more interface circuits and one or more processors; the interface circuits and the processors are interconnected by lines; the interface circuits are used to receive signals from the memory of the task processing device and send signals to the processor, and the signals include computer instructions stored in the memory; when the processor executes the computer instructions, the task processing device executes the method in the aforementioned first aspect or any possible implementation of the first aspect.
[0035] In a fifth aspect, the present application provides a computer-readable storage medium having a computer program or instruction stored thereon. When the computer program or instruction is executed on a computer device, the computer device executes the method in the aforementioned first aspect or any possible implementation of the first aspect.
[0036] In a sixth aspect, the present application provides a computer device program product, which includes a computer device program code. When the computer device program code is executed on a computer device, the computer device executes the method in the aforementioned first aspect or any possible implementation of the first aspect.
[0037] In a seventh aspect, the present application provides a computer device cluster, comprising at least one computer device, each computer device comprising a processor and a memory; the processor of at least one computer device is used to execute instructions stored in the memory of at least one computer device, so that the computer device cluster executes the method in the aforementioned first aspect or any possible implementation of the first aspect.
[0038] In an eighth aspect, the present application provides a cloud system, comprising: a device for processing tasks according to the second or third aspect above.
[0039] Among them, the technical effects brought about by the second to eighth aspects can refer to the technical effects brought about by the first aspect or different possible implementation methods of the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] FIG1A is a schematic diagram of a structure of a cloud system provided in an embodiment of the present application;
[0041] FIG1B is another schematic diagram of the structure of the cloud system provided in an embodiment of the present application;
[0042] FIG2 is a schematic diagram of a structure of a task processing device provided in an embodiment of the present application;
[0043] FIG3 is a schematic diagram of an embodiment of a task processing method provided in an embodiment of the present application;
[0044] FIG4 is a schematic diagram of an example of a directed acyclic graph provided in an embodiment of the present application;
[0045] FIG5 is a schematic diagram illustrating an example process of dividing a directed acyclic graph according to an embodiment of the present application;
[0046] FIG6 is a schematic diagram of a correspondence relationship between tasks and instances provided in an embodiment of the present application;
[0047] FIG7 is another schematic diagram of the structure of the cloud system provided in an embodiment of the present application;
[0048] FIG8 is a schematic diagram of an architecture of a control plane in a cloud system provided by an embodiment of the present application;
[0049] FIG9 is a schematic diagram of an example of attributes of a task provided in an embodiment of the present application;
[0050] FIG10 is a schematic diagram illustrating an example of dependency relationships of tasks provided in an embodiment of the present application;
[0051] FIG11 is a schematic diagram of an example scenario provided by an embodiment of the present application;
[0052] FIG12 is another structural diagram of the task processing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0053] The following describes the embodiments of the present application in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present application, not all of the embodiments. Those skilled in the art will appreciate that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0054] The terms "first," "second," and the like in the specification and claims of this application and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatus.
[0055] The present application provides a task processing method for reducing memory usage and improving system performance during big data task processing. The present application also provides corresponding devices, systems, computer-readable storage media, and computer program products. These are described in detail below.
[0056] To facilitate understanding, the technology and background involved in the embodiments of this application are explained below.
[0057] Virtualization is a resource management technology that abstracts and transforms a host's physical resources, such as computing, networking, and storage, to create a more tangible representation. This breaks down the barriers between the host's physical structure and allows users to utilize these resources in a more efficient manner than their original configuration. Virtualized resources are not restricted by the configuration, location, or physical configuration of existing physical resources.
[0058] Physical machine (PM): A physical resource used to host virtualization technology. A host is also called a physical machine. Typically, a physical server is used to deploy virtual instances. A physical machine has multiple physical devices. For example, a physical server has physical devices such as a processor and memory. Multiple virtual instances can be deployed on a single host. Multiple virtual instances deployed on the same host share the host's physical resources. Depending on the usage scenario, a single host can host virtual instances belonging to only one user or multiple users.
[0059] Instances typically run on a host's operating system and can use the host's hardware resources. Instances are isolated from each other. Typically, instances can be software resources such as processes or threads that execute task groups, or hardware resources such as virtual machines or containers.
[0060] A virtual machine (VM) is a complete computer system that uses virtualization technology to simulate the full functionality of a hardware system and runs in a completely isolated environment. A subset of the VM's instructions can be processed on the host computer, while other instructions can be executed in an emulated manner. A VM is also called a virtual server.
[0061] A virtual machine can be considered a collection of several virtual devices, which is a complete computer system with complete hardware system functions and running in a completely isolated environment. Virtual devices are virtualized based on physical devices that can share resources using virtualization technology. For example, a virtual processor virtualized based on a processor using virtualization technology is a virtual device. For another example, a training card virtualized based on a field-programmable gate array (FPGA) using virtualization technology is also a virtual device.
[0062] Containers provide a lightweight virtual runtime environment. Containers can be obtained by packaging all the code, libraries, and dependencies of the user's application into an image. When the image is executed, the image runs in the virtual runtime environment. At this time, the container is a runtime instance of the image, similar to a lightweight sandbox, which can be started, started, stopped, and deleted. The image does not share the host's memory, processor (such as the central processing unit (CPU)), and disk resources with other images, achieving container isolation between the image and the host, and between the image and other images, ensuring that the process in the container cannot monitor any process or resources outside the container. Container technologies include Docker, Kubernetes, CoreOS, and other container technologies.
[0063] Task scheduling refers to the rational arrangement of the execution order of multiple tasks based on the dependencies and priorities between tasks to improve resource utilization efficiency and overall processing efficiency.
[0064] A slowdown refers to the task that takes the longest to execute in a parallel processing process and slows down the entire processing flow. Such a task or process that takes a long time to execute and affects the overall progress is called a slowdown.
[0065] The purpose of dependency decoupling is to reduce unnecessary dependencies between different tasks, thereby improving the flexibility of task scheduling and preventing the delay of certain tasks from affecting subsequent tasks that depend on them.
[0066] The virtualization technology introduced above is usually applied to a cloud system, which can be a public cloud, a private cloud, or a hybrid cloud. For more information about the cloud system, please refer to Figure 1A.
[0067] As shown in Figure 1A, a structure of the cloud system provided in an embodiment of the present application may include: a terminal device 01, a scheduling node 02, a working node cluster 03, and a working node cluster 04. Communication connections can be established between the client 01 and the scheduling node 02, between the scheduling node 02 and the working node cluster 03, between the scheduling node 02 and the working node cluster 04, and between the working node cluster 03 and the working node cluster 04. For example, a communication connection can be established between the terminal device 01 and the scheduling node 02 via a network. Optionally, the network can be a local area network, the Internet, or other networks, which are not limited by the embodiments of the present application.
[0068] In this implementation environment, cloud system operators can interact with scheduling node 02 via terminal device 01. For example, operators can send cloud system deployment instructions to scheduling node 02 via terminal device 01, instructing scheduling node 02 to deploy the cloud system based on the resources of scheduling node 02, worker node cluster 03, and worker node cluster 04. The cloud system is used to manage the resources of worker node cluster 03 and worker node cluster 04 and provide cloud services to users based on the resources of worker node cluster 03 and worker node cluster 04. Users can also send data processing requirements via terminal device 01. Scheduling node 02 can generate corresponding data processing tasks based on these data processing requirements and schedule the tasks to the worker nodes in worker node cluster 03 and / or worker node cluster 04 for processing.
[0069] In the task processing solution provided by the embodiment of the present application, the scheduling node 02 can determine the tasks to be executed according to user needs; analyze the dependencies between the tasks to be executed, construct a directed acyclic graph (DAG), and then divide the DAG into multiple rounds, divide the nodes on the DAG into at least two task groups, and orchestrate the tasks in the task groups. The orchestrated tasks are then assigned to instances for execution. The instance can be located on the scheduling node or in the worker node cluster 03 and the worker node cluster 04.
[0070] Optionally, the terminal device 01, also referred to as user equipment (UE), mobile station (MS), mobile terminal (MT), etc., is a device that includes wireless communication functionality (providing voice / data connectivity to users), for example, a handheld device with wireless connection functionality. Currently, some examples of terminal devices include: mobile phones, tablet computers, laptops, PDAs, notebook computers, wireless routers, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving cars, wireless terminals in the Internet of Vehicles, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, or wireless terminals in smart homes.
[0071] The scheduling node 02 can be a cloud server or a cloud physical machine, or a cloud server cluster or a physical machine cluster composed of several cloud servers, or a cloud computing service center. The functions of the scheduling node can be implemented by software or hardware.
[0072] As an example of a software functional unit, a scheduling node may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, a scheduling node may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.
[0073] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Inter-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.
[0074] As an example of a hardware functional unit, a scheduling node may include at least one computing device, such as a server. Alternatively, the scheduling node may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0075] The multiple computing devices included in a scheduling node can be distributed in the same zone or in different zones. The multiple computing devices included in a scheduling node can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in a scheduling node can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.
[0076] Worker node cluster 03 and worker node cluster 04 can be a server cluster consisting of several servers, or a cloud computing service center. A large number of basic resources owned by the cloud service provider are deployed in the cloud computing service center. For example, computing resources, storage resources, and network resources are deployed in the cloud computing service center.
[0077] It should be noted that the scheduling node 02, working node cluster 03 and working node cluster 04 in this implementation environment can also be implemented through other resource platforms besides the cloud computing service center, and this embodiment of the application does not specifically limit it.
[0078] It should be understood that the above content is an illustrative description of the cloud system provided in the embodiment of the present application, and does not constitute a limitation on the application scenarios of the deployment method of the cloud resource management system. Ordinary technicians in this field know that as business needs change, its application scenarios can be adjusted according to application needs.
[0079] In the embodiments of the present application, the architecture of the cloud system can also be understood with reference to Figure 1B . As shown in Figure 1B , the cloud system includes a cloud platform and basic resources. The cloud platform includes a cloud platform manager, and the scheduling node described above can be the cloud platform manager in Figure 1B . The basic resources can include multiple servers, each of which can be a worker node, or each server can include multiple worker nodes.
[0080] The working node in FIG1B may be a computing device card or a virtual machine (VM), wherein the computing device card may be at least one of a central processing unit (CPU), a graphics processing unit (GPU), and a neural network processing unit (NPU).
[0081] The cloud platform manager maintains or regularly collects information about each worker node in the underlying resources, such as resource usage (resource utilization or idleness, memory usage), etc. This information can be used as an auxiliary decision-making tool for the number of tasks in a task group.
[0082] The cloud platform manager receives data processing requests from users and then determines the corresponding tasks and the dependencies between them based on the data processing requirements. It then generates a DAG, divides the task groups, and orchestrates the tasks within the groups. The cloud platform manager can also assign the task groups to one or more worker nodes in the cloud system, which will then execute the corresponding tasks. After the worker nodes complete the tasks, the cloud platform manager returns the data processing results to the user.
[0083] FIG2 is a schematic diagram of a possible logical structure of a task processing device provided in an embodiment of the present application. The task processing device may be the scheduling node 02 in FIG1A or the cloud platform manager in FIG1B , or the task processing device may be included in the scheduling node 02 in FIG1A or the cloud platform manager in FIG1B . As shown in FIG2 , the task processing device 20 provided in an embodiment of the present application includes: a processor 201, a communication interface 202, a memory 203, and a bus 204. The processor 201, the communication interface 202, and the memory 203 are interconnected via the bus 204. In an embodiment of the present application, the processor 201 is used to control and manage the operations of the task processing device 20. For example, the processor 201 is used to generate a DAG based on the user's data processing requirements, and to divide the task groups and arrange the tasks within the groups. The communication interface 202 is used to support communication between the task processing device 20. For example, the communication interface 202 can schedule tasks to work nodes or receive data processing requirements from users. The memory 203 is used to store the program code and data of the task processing device 20.
[0084] The processor 201 may be a central processing unit (CPU), a general-purpose processor (GPOR), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic device (PLD), a transistor logic device (TLD), a hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. A processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like. The bus 204 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, for example. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG. 2 shows only one thick line, but this does not imply that there is only one bus or only one type of bus.
[0085] The following describes a method for task processing provided by an embodiment of the present application. The method can be executed by a task processing device, which can be an independent physical machine, a virtual machine, a processor, a chip, or a chip system.
[0086] FIG3 is a schematic diagram of an embodiment of a task processing method provided in an embodiment of the present application.
[0087] 301. Obtain a first directed acyclic graph for describing a data processing task.
[0088] The first directed acyclic graph includes multiple nodes and multiple edges for connecting the nodes, wherein each node is used to represent a task and the edge is used to represent the dependency relationship between the two connected nodes; each node in the first directed acyclic graph can represent a task or a task set, and there is no dependency relationship between the tasks in the task set.
[0089] The first directed acyclic graph can be understood by referring to FIG4 . As shown in FIG4 , the first directed acyclic graph includes 16 nodes, namely, node 1, node 2, ..., node 16. The arrows between the nodes represent the dependency relationship between the nodes, that is, the dependency relationship between the tasks represented by the nodes. Each arrow can represent an in-degree. Each node can correspond to a task or a set of tasks. Node 1, node 2, and node 3 are starting nodes. The node in-degrees of these three nodes are all 0, the node in-degree of node 10 is 2, and the node in-degrees of the other nodes are all 1.
[0090] FIG4 may correspond to a node list, which may be understood by referring to the following Table 1, as shown in Table 1:
[0091] Table 1: Node list of the first directed acyclic graph
[0092] When the first directed acyclic graph changes, the node list may be updated according to the change.
[0093] 302. Divide multiple nodes into at least two task groups, wherein the number of tasks in each task group is no more than N, and the in-degree of the nodes associated with the tasks in each task group is zero, where N is determined by the size of the machine memory, the size of the memory occupied by the instance, and the size of the memory occupied by the output data of the task, and N is a positive integer.
[0094] N in this application may be configured based on experience or determined dynamically.
[0095] In this application, the process of determining N may be to obtain the maximum number of instances, the memory size occupied by each instance, the memory size occupied by the output data of a single task, and the size of the machine memory; and determine N using the following relationship:
[0096] Among them, X is the maximum number of instances, MemAllocate insi ResSize is the memory size occupied by each instance taskj The size of the memory occupied by the output data of a single task, Limit mem The size of the machine's memory.
[0097] The maximum number of instances, the memory size of each instance, the memory size of the output data of a single task, and the machine memory size can be obtained by collecting system resources. The value of N can be the maximum value that satisfies the above relationship. When the value of N increases by 1, the above relationship is no longer satisfied.
[0098] In the present application, the process of dividing the first directed acyclic graph into at least two task groups can be carried out in multiple rounds, and each round only divides the tasks associated with the nodes with an in-degree of zero. The specific process can be: dividing the nodes with an in-degree of zero in the first directed acyclic graph into at least one task group to obtain a second directed acyclic graph, and updating the in-degree of the nodes in the second directed acyclic graph; dividing the nodes with an in-degree of zero in the second directed acyclic graph into at least one task group to obtain a new directed acyclic graph, and then updating the node list corresponding to the second directed acyclic graph, and then continuing the process of dividing the task groups until all nodes on the first directed acyclic graph are divided into task groups.
[0099] In dividing the nodes with zero in-degree in the first directed acyclic graph into at least one task group, when the number of nodes with zero in-degree in the first directed acyclic graph is not greater than N, the nodes with zero in-degree can be divided into one task group; when the number of nodes with zero in-degree in the first directed acyclic graph is greater than N, no more than N nodes that meet the affinity requirements are divided into the same task group. Tasks that meet the affinity requirements can be understood as tasks that have a dependency relationship with the same subsequent task. The execution of these tasks is conducive to better processing of subsequent tasks, and there is no dependency relationship between these tasks. Dividing task groups according to affinity is conducive to improving the processing efficiency of related subsequent tasks. For example: in Figure 4, the task associated with node 10 is a subsequent task of the tasks associated with nodes 5 and 6. The execution of the tasks associated with nodes 5 and 6 will affect the task associated with node 10. Nodes 5 and 6 are tasks that meet the affinity requirements.
[0100] It should be noted that a new task group will be enabled in each round of division, and the nodes divided in the second round will not be assigned to the task group that has been divided in the first round. This ensures that the tasks between and within the task groups are fully decoupled.
[0101] The number of tasks in each task group may be equal or unequal, as long as the number of tasks in each task group is not greater than N.
[0102] 4 , taking N=2 and 2 tasks being divided into each task group as an example, refer to FIG5 , which illustrates the process of task division in the DAG of FIG4 .
[0103] As shown in Figure 5, in the first round of partitioning, nodes 1 and 2 can be assigned to the same task group. For example, the assignment to task group 1 can be represented as: task group 1 {node 1, node 2}. After assigning nodes 1 and 2 to task group 1, Table 1 needs to be updated. This update process can involve deleting the rows for nodes 1 and 2 and updating the in-degrees of the nodes associated with them. Specifically, the in-degrees of nodes 4 and 5 are updated from 1 to 0.
[0104] In the second round of partitioning, nodes 3 and 4 can be assigned to task group 2, which can be represented as: Task group 2 {node 3, node 4}. After assigning nodes 3 and 4 to task group 2, the node list needs to be updated. This update process can be done by deleting the rows of nodes 3 and 4 and updating the in-degrees of the nodes associated with them. That is, the in-degrees of nodes 6, 7, 8, and 9 are updated from 1 to 0. The node list at this point can be understood by referring to Table 2:
[0105] Table 2: Node list of the directed acyclic graph after the second round of partitioning
[0106] In the third round of division, considering affinity, nodes 5 and 6 can be divided into task group 3, which can be expressed as: task group 3 {node 5, node 6}. After dividing nodes 5 and 6 into task group 3, it is necessary to continue to update the above node list 2. The update process can be to delete the rows of nodes 5 and 4, and update the in-degree of nodes associated with nodes 5 and 6, that is, the in-degree of node 10 is updated from 2 to 0. The third round can also simultaneously divide nodes 8 and 9 into task group 4, task group 4 {node 8, node 9}. After dividing nodes 8 and 9 into task group 4, it is necessary to continue to update the above node list. The update process can be to delete the rows of nodes 8 and 9, and update the in-degree of nodes associated with nodes 8 and 9, that is, the in-degree of nodes 15 and 16 is updated from 1 to 0.
[0107] The node list after dividing task group 3 and task group 4 can be understood by referring to Table 3:
[0108] Table 3: Node list of the directed acyclic graph after the third round of partitioning
[0109] After the third round of partitioning, the remaining nodes can be partitioned according to the above partitioning ideas until all remaining nodes are divided into task groups. After the first DAG partition, there can be eight task groups, for example: Task Group 1 {node 1, node 2}, Task Group 2 {node 3, node 4}, Task Group 3 {node 5, node 6}, Task Group 4 {node 8, node 9}, Task Group 5 {node 7, node 10}, Task Group 6 {node 13, node 11}, Task Group 7 {node 15, node 16}, Task Group 8 {node 14, node 12}.
[0110] 303. Process each task group.
[0111] In this application, tasks within the same task group are fully decoupled and independent of each other, and can be processed in parallel; if there is no dependency between different task groups, then the task groups without dependencies can also be processed in parallel, but task groups with dependencies cannot be processed in parallel.
[0112] For example, among the eight task groups above, the execution order of task group 1 {node 1, node 2} must take precedence over that of task group 2 {node 3, node 4}, because the task associated with node 4 depends on the execution result of the task associated with node 1. The execution order of task group 2 {node 3, node 4} must take precedence over that of task group 3 {node 5, node 6}, because the task associated with node 6 depends on the execution result of the task associated with node 3. If task groups 1, 2, and 3 are all executed, task group 4 {node 8, node 9} and task group 5 {node 7, node 10} can be processed in parallel, because the tasks on which the nodes in task group 4 {node 8, node 9} and task group 5 {node 7, node 10} depend have no dependencies.
[0113] In an embodiment of the present application, the in-degree of the node associated with the task in each task group is 0, indicating that the tasks in the same task group are independent of each other and do not need to generate intermediate data, which can reduce the occupancy of machine memory; in addition, the number of tasks in the task group is no more than N, which can limit the amount of output data of the task, thereby reducing the memory occupancy of the task output data; moreover, the tasks in the same task group are fully decoupled and independent of each other, and can be processed in parallel, which can improve the efficiency of task processing and reduce the end-to-end (edge to edge, E2E) delay of task processing, thereby improving the system performance of task processing.
[0114] Optionally, the above step 303 may include: establishing a correspondence between the P tasks in the first task group and the M instances according to the task scheduling goal, P≤N, and P is a positive integer, M is a positive integer, M instances are used to execute P tasks, the first task group is any one of at least two task groups, and the task scheduling goal is the shortest sum of the time for M instances to execute P tasks.
[0115] In the embodiment of the present application, the sum of the time for M instances to execute P tasks can be obtained based on the statistical information of the previous operation. It can be that the time for each instance to execute different tasks is counted in advance, and then the sum of the time for the P tasks to be executed by M instances is calculated. For each task in the task group, when assigning the corresponding instance to the task, the shortest sum of the time for the M instances to execute P tasks (min L,P finish_time i , 1≤i≤P is minimized) to schedule P tasks, which can avoid the long processing time of a task and slow down the processing of the entire task group, thereby reducing the end-to-end delay and improving resource utilization.
[0116] When P ≤ M, each task can be associated with an instance. When P > M, the P tasks within any task group can be orchestrated in at least two ways, that is, a correspondence between P tasks and M instances can be established. These two methods are described below.
[0117] 1. For each task group, divide it into subgroups;
[0118] The process may include: dividing P tasks in a first task group into M subgroups according to a task scheduling goal; wherein the absolute value of the difference between the execution time of the tasks in at least one subgroup and a time reference value is less than a first threshold, and the time reference value includes the ratio of the time for executing the P tasks and M; and establishing a correspondence between the M subgroups and the M instances.
[0119] The process may first determine a time reference value (mean_value), which may be a ratio of the sum of the times for executing P tasks to M.
[0120] If P tasks include tasks with execution times exceeding the time base value, these tasks are grouped separately, with each task with execution times exceeding the time base value being a separate subgroup. This way, when P tasks are divided into M subgroups, these M subgroups will include at least one first subgroup containing a task with execution times exceeding the time base value. In other words, when a task's execution time exceeds the time base value, it can be assigned to a separate subgroup and executed exclusively by a single instance. This prevents it from slowing down other tasks.
[0121] For tasks whose execution time is less than the time reference value, a traversal approach can be used to divide the subgroups. For example, one task can be found in the first task group, and then the other tasks can be traversed, looking for an execution time that is close to the time reference value. If the execution time of the two tasks still does not reach the time reference value, the traversal can be continued, and a third task can be added to the subgroup. The execution time of these three tasks is close to the time reference value. Whether to add new tasks to a subgroup that has not yet been divided can be determined using a distance method. The process can be that if a new task is added to the incomplete subgroup and the sum of the execution times of the tasks in the subgroup is closer to the time reference value, the new task is added to the subgroup. If the addition of the new task causes the sum of the execution times of the tasks in the subgroup to move away from (possibly exceeding) the time reference value, the new task is not added to the subgroup, and the traversal continues with the next task until all tasks in the first task group are traversed. At this point, the subgroup is divided, removed from the first task group, and the remaining tasks in the first task group are divided again. Thus, when P tasks are divided into M subgroups, these M subgroups will include at least one second subgroup, which includes at least two tasks, and the absolute value of the difference between the execution time and the time reference value of at least two tasks is less than the first threshold. In other words, when some tasks have a short execution time, at least two such tasks can be grouped into a subgroup, and at least two tasks in the subgroup can be executed by the same instance. This ensures that each instance executes tasks at roughly the same time, eliminating the need to enter a waiting state prematurely, thereby improving resource utilization.
[0122] In an embodiment of the present application, the execution time of tasks in each subgroup is as close to the time reference value as possible. That is, the absolute value of the difference between the execution time of tasks in at least one subgroup and the time reference value is as close to 0 as possible. The first threshold value can be a very small positive number close to 0, such as 0.1, 0.01, and other possible values. In the present application, P tasks are arranged into different subgroups according to the requirement that the absolute value of the difference between the execution time of tasks in at least one subgroup and the time reference value is less than the first threshold value, which is conducive to improving the efficiency of task processing.
[0123] 2. Directly establish a correspondence between each task and instance, and configure priorities for at least two tasks associated with the same instance;
[0124] The process can be: if P tasks include a task whose execution time is greater than the time reference value, an allocation relationship with a separate instance will be established for such tasks, that is, each task whose execution time is greater than the time reference value corresponds to a different instance, and these instances are no longer allocated to other tasks.
[0125] For tasks whose execution time is less than the time base value, you can first configure an instance from an instance that has not yet established an allocation relationship. Then, for this instance, use the task traversal method to continue traversing other tasks, looking for another task whose execution time of the two tasks is as close as possible to the time base value. If the execution time of the two tasks still does not reach the time base value, you can continue traversing and try to add a third task to the subgroup. The execution time of these three tasks is as close as possible to the time base value. This concept is basically the same as the concept of dividing subgroups above, so it will not be repeated here.
[0126] When two or more tasks are configured with the same instance, it is necessary to determine the execution priority of these tasks. In this way, when tasks are subsequently assigned to instances, the tasks can be assigned to the corresponding instances for execution based on the instance-task assignment relationship and priority.
[0127] This process can be understood by referring to Figure 6. As shown in Figure 6, the first task group includes four tasks: Task 1, Task 2, Task 3, and Task 4. These four tasks are assigned to three instances for execution: Instance 1, Instance 2, and Instance 3. The process for orchestrating these four tasks can be as follows: Determine the execution times of the four tasks as 9 milliseconds (ms), 6 ms, 3 ms, and 3 ms, respectively. The time base value is (9 + 6 + 3 + 3) / 3 = 7 ms.
[0128] After traversal, it is found that the execution time of Task 1, 9ms, is greater than the time base value of 7ms, while the execution times of Task 2, 6ms, 3ms, and 4, 3ms, are all less than the time base value of 7ms. In this case, Task 1 can be assigned to Instance 1. Task 2's execution time of 6ms is closer to the time base value of 7ms, and after adding the execution time of Task 3 or Task 4, the sum of the two tasks is further away from the time base value of 7ms relative to Task 2's 6ms. In this case, Task 2 is assigned to Instance 2, and Task 3 and Task 4 are then assigned to Instance 3. Since two tasks point to Instance 3, the priorities of Task 3 and Task 4 need to be configured. For example, if Task 3's priority is configured to 1 and Task 4's priority is configured to 2, Task 3 will be executed before Task 4.
[0129] The system structure of the task processing method introduced above can also be understood by referring to Figure 7. As shown in Figure 7, the system structure includes a system resource collection module 701, a task dependency analysis module 702, a task group division module 703 and an intra-group task scheduling module 704.
[0130] The system resource collection module 701 is used to collect information about system resources, such as the size of machine memory in the system, the maximum number of instances, the size of memory occupied by each instance, the size of memory occupied by output data of a single task, and other information.
[0131] The task dependency parsing module 702 may be used to determine related tasks based on the data processing requirements submitted by the user, and to parse the dependencies between the tasks to construct a directed acyclic graph.
[0132] The task group division module 703 is used to perform the task group division process, dividing the tasks associated with the nodes in the directed acyclic graph into at least two task groups. This process can be understood by referring to the relevant introduction of the above step 302.
[0133] The intra-group task scheduling module 704 is used to schedule tasks within each task group. This process can be understood by referring to the introduction to step 303 above.
[0134] The scheduled tasks in each task group are dispatched to the corresponding instances according to their priority. For more information, refer to the previous section.
[0135] The above-mentioned system resource collection module 701, task dependency analysis module 702, task group division module 703 and intra-group task arrangement module 704 can all be deployed on the control plane of the cloud system.
[0136] The task processing solution provided in the embodiment of the present application can be applied to a workflow scheduling system. As shown in FIG8 , in the workflow scheduling system, the software can provide a set of workflow description interfaces Link and Build_graph_from_json for upper-layer distributed applications. In this way, users can describe task workflows through an application programming interface (API) or a DAG workflow description file in JS object notation (JSON) format and submit it to a dynamic orchestration scheduling engine 80. The dynamic orchestration scheduling engine 80 includes a resource indicator collection module 800, a directed acyclic graph construction module 801, a task group partitioning module 802, an intra-group task orchestration module 803, and a task scheduler 804. The resource indicator collection module 800 can collect resource information from the resource management module 805. The resource information can include information such as the size of the machine memory in the system, the maximum number of instances, the size of the memory occupied by each instance, and the size of the memory occupied by the output data of a single task.
[0137] The Directed Acyclic Graph (DAG) Construction Module 801 interfaces with users, providing a Link interface for describing dependencies through an API. This Link interface is used to construct nodes in a Directed Acyclic Graph (DAG). This Link interface accepts a callable task function and its input parameters and returns a node representing the task. Unlike directly calling a function, the node created by Link is inert and does not immediately execute the task function. This allows users to freely combine multiple tasks to build a large DAG without having to execute them immediately. The DAG nodes returned by this Link interface can be used as input parameters for other tasks to define inter-task dependencies. Multiple Link calls can be nested to generate a complex DAG. The system also provides a Build_graph_from_json interface to receive a description of the entire task DAG graph in one go, eliminating the need to call the API multiple times to build dependencies. The functionality of the Directed Acyclic Graph Construction Module 801 can be further understood by referring to the functionality of the Task Dependency Parsing Module 702 in Figure 7. After the DAG is constructed, the Run interface can be called to start the DAG execution. The Run interface accepts a DAG starting node or the entire DAG as input and executes the corresponding task flow. Run will persist the execution state so that progress can be recovered in the event of a failure. The system will save the execution results of each node so that nodes that have been successfully executed do not need to be rerun.
[0138] The Run interface can connect to the task group division module 802 to trigger the preparation of the DAG scheduling plan, including task group division and intra-task group scheduling. The process of task group division and intra-task scheduling can be understood by referring to the functions of the task group division module 703 in Figure 7 and the functions of the intra-task scheduling module 704 in Figure 7.
[0139] After the orchestration is completed, it means that each task is configured with the resource instance to be allocated and the priority. After the task orchestration module 803 in the group completes the task orchestration, the orchestrated task is handed over to the task scheduler 804. The task scheduler 804 will select one task group at a time to start scheduling based on the divided task groups, and send the task to the corresponding resource instance to queue. The resource instance will execute the task.
[0140] The scheduling process can be understood by referring to FIG9 . As shown in FIG9 , there are two tasks, task 1 (task_1) and task 2 (task_2), and both task 1 and task 2 point to instance 1 (ins_1). Among them, the priority of task 1 is 1 (priority: 1) and the priority of task 2 is 2 (priority: 2). Then, when scheduling task 1 and task 2, the task scheduler 804 can schedule both task 1 and task 2 to the task queue of ins_1, and schedule task 1 first and then task 2 in order of priority.
[0141] The resource management module 805 manages specific resource instances, such as: pulling up instance 1 to instance M. The resource management module 805 executes tasks according to the tasks sent by the task scheduler 804. The resource management module 805 can put tasks into the task queue in order of priority. The instance continuously takes out tasks from the queue from the loop and starts execution.
[0142] The system structure shown in FIG8 can shield the underlying hardware differences and achieve resource isolation.
[0143] The task processing solution provided in the embodiment of the present application can also be applied to financial systems. Taking the offline calculation of financial quantitative factors of the financial system as an example, the dependencies between tasks in this scenario are complex. As shown in Figure 10, multiple layers of dependent tasks may be nested under one task. Figure 10 illustrates three layers of nested tasks for one task. Figure 10 is just an example. In practice, more layers of tasks may be nested.
[0144] The task processing process in this scenario can be understood by referring to Figure 11. As shown in Figure 11, the user program (user function) can be distributed on multiple machines. Different machines can call the source data reading method read_highfreq_from_assets to read the financial minute-level source data and store it for subsequent factor calculation.
[0145] In the formal factor calculation phase, users begin to submit tasks in batches in a loop, calling the Link interface. The input parameters are the upstream factors that the factor depends on, as well as the source data it depends on. The output results can be used as input parameters for downstream factors. The runtime function at the system level can submit tasks to the dependency management module (Dependency Manager) to resolve upstream and downstream dependencies, build a DAG, and then trigger the workflow management (Workflow Manager) module to perform task group division and intra-group task orchestration on the DAG. After executing the intra-group orchestration, each task will be assigned the priority attribute and the ins attribute, which specify the relative order of the tasks and the location of the resource node to be sent. The Workflow Manager module calls the task submission (submit_task) internal method to send the task to the task queue of the task management (task manager) module for resource allocation.
[0146] In the Task Manager module, the task queue continuously tries to pop out tasks. After successfully popping out a task, it will try to call the internal method ins_mgr.schedule to pop out a corresponding instance. If it is also successfully popped out, it will call the internal ins.add_task method to insert the task into the instance's queue according to priority.
[0147] Then, through the Invoke Client module, which connects the task scheduling layer with the resource allocation layer, all relevant task arrangements can be sent to the resource management kernel for task execution.
[0148] The above describes the task processing method. The following describes the task processing device provided in the embodiment of the present application with reference to the accompanying drawings.
[0149] As shown in FIG12 , the task processing apparatus 120 provided in an embodiment of the present application includes:
[0150] The acquisition unit 1201 is used to acquire a first directed acyclic graph for describing a data processing task; the first directed acyclic graph includes multiple nodes and multiple edges for connecting the nodes, wherein each node is used to represent a task, and the edge is used to represent a dependency relationship between two connected nodes.
[0151] The first processing unit 1202 is used to divide the multiple nodes into at least two task groups; wherein the number of tasks in each task group is not greater than N, and the in-degree of the nodes associated with the tasks in each task group is zero, N is determined by the size of the machine memory, the size of the memory occupied by the instance, and the size of the memory occupied by the output data of the task, and N is a positive integer.
[0152] The second processing unit 1203 is configured to process each task group.
[0153] Optionally, the first processing unit 1202 is specifically used to: divide the nodes with zero in-degree in the first directed acyclic graph into at least one task group to obtain a second directed acyclic graph, and update the in-degree of the nodes in the second directed acyclic graph; divide the nodes with zero in-degree in the second directed acyclic graph into at least one task group.
[0154] Optionally, the first processing unit 1202 is specifically used to: when the number of nodes with zero in-degree in the first directed acyclic graph is not greater than N, divide the nodes with zero in-degree into a task group; when the number of nodes with zero in-degree in the first directed acyclic graph is greater than N, divide no more than N nodes that meet the affinity requirements into the same task group.
[0155] Optionally, the second processing unit 1203 is specifically used to: establish a corresponding relationship between the P tasks in the first task group and the M instances according to the task scheduling target, P≤N, and P is a positive integer, M is a positive integer, M instances are used to execute P tasks, the first task group is any one of at least two task groups, and the task scheduling target is the shortest sum of the time for M instances to execute P tasks.
[0156] Optionally, the second processing unit 1203 is specifically used to: when P>M, divide the P tasks in the first task group into M subgroups according to the task scheduling target; wherein the absolute value of the difference between the execution time of the tasks in at least one subgroup and the time reference value is less than a first threshold, and the time reference value includes the time for executing the P tasks and the ratio to M; establish a correspondence between the M subgroups and the M instances.
[0157] Optionally, the M subgroups include at least one first subgroup, the first subgroup includes a task, and the execution time of the task is greater than the time reference value.
[0158] Optionally, the M subgroups include at least one second subgroup, the second subgroup includes at least two tasks, and the absolute value of the difference between the execution time of the at least two tasks and the time reference value is smaller than the first threshold.
[0159] Optionally, the second processing unit 1203 is specifically used to: when P>M, establish an allocation relationship between each of the P tasks and the M instances; wherein the absolute value of the difference between the execution time of at least two tasks pointing to the same instance and the time reference value is less than a first threshold, and the time reference value includes the time for executing the P tasks and the ratio to M; configure priority for at least two tasks pointing to the same instance, and the priority is used to indicate the order in which at least two tasks pointing to the same instance are allocated to the corresponding instance.
[0160] Optionally, the acquisition unit 1201 is further configured to acquire the maximum number of instances, the memory size occupied by each instance, the memory size occupied by output data of a single task, and the size of the machine memory; and determine N using the following relationship:
[0161] Among them, X is the maximum number of instances, MemAllocate insi ResSize is the memory size occupied by each instance taskj The size of the memory occupied by the output data of a single task, Limit mem The size of the machine's memory.
[0162] Optionally, the acquisition unit 1201 is specifically configured to receive a demand for processing data, where the demand is used to determine tasks to be executed; parse dependencies between tasks to be executed, and construct a first directed acyclic graph.
[0163] In another embodiment of the present application, a computer-readable storage medium is also provided, in which computer execution instructions are stored. When the processor of a computer device executes the computer execution instructions, the computer device executes the steps performed by the task processing device in Figures 3 to 11 above.
[0164] In another embodiment of the present application, a computer program product is provided. The computer program product includes computer program code. When the computer program code is executed on a computer, the computer device executes the steps executed by the task processing apparatus in Figures 3 to 11 above.
[0165] In another embodiment of the present application, a chip system is also provided, which includes one or more interface circuits and one or more processors; the interface circuits and the processors are interconnected by lines; the interface circuits are used to receive signals from the memory of the computer device and send signals to the processor, and the signals include computer instructions stored in the memory; when the processor executes the computer instructions, the computer device executes the steps performed by the task processing device in Figures 3 to 11 above. In one possible design, the chip system may also include a memory, which is used to store program instructions and data necessary for the client. The chip system can be composed of chips, or it can include chips and other discrete devices.
[0166] In another embodiment of the present application, a computer device cluster is also provided, including at least one computer device, each computer device including a processor and a memory; the processor of the at least one computer device is used to execute instructions stored in the memory of the at least one computer device, so that the computer device cluster performs the steps performed by the task processing device in Figures 3 to 11 above.
[0167] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0168] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0169] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in whole or in part through software, hardware, firmware, or any combination thereof.
[0170] When software is used to implement the integrated unit, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state drive (SSD)).
Claims
1. A task processing method, characterized in that: include: Obtaining a first directed acyclic graph for describing a data processing task; the first directed acyclic graph includes a plurality of nodes and a plurality of edges for connecting the nodes, wherein each node is used to represent a task, and the edge is used to represent a dependency relationship between two connected nodes; Divide the multiple nodes into at least two task groups; wherein the number of tasks in each task group is not greater than N, and the in-degree of the nodes associated with the tasks in each task group is zero, wherein N is determined by the size of the machine memory, the size of the memory occupied by the instance, and the size of the memory occupied by the output data of the task, and N is a positive integer; Each of the task groups is processed.
2. The method according to claim 1, characterized in that The dividing the plurality of nodes into at least two task groups comprises: Dividing the nodes with zero in-degree in the first directed acyclic graph into at least one task group to obtain a second directed acyclic graph, and updating the in-degree of the nodes in the second directed acyclic graph; The nodes with zero in-degree in the second directed acyclic graph are divided into at least one task group.
3. The method according to claim 2, characterized in that The dividing the nodes with zero in-degree in the first directed acyclic graph into at least one task group includes: When the number of nodes with zero in-degree in the first directed acyclic graph is not greater than N, the nodes with zero in-degree are divided into a task group; When the number of nodes with zero in-degree in the first directed acyclic graph is greater than N, no more than N nodes meeting the affinity requirements are divided into the same task group.
4. The method according to any one of claims 1 to 3, characterized in that: The processing of each task group includes: Establish a corresponding relationship between the P tasks in the first task group and the M instances according to the task scheduling target, P≤N, and P is a positive integer, M is a positive integer, the M instances are used to execute the P tasks, the first task group is any one of the at least two task groups, and the task scheduling target is the shortest sum of the time for the M instances to execute the P tasks.
5. The method according to claim 4, characterized in that When P>M, the step of establishing a corresponding relationship between the P tasks in the first task group and the M instances according to the task scheduling strategy includes: Divide the P tasks in the first task group into M subgroups according to the task scheduling target; wherein the absolute value of the difference between the execution time of the tasks in at least one subgroup and the time reference value is less than a first threshold, and the time reference value includes the ratio of the time for executing the P tasks and M; A corresponding relationship between the M subgroups and the M instances is established.
6. The method according to claim 5, characterized in that The M subgroups include at least one first subgroup, the first subgroup includes a task, and the execution time of the task is greater than the time reference value.
7. The method according to claim 5 or 6, characterized in that: The M subgroups include at least one second subgroup, the second subgroup includes at least two tasks, and the absolute values of the execution times of the at least two tasks and the differences with the time reference value are less than a first threshold.
8. The method according to claim 4, characterized in that When P>M, the step of establishing a corresponding relationship between the P tasks in the first task group and the M instances according to the task scheduling strategy includes: Establishing an allocation relationship between each of the P tasks and the M instances; wherein the absolute value of the difference between the execution time of at least two tasks pointing to the same instance and a time reference value is less than a first threshold, and the time reference value includes a ratio of the execution time of the P tasks to M; A priority is configured for the at least two tasks pointing to the same instance, where the priority is used to indicate the order in which the at least two tasks pointing to the same instance are allocated to the corresponding instance.
9. The method according to any one of claims 1 to 8, characterized in that: The method further comprises: Get the maximum number of instances, the memory size of each instance, the memory size of the output data of a single task, and the size of the machine memory; The N is determined by the following relationship: Among them, X is the maximum number of instances, MemAllocate insi The memory size occupied by each instance, ResSize taskj The size of the memory occupied by the output data of a single task, Limit mem The size of the machine's memory.
10. The method according to any one of claims 1 to 8, characterized in that: The obtaining of a first directed acyclic graph for describing the data processing task includes: receiving a requirement for processing the data, wherein the requirement is used to determine a task to be performed; The dependencies between the tasks to be executed are parsed to construct the first directed acyclic graph.
11. A task processing device, characterized in that: include: An acquisition unit, configured to acquire a first directed acyclic graph for describing a data processing task; the first directed acyclic graph includes a plurality of nodes and a plurality of edges for connecting the nodes, wherein each node is used to represent a task, and the edge is used to represent a dependency relationship between two connected nodes; A first processing unit is used to divide the multiple nodes into at least two task groups; wherein the number of tasks in each task group is not greater than N, and the in-degree of the node associated with the task in each task group is zero, wherein N is determined by the size of the machine memory, the size of the memory occupied by the instance, and the size of the memory occupied by the output data of the task, and N is a positive integer; The second processing unit is used to process each task group.
12. The device according to claim 11, characterized in that The first processing unit is specifically configured to: Dividing the nodes with zero in-degree in the first directed acyclic graph into at least one task group to obtain a second directed acyclic graph, and updating the in-degree of the nodes in the second directed acyclic graph; The nodes with zero in-degree in the second directed acyclic graph are divided into at least one task group.
13. The device according to claim 12, characterized in that The first processing unit is specifically configured to: When the number of nodes with zero in-degree in the first directed acyclic graph is not greater than N, the nodes with zero in-degree are divided into a task group; When the number of nodes with zero in-degree in the first directed acyclic graph is greater than N, no more than N nodes meeting the affinity requirements are divided into the same task group.
14. The device according to any one of claims 11 to 13, characterized in that: The second processing unit is specifically used for: Establish a corresponding relationship between the P tasks in the first task group and the M instances according to the task scheduling target, P≤N, and P is a positive integer, M is a positive integer, the M instances are used to execute the P tasks, the first task group is any one of the at least two task groups, and the task scheduling target is the shortest sum of the time for the M instances to execute the P tasks.
15. The device according to claim 14, characterized in that When P>M, the second processing unit is specifically used for: Divide the P tasks in the first task group into M subgroups according to the task scheduling target; wherein the absolute value of the difference between the execution time of the tasks in at least one subgroup and the time reference value is less than a first threshold, and the time reference value includes the ratio of the time for executing the P tasks and M; A corresponding relationship between the M subgroups and the M instances is established.
16. The device according to claim 14, characterized in that When P>M, the second processing unit is specifically used for: Establishing an allocation relationship between each of the P tasks and the M instances; wherein the absolute value of the difference between the execution time of at least two tasks pointing to the same instance and a time reference value is less than a first threshold, and the time reference value includes a ratio of the execution time of the P tasks to M; A priority is configured for the at least two tasks pointing to the same instance, where the priority is used to indicate the order in which the at least two tasks pointing to the same instance are allocated to the corresponding instance.
17. The device according to any one of claims 11 to 16, characterized in that: The acquisition unit is further used to acquire the maximum number of instances, the size of memory occupied by each instance, the size of memory occupied by output data of a single task, and the size of machine memory; and determine N through the following relationship; Among them, X is the maximum number of instances, MemAllocate insi The memory size occupied by each instance, ResSize taskj The size of the memory occupied by the output data of a single task, Limit mem The size of the machine's memory.
18. A task processing device, characterized in that: include: A communication interface, a processor and a memory, wherein the communication interface and the processor are coupled to the memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the task processing device executes the method as described in any one of claims 1-10.
19. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed on a computer device, the computer device executes the method according to any one of claims 1 to 10.
20. A computer program product, characterized in that The computer program product comprises a computer program code, and when the computer program code is run on a computer device, the computer device is caused to perform the method according to any one of claims 1 to 10.
21. A chip system, characterized in that: The chip system includes one or more interface circuits and one or more processors; the interface circuits and the processors are interconnected by lines; the interface circuits are used to receive signals from the memory of the task processing device and send signals to the processor, the signals including computer instructions stored in the memory; when the processor executes the computer instructions, the task processing device executes the method as described in any one of claims 1-10.
22. A computer equipment cluster, characterized in that: comprising at least one computer device, each computer device comprising a processor and a memory; The processor of the at least one computer device is configured to execute instructions stored in the memory of the at least one computer device, so that the computer device cluster executes the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Task processing method and corresponding device
CN120162121A
Method and device for server task scheduling
CN106897132A
Task scheduling method, device and equipment, and storage medium
CN111984390A
Task calling method and device, electronic equipment and storage medium
CN113535363A
Task scheduling method, system and equipment based on directed acyclic graph and storage medium
CN114625507A
Cited By
Cloud platform operation and maintenance method, device, equipment, medium and product
CN120872745A
Dynamic DAG arrangement method and system based on operator capability portrait
CN121387519A