Feedback-based task allocation
By using the controller to receive feedback information in the task management framework, learning to evaluate computing resources and workloads, the problem of inefficient task allocation in the existing technology is solved, and more efficient and fair task allocation is achieved.
Patent Information
- Application Number
- CN202411740518.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-11-29
- Publication Date
- 2025-07-01
AI Technical Summary
The existing task management framework lacks effective visibility into the work nodes or compute nodes when allocating tasks, resulting in inefficient task allocation.
The controller receives feedback information from work nodes or computing nodes, and learns to evaluate the amount of computing resources required to complete the task and the workload of the node, thereby allocating new tasks more accurately.
Improve the efficiency and fairness of task allocation, ensure that tasks are optimally allocated to work nodes, and adapt to changes in task types and resource requirements.
Smart Images

Figure CN120234110A_ABST
Abstract
Description
[0001] This application claims priority to U.S. Patent Application No. 18 / 400,349, filed on December 29, 2023, the entire content of which is incorporated herein by reference. Technical Field
[0002] This disclosure relates to computing systems, and more particularly, to techniques for allocating tasks in a task or job management framework. Background Art
[0003] In a typical task management framework, a controller allocates tasks and / or workloads to computing nodes for execution and attempts to efficiently use the available computing resources. Tasks can be allocated based on information such as the memory usage or CPU usage of the computing nodes included within the computing resources. Additionally, in some cases, tasks may be allocated according to a partitioning system (e.g., Apache Kafka), in which the tasks to be executed are divided or partitioned before being allocated across the available computing nodes.
[0004] Task management frameworks can be helpful in a variety of situations, including when collecting data, such as in the case of providing full-stack observability by integrating with various application performance monitoring tools and performing cross-layer analysis in a data center or other computing environment. In such a case, the controller can collect data from the hardware and software systems within the data center. Typically, the amount of data to be collected is large, which presents challenges. Summary of the Invention
[0005] This disclosure describes techniques for allocating tasks to computing resources in a computing environment, such as a multi-tenant data center. Generally, at least some of the techniques described herein can be characterized as improvements to the process of allocating tasks in a task management framework.
[0006] Task management solutions typically involve a controller allocating tasks to worker nodes or computing nodes in an efficient manner. The techniques described herein provide the controller with improved visibility into the work being performed by the worker nodes or computing nodes, which allows the controller to more effectively allocate tasks to the nodes. As described herein, the controller obtains such visibility through feedback received from the worker nodes or computing nodes. Based on this feedback, the controller learns over time how to accurately evaluate the amount of computing resources required to complete a given type of task. The controller can also learn over time how to accurately evaluate the workload of the nodes available for task execution. As a result of evaluating the feedback on completed tasks and how the feedback varies according to task type, the controller is able to better allocate new tasks to the available nodes efficiently.
[0007] In some examples, the present disclosure describes operations performed by a computing system in accordance with one or more aspects of the present disclosure. In a particular example, the present disclosure describes a method, including: receiving, by a controller, a first set of tasks, wherein each task in the first set of tasks has one of a plurality of task types; assigning, by the controller, each task in the first set of tasks to a worker node for processing by the worker node; receiving, by the controller and for at least some of the tasks in the first set of tasks, feedback information regarding the processing by the worker node; determining, by the controller and based on the feedback information, an expected processing volume associated with each of the plurality of task types; receiving, by the controller, a second set of tasks, wherein each task in the second set of tasks has one of a plurality of task types; and assigning, by the controller and based on the expected processing volume associated with each task type, each task in the second set of tasks to a worker node for processing.
[0008] In another example, the present disclosure describes a system including a storage system and processing circuitry having access to the storage system, wherein the processing circuitry is configured to perform the operations described herein. In yet another example, the present disclosure describes a computer-readable storage medium including instructions that, when executed, configure processing circuitry of a computing system to perform the operations described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 is a conceptual diagram illustrating an exemplary system for task allocation in a job management framework in a data center in accordance with one or more aspects of the present disclosure.
[0010] Figure 2 is a block diagram illustrating an exemplary system for task allocation to worker nodes in a computing environment in accordance with one or more aspects of the present disclosure.
[0011] Figure 3 is a flowchart illustrating operations performed by an exemplary controller in accordance with one or more aspects of the present disclosure. DETAILED DESCRIPTION
[0012] This disclosure outlines techniques that include providing a controller in a task management framework with better visibility into the work being performed by worker nodes or compute nodes. This visibility enables the controller to more effectively assign tasks to the nodes. Such techniques improve upon traditional techniques where the controller uses only externally visible information, such as CPU utilization and memory usage, when assigning tasks to worker nodes or compute nodes. As described herein, a worker node can correspond to a virtual execution instance running on a compute node. In such an example, multiple worker nodes can execute on a single compute node (e.g., as multiple container pools execute on a Kubernetes compute node). Although the techniques are primarily described herein in terms of worker nodes executing as instances on compute nodes, the techniques described herein can be applied to other scenarios. Such other contexts can include allocating work across compute nodes or across computing resources in other computing architectures or arrangements.
[0013] As described herein, the controller can learn to make accurate assessments of the amount of computing resources required to complete a task (i.e., the "weight" of the task). Additionally, and also as described herein, the controller can more effectively evaluate the nature of the work already being performed on a worker node or compute node, enabling the controller to more accurately determine the availability of any given worker node or compute node to be assigned a new task of a given type. Further, techniques are described herein for ensuring fair access to computing resources in a multi-tenant computing environment, preventing any one tenant from having an adverse impact on the workloads that another tenant is attempting to complete.
[0014] These techniques are primarily described herein in the context of solutions related to providing full-stack observability and cross-layer analysis in a computing system (such as a data center). Such observability is typically a key component in providing sufficient visibility into network operations, performance, and health. In some examples, the systems described herein employ a telemetry data collector framework that retrieves telemetry data from multiple systems and stores the data for further processing. A large amount of data typically needs to be collected, but the techniques described herein enable the framework to scale as needed. In at least some examples, it may be important to collect data in a near or seemingly near real-time manner because the analysis process and other processes can evaluate each data point and may need to act quickly on the data. Thus, the ability of the framework to collect data efficiently can be very important. Although the techniques described herein are generally described in the context of data collection for observability solutions, such techniques can also be applied to other scenarios, such as those related to task management frameworks operating in other domains.
[0015] Figure 1FIG. 0 is a conceptual diagram illustrating an exemplary system for task allocation in a job management framework in a data center in accordance with one or more aspects of the present disclosure. Generally, data center 110 provides an operating environment for system 100. Data center 110 may represent an on-premises environment, a private cloud, a hybrid cloud, or a multi-tenant cloud environment. In a multi-tenant cloud environment, applications and services operate on behalf of one or more tenants or customer sites 11 (shown as "Customer 11"), which have one or more customer networks coupled to the data center via a service provider network 7. For example, data center 110 may host infrastructure equipment such as networking and storage systems, redundant power supplies, and environmental controls. Service provider network 7 is coupled to a public network 4 that may represent one or more networks managed by other providers, and through which a large-scale public network infrastructure (e.g., the Internet) may be formed. For example, public network 4 may represent a local area network (LAN), a wide area network (WAN), the Internet, a virtual LAN (VLAN), an enterprise LAN, a Layer 3 virtual private network (VPN), an Internet protocol (IP) intranet operated by a service provider that operates service provider network 7, an enterprise IP network, or some combination thereof.
[0016] In some examples, data center 110 may represent one of a number of geographically distributed data centers. As Figure 1 illustrated in the example of, data center 110 may be a facility that provides network services to customers. Customers of the service provider may be general entities such as enterprises and governments or individuals. For example, the data center may host network services for several enterprises and end users. Other exemplary services may include data storage, virtual private networks, traffic engineering, file services, data mining, scientific or supercomputing, etc. Although shown as a separate edge network of service provider network 7, elements of data center 110 such as one or more physical network functions (PNFs) or virtualized network functions (VNFs) may be included within the core of service provider network 7.
[0017] In Figure 1In the example shown, data center 110 includes storage and / or compute servers interconnected via a fabric 14 provided by one or more layers of physical network switches and routers, and depicts any number of servers 112A through 112J (referred to herein as "servers 112") coupled to top-of-rack (TOR) switches 116A through 116M. Servers 112 may also be referred to herein as "hosts" or "host devices". Data center 110 may include many additional servers coupled to other TOR switches 116 of data center 110. Each of servers 112 may host one or more virtual computing or execution instances 113 (e.g., instances 113A through 113N, referred to herein as "instances 113"), which may be containers, a Kubernetes container pool of one or more containers, virtual machines, or other instances. In some cases, as referred to herein, a "compute node" or "worker node" may correspond to one or more instances 113 and / or one or more servers 112.
[0018] In the example shown, fabric 14 includes interconnecting top-of-rack (or other "leaf") switches 116A through 116N (collectively referred to as "TOR switches 116") of a distribution layer coupled to chassis (or "spine" or "core") routers or switches 18A through 18F (collectively referred to as "chassis switches 18"). Although Figure 1 not specifically shown in, data center 110 may also include, for example, one or more non-edge switches, routers, hubs, gateways, security devices (such as firewalls, intrusion detection, and / or intrusion prevention devices), servers, computer terminals, laptops, printers, databases, wireless mobile devices such as cellular phones or personal digital assistants, wireless access points, bridges, cable modems, application accelerators, or other network devices. The techniques described herein may be applied to any of these systems or devices.
[0019] In Figure 1In the example shown, the TOR switch 116 and the chassis switch 18 provide redundant (multi-homed) connections to the IP fabric 20 and the service provider network 7 to the servers 112. The chassis switch 18 aggregates traffic flows and provides connections between the TOR switches 116. The TOR switch 116 can be a network device that provides Layer 2 (MAC) and / or Layer 3 (e.g., IP) routing and / or switching functions. The TOR switch 116 and the chassis switch 18 can each include one or more processors and memories, and can execute one or more software processes. The chassis switch 18 is coupled to the IP fabric 20, which can perform Layer 3 routing to route network traffic between the data center 110 and the customer site 11 through the service provider network 7. The switching fabric of the data center 110 is merely an example, and other architectures are possible. For example, other switching architectures can have more or fewer switching layers.
[0020] Each of the servers 112 can be a compute node, an application server, a storage server, or other types of servers. For example, each of the servers 112 can represent a computing device configured to operate according to the techniques described herein. The servers 112 can provide network function virtualization infrastructure (NFVI) for an NFV architecture. The servers 112 can host endpoints of one or more virtual networks operating on the physical network represented herein by the IP fabric 20 and the switching fabric 14. Although described primarily with respect to a data center-based switching network, other physical networks such as the service provider network 7 can underlie one or more virtual networks.
[0021] Figure 1 A system is shown in which a controller 140 receives a task flow 101 (i.e., tasks 101A to 101N, referred to herein as "task 101") and dispatches these tasks to worker nodes or other computing resources within the data center 110. In Figure 1 the example shown, the controller 140 can dispatch each of the tasks 101 to one or more of the instances 113 each executing on one of the servers 112. For example, the controller 140 can assign task 101A to the instance 113A executing on the server 112A. Generally, the controller 140 can assign task 101N to the instance 113N that can operate on any of the servers 112. Once task 101A is assigned to the instance 113A, the instance 113A processes task 101A. Similarly and generally, the instance 113N processes task 101N.
[0022] Instance 113 completes tasks 101A through 101N (referred to herein as "task 101") and outputs feedback 102A through 102N (referred to herein as "feedback 102") to controller 140. For example, when instance 113A completes underlying task 101A, instance 113A outputs feedback 102A to controller 140. Similarly, when instance 113N completes underlying task 101N, instance 113N outputs feedback 102N to controller 140. In each case, feedback 102 can include information about the computing resources consumed in performing underlying task 101. Feedback 102 can include information about processing cycles, memory used, other resources used, the time to complete underlying task 101, and other information about processing a given task 101. Generally, after instance 113 completes underlying task 101, each of instances 113 transmits feedback 102 to controller 140. However, in some cases, instance 113 can alternatively or additionally provide feedback 102 before a given task 101 is fully complete.
[0023] Controller 140 collects each instance of feedback 102 and associates each instance with the corresponding task 101. Controller 140 also determines the type of task being performed (or "task type"). As described herein, task 101 can be classified as one of multiple types, where the type is associated with the nature of the task being performed. In the case where each of tasks 101 represents a process for collecting data, the type associated with a given task 101 can correspond to the type of data being collected. However, in other cases, the type of each task 101 may otherwise depend on the nature of the task or the overall goal of the work that controller 140 or a given tenant is attempting to complete.
[0024] The controller 140 evaluates the feedback 102 associated with each of the tasks 101 and determines the amount of computing resources consumed by the underlying task 101. Based on this information, the controller 140 generates or updates a model that enables the controller 140 to predict the expected amount of computing resources that will be consumed by a given type of task. As described herein, this expected amount of computing resources can be regarded as the "weight" of a given task. When the controller 140 receives additional tasks 101, the controller 140 predicts the weight of each additional task 101 to determine how to dispatch and / or assign the new tasks 101 to one of the instances 113. In some cases, the controller 140 may also predict the extent to which each instance 113 is available to process the additional tasks 101. To make such predictions, the controller 140 uses information about the tasks previously assigned to the instances 113 to predict or determine the workload (or total weight) of each instance 113. Based on these evaluations performed by the controller 140, the controller 140 assigns the new tasks 101 to the server 112 and / or the instances 113 in an efficient manner and / or in a manner that can efficiently process the tasks 101.
[0025] Then, the newly assigned tasks 101 are executed, the corresponding instances 113 generate additional feedback 102, and the instances 113 send the additional feedback 102 about these newly assigned tasks 101 to the controller 140. Then, the controller 140 uses the additional feedback 102 to further improve its ability to predict the weight of each incoming task 101 based on, for example, the type of each of the tasks 101. This process may continue indefinitely. Thus, the controller 140 can use both the tasks 101 and the corresponding feedback 102 as part of a closed-loop system that is capable of efficiently allocating tasks within the data center 110 based on historical information about the processing of the tasks 101 by the instances 113.
[0026] In addition, the controller 140 can also use the ability to predict the weights of the various tasks 101 to evaluate when and how to scale the computing resources within the data center 110. For example, the controller 140 can determine that an incoming task stream 101 will require additional resources (e.g., additional instances 113). In such an example, the controller 140 can instantiate additional instances 113 within the data center 110. Similarly, the controller 140 can determine that the available resources within the data center 110 are sufficient for the incoming task stream 101. In this latter example, the controller 140 can de-dispatch the various instances 113 so that their resources are available for other uses within the data center 110.
[0027] In some cases, the controller 140 generates predictions using a machine learning model that is trained to predict the weight of task 101 based on attributes of the task 101, such as the "type" of the task. As described herein, the "type" of a task may relate to the nature of the task to be performed or the type of resources used to complete the task.
[0028] In an example where the controller 140 applies a machine learning model to predict task weights, the machine learning model used by the controller 140 may be trained and / or maintained by a separate computing system (the "training system" - Figure 1 not specifically shown in the figure). The machine learning model may include one or more neural networks, such as one or more of a deep neural network (DNN) model, a recurrent neural network (RNN) model, and / or a long short-term memory (LSTM) model. Generally, DNNs and RNNs learn from data as feature vectors, and LSTMs learn from sequential data.
[0029] The training system may implement other types of machine learning to train the machine learning model. For example, the training system may apply one or more of nearest neighbor, naive Bayes, decision tree, linear regression, support vector machine, neural network, k-means clustering, Q-learning, temporal difference, deep adversarial network, or other supervised, unsupervised, semi-supervised, or reinforcement learning algorithms to train the machine learning model.
[0030] The machine learning model processes training data for training the machine learning model, data for prediction, or other data. The machine learning model may be trained in this way to identify patterns in task attributes, instance attributes, or other factors that may affect resource utilization and workload over time.
[0031] The techniques described herein may provide certain technical advantages. For example, by assigning different weights to different tasks and tracking the weight-based load of each worker node, the controller 140 may be able to achieve more efficient and fair work distribution. Such efficiency may be achieved through a closed-loop mechanism that enables the controller 140 to re-calibrate the weights of various tasks and efficiently allocate the workload based on these calculated weights.
[0032] In addition, the closed-loop mechanism operates as a self-learning system that improves over time and tends to optimally assign tasks to worker nodes. Such a system is also able to efficiently and accurately determine when and to what extent the framework should be scaled in response to increased or decreased demand or workload (e.g., increasing or decreasing the resources or nodes available to the framework).
[0033] The techniques described herein can also be used in multi-tenant management applications that require features such as task prioritization (with fairness) and dynamic scalability. Task prioritization can be used to ensure that resources consumed by the workload of one tenant do not starve another tenant of resources.
[0034] Existing solution task management systems often cannot efficiently allocate load while taking into account real-time processing limitations. For example, in traditional solutions, partitioning / message assignment to consumers within a consumer group (e.g., in Kafka) is typically implemented in a round-robin fashion, which makes it relatively fixed over the lifetime of the consumers unless the number of consumers changes. Thus, in at least some prior solutions, prioritizing messages that need to be processed urgently (i.e., having a higher priority than other messages) is a challenge because it is difficult or impossible to change priorities over time.
[0035] Figure 2 is a block diagram illustrating an exemplary system for assigning tasks to worker nodes in a computing environment according to one or more aspects of the present disclosure. Figure 2 The data center 210 includes some of the same elements as the data center 110 described in Figure 1 In Figure 2 the controller 240 can be regarded as an example or alternative implementation of the controller 140 in Figure 1 although other implementations are possible.
[0036] Figure 2 The collector 213 shown in Figure 1 can be regarded as an exemplary implementation of the instance 113 (or server 112) in Figure 2 Similarly, the other elements shown in
[0037] In Figure 2 the example of, the data center 210 provides full-stack observability and cross-layer analysis for the work done by worker nodes in the data center 210. In Figure 2 the worker nodes correspond to collectors 213A through 213N (“collectors 213”). Figure 2 The data collection performed in Figure 2In this case, the controller 240 assigns task 101 to the collector 213, where each task specifies the type of information that should be collected from a specific APM 225. Multiple collector instances (i.e., collectors 213A through 213N) can retrieve different types of data (such as metrics, logs, traces) from the APM 225 (or other systems). Some examples of APM 225 are applications such as NewRelic or DynaTrace.
[0038] Each collector 213 communicates with the controller 240 via a two-way connection (e.g., a remote procedure call or RPC / gRPC), enabling the exchange of control information and other information. Whenever a new collector 213 is instantiated, the collector 213 establishes a control channel with the controller 240 (e.g., via a gRPC connection). Thus, such a connection is used to transfer task 101 to the collector 213, and the collector 213 can use this connection to transfer feedback 102 back to the controller 240.
[0039] Figure 2 The controller 240 is shown in order to facilitate the description of certain components, modules, and other aspects of a computing system in which a system for managing task assignment can be implemented. Figure 2 The controller 240 is also shown in order to facilitate the description of how such a computing system can operate in accordance with the techniques described herein.
[0040] For ease of illustration, the controller 240 is depicted in Figure 2 as a single computing system. However, in other examples, the controller 240 can be implemented by multiple devices or computing systems distributed across one data center, multiple data centers, or multiple cloud networks. For example, separate computing systems can implement the functions performed by each of the receiving module 251, the calibration module 252, and the assignment module 253 described herein. Alternatively or additionally, Figure 2 the modules included within the controller 240 as shown in can be implemented by distributed virtualized computing instances (e.g., virtual machines, containers) of a data center, a cloud computing system, a server farm, and / or a server cluster. Although shown operating within a single data center 210, the controller 240 can operate outside of the data center 210 and can operate to manage and / or distribute tasks across multiple data centers 210 or computing environments.
[0041] In Figure 2In [the figure], controller 240 is shown as having underlying physical hardware that includes a power supply 242, one or more processors 244, one or more communication units 245, one or more input devices 246, one or more output devices 247, and one or more storage devices 250. The storage device 250 may include a receiving module 251, a calibration module 252, an assignment module 253, and a data memory 259. The storage device 250 may also be used to store a task list 258, although in some cases, such a task list 258 may be included within the data memory 259. The storage device 250 may also be used to store task 101 and feedback 102.
[0042] One or more of the devices, modules, storage areas, or other components of controller 240 may be interconnected to enable communication (physically, communicatively, and / or operatively) between components. In some examples, such a connection may be provided via a communication channel that may include a system bus (e.g., communication channel 243), a network connection, an interprocess communication data structure, or any other method for transferring data.
[0043] The power supply 242 of controller 240 may supply power to one or more components of controller 240. The power supply 242 may receive power from a main alternating current (AC) power supply in a building, data center, or other location. In some examples, the power supply 242 may include a battery or device that supplies direct current (DC). The power supply 242 may have intelligent power management or consumption capabilities, and such features may be controlled, accessed, or adjusted by the processor 244 to intelligently consume, allocate, supply, or otherwise manage power.
[0044] One or more processors 244 of controller 240 may implement functions associated with controller 240 or with one or more of the modules shown and / or described herein and / or execute instructions associated therewith. One or more processors 244 may be a processing circuit that performs operations in accordance with one or more aspects of the present disclosure, may be a part of the processing circuit, and / or may include the processing circuit.
[0045] One or more communication units 245 of controller 240 may communicate with devices external to controller 240 by sending and / or receiving data and, in some aspects, may operate as both an input device and an output device. In some or all cases, the communication unit 245 may communicate with other devices or computing systems via a network.
[0046] One or more input devices 246 may represent any input device of the controller 240, and one or more output devices 247 may represent any output device of the controller 240. The input device 246 and / or the output device 247 may generate, receive, and / or process outputs from any type of device capable of outputting information to a human or a machine. For example, one or more input devices 246 may generate, receive, and / or process inputs in the form of electrical, physical, audio, image, and / or visual inputs (e.g., peripherals, keyboards, microphones, cameras). Correspondingly, one or more output devices 247 may generate, receive, and / or process outputs in the form of electrical and / or physical outputs (e.g., peripherals, actuators).
[0047] One or more storage devices 250 within the controller 240 may store information for processing during the operation of the controller 240. The storage device 250 may store program instructions and / or data associated with one or more of the modules described in accordance with one or more aspects of the present disclosure. One or more processors 244 and one or more storage devices 250 may provide an operating environment or platform for such modules, which may be implemented as software, but in some examples may include any combination of hardware, firmware, and software. One or more processors 244 may execute instructions, and one or more storage devices 250 may store the instructions and / or data of one or more modules. The combination of the processor 244 and the storage device 250 may retrieve, store, and / or execute the instructions and / or data of one or more applications, modules, or software. The processor 244 and / or the storage device 250 may also be operably coupled to one or more other software and / or hardware components, including but not limited to one or more of the components of the controller 240 and / or one or more devices or systems shown or described as being connected to the controller 240.
[0048] The receiving module 251 may perform functions related to receiving the task 101 from a client or tenant in a multi-tenant environment or from other systems or entities. In some examples, the receiving module 251 receives the task and stores the task in one or more of the data memory 259 or the task list 258.
[0049] The calibration module 252 may perform functions related to determining weights for various types of tasks and / or calibrating parameters used by the controller 240 for task allocation to worker nodes, in Figure 2 the context where the worker nodes include the collector 213. In some examples, the calibration module 252 may use the feedback 102 to occasionally or continuously update a machine learning model to predict appropriate weights for various task types.
[0050] The assignment module 253 may perform functions related to assigning tasks to the collector 213 or other worker nodes within the data center 210. In some examples, the assignment module 253 may use information about task weights to determine how to assign such tasks. Additionally, the assignment module 253 may use information about the extent to which a tenant has processed tasks within the data center to ensure that no tenant starves another tenant of data center resources.
[0051] The data memory 259 of the controller 240 may represent any suitable data structure or storage medium for storing information related to the assignment and / or management of tasks and / or task types. The information stored in the data memory 259 may be searchable and / or categorized such that one or more modules within the controller 240 may provide an input requesting information from the data memory 259 and, in response to the input, receive the information stored within the data memory 259. The data memory 259 may be maintained primarily by the receiving module 251.
[0052] In operation, and in accordance with one or more aspects of the present disclosure, the controller 240 may receive tasks to be executed. For example, in the example that may be described in the context of Figure 2 the communication unit 245 of the controller 240 detects an input and outputs information about the input to the receiving module 251. The receiving module 251 determines that the input corresponds to one or more tasks 101 to be processed by the computing environment (e.g., the data center 210). The receiving module 251 identifies the type of task to be executed for each task. In the case where the data center 210 is a multi-tenant data center, the receiving module 251 may also determine the identity of the tenant for each of the tasks 101 for which the task 101 is to be executed.
[0053] The controller 240 may store the tasks 101 in a data structure. For example, the receiving module 251 may store the tasks in one of the task lists 258 within the storage device 250. Each of the task lists 258 (i.e., task lists 258A to 258C) may store tasks of a particular type. For example, if the controller 240 is assigning tasks related to collecting data from the APM 225, the task list 258A may correspond to a task list involving collecting metrics, the task list 258B may correspond to a task list involving collecting log information, and the task list 258C may correspond to a task list involving collecting trace information. Although only three task lists 258 are shown in the storage device 250 of the controller 240, any number of lists or other suitable data structures may be used. Each of the task lists 258 is shown as being included within the storage device 250, but in some examples, one or more of the task lists 258 may be within the data memory 259.
[0054] Over time, the controller 240 may receive a number of tasks 101 (e.g., tasks 101A through 101N, representing any number of tasks). For each such task received by the controller 240, the receiving module 251 may determine the type and tenant associated with the task and store the task 101 appropriately. In some examples, the tasks 101 received by the communication unit 245 of the controller 240 in Figure 2 may correspond to tasks 101A through 101N received by the controller 140 of Figure 1 from various tenants or customers 11 in the system 100 of Figure 1 .
[0055] The controller 240 may assign tasks to the collectors 213. For example, continuing with the example still being described in the context of Figure 2 , the assignment module 253 accesses a task 101 from one of the task list 258 (or data store 259). In one example, the assignment module 253 accesses task 101A. The assignment module 253 determines that collector 213A has the lowest load among all the collectors 213. Thus, the assignment module 253 assigns task 101A to collector 213A. To do so, the assignment module 253 causes the communication unit 245 to output information about task 101A to collector 213A over the network. Collector 213A receives task 101A and queues task 101A for processing (or processes task 101A immediately). In a similar manner, the assignment module 253 of the controller 240 assigns each of the tasks 101 by selecting the collector 213 with the lowest load and outputting information about each task 101 to the selected collector 213. In some examples, particularly for the initial assignment of tasks 101, the assignment module 253 may assume that each task 101 corresponds to an equal amount of work or has an equal impact on the processing resources of the collector 213 that executes the task 101. As described herein, this amount of work or impact on the processing resources may be referred to as the "weight" of the task 101. Thus, the load on each of the collectors 213 may be considered the sum of the weights of the tasks 101 assigned to that collector 213.
[0056] The controller 240 may receive feedback 102 about the task 101. For example, continuing again with the example being described in the context of Figure 2 , collector 213A finishes processing task 101A. Collector 213A outputs feedback 102A to the controller 240, providing information about the processing required to complete task 101A. In some examples, such as when the task involves collecting information from the APM 225 in Figure 2Among them, such information may include the amount of data collected from one or more APMs 225, the time taken to collect the data, the amount of processing resources required to perform the collection (e.g., average utilization or memory), and other information regarding the processing of task 101A. Similarly, for each of the tasks 101 assigned to one of the collectors 213, the controller 240 may receive an instance of feedback 102 when the task 101 is completed. The calibration module 252 stores information regarding each instance of feedback 102 in the storage device 250, where each instance of feedback 102 corresponds to a different underlying task 101. For example, when the collector 213A completes processing task 101A, the collector 213A sends feedback 102A to the controller 240. Similarly, when the collector 213B completes processing task 101B, the collector 213B sends feedback 102B to the controller 240. And generally, when the collector 213N completes processing task 101N, the collector 213N sends feedback 102N to the controller 240.
[0057] The controller 240 may calibrate based on the feedback 102. The calibration module 252 accesses the instances of feedback 102 in the storage device 250. The calibration module 252 associates each received instance of feedback 102 with one of the tasks 101. Based on the association, the calibration module 252 determines the task type (e.g., metric collection, log data collection, trace data collection) associated with each instance of feedback 102. The calibration module 252 determines the expected amount of processing resources consumed by each task type (e.g., metric collection, log data collection, or transaction data collection). For example, in some cases, certain task types may require significantly more processing resources than other task types. Initially, before sufficient feedback 102 from the instance 113 has been analyzed, the calibration module 252 may assume that each of the tasks 101 will consume approximately equal processing resources (or require the same amount of time to complete). However, after analyzing the feedback 102, the calibration module 252 may determine that certain task types require more processing resources (or more time) than others. Thus, based on the feedback 102, the calibration module 252 recalibrates the expected weights associated with each task type. The calibration module 252 determines the updated weights associated with each task type based on the feedback 102. In this way, the calibration module 252 recalibrates the weights associated with each task type (e.g., metric collection, log data collection, or transaction data collection).
[0058] The controller 240 may assign the tasks 101 based on the calibration. For example, still referring to Figure 2, the controller 240 continues to receive new task flows 101. The receiving module 251 of the controller 240 stores each received task 101 in the data memory 259 (and / or one of the task lists 258). The assignment module 253 assigns each of the tasks 101 to the collector 213 with the least load. However, to perform the assignment, the assignment module 253 uses recalibration information regarding the weights associated with each task type. Thus, some collectors 213 may handle fewer tasks 101 after recalibration, where these tasks have a higher expected weight, while other collectors 213 may handle more tasks 101 with a lower expected weight.
[0059] When additional feedback 102 is received, the controller 240 may continue to recalibrate. For example, referring again to Figure 2 , when the controller 240 continues to assign additional tasks 101 to the collectors 213, the controller 240 receives new instances of the feedback 102 associated with the assigned additional tasks 101. The calibration module 252 recalculates the weights for each task type based on the received additional feedback. Over time, the calibration module 252 may improve its ability to predict the weight of each task 101 and / or the amount of time and / or processing resources required to process each task type. Additionally, in the case where the amount of time and / or processing resources required to process each task type changes over time, if the calibration module 252 continues to recalibrate based on the new feedback 102 received for the additional tasks 101, the calibration module may appropriately adjust the weights associated with each task as the conditions change.
[0060] The controller 240 may scale out the collectors 213. For example, the assignment module 253 may determine that the number of tasks 101 to be processed by the collectors 213 is large, and that the tasks 101 can be processed faster and / or more efficiently if one or more additional collectors 213 are instantiated. Thus, if the assignment module 253 can accurately calibrate and / or predict the weights associated with each new task, the assignment module 253 may also be able to accurately determine how many collectors 213 should be added to meet the increased demand. In such an example, the assignment module 253 identifies the number of additional collectors 213 that should be instantiated, and the assignment module 253 causes the controller 240 to instantiate these new collectors 213. Similarly, the assignment module 253 may be able to accurately determine when the number of collectors 213 exceeds the number required for the current processing of the tasks 101. In such an example, the assignment module 253 identifies the number of collectors 213 that are not needed, and the assignment module 253 causes the controller 240 to deallocate or scale down the number of collectors 213 used in the data center 210.
[0061] Thus, based on the closed-loop mechanism of assigning tasks and receiving feedback on those tasks, the controller 240 can be accurately recalibrated over time and efficiently process tasks as needed. Such a process can operate in a relatively self-driven manner because as the recalibration process indicates that the weights of tasks increase over time, appropriate additional computing resources can be used to process those tasks with higher weights. Similarly, since the recalibration process indicates that the weights of tasks decrease over time, fewer computing resources can be assigned to process those tasks with lower weights.
[0062] The controller 240 can also ensure fairness among multiple tenants. For example, the receiving module 251 can receive tasks 101 from many different tenants in a multi-tenant data center. In some cases, a tenant that generates many tasks 101 (or tasks 101 with higher weights) may consume more resources than other tenants. For example, if a first tenant has more tasks to be executed compared to any one of many other tenants, the tasks of the first tenant can be assigned to the collector 213 more frequently than the tasks of any other tenant. And if the tasks of the first tenant require a large amount of processing, this may affect the availability of processing resources for other tenants. Therefore, when assigning a specific task 101 to the collector 213, the assignment module 253 can consider the tenant associated with that task. In some examples, the assignment module 253 can delay the assignment or assign a lower priority to the tasks 101 associated with tenants that may consume too many resources within the data center 210. Such a policy can mitigate or prevent one or more tenants from starving other tenants of processing resources.
[0063] As described above, the collector 213 maintains a queue of tasks 101 and executes the tasks 101. Once a task is completed, each collector shares information about the task with the controller, which can include the time taken to complete the task. In some examples, the heuristic may be as follows: within a given time interval, the time taken to complete tasks related to metrics, logs, and traces has values "MT", "LT", and "TT", respectively. The controller 240 calculates the weight of a task (for each different type) by considering the relative time taken for each task type (e.g., for metrics, the weight may be = MT / (MT + LT + TT)).
[0064]
[0065] In one example, the load of collector 213 can be represented by the sum of the weights of the individual tasks 101 assigned to that collector 213. The weight of a task may correspond to the weight of the data it is supposed to collect. The controller 240 uses this information to distribute tasks 101 among different collectors 213 by selecting the collector with the lowest load. If the controller 240 needs to obtain metrics, logs, and traces between timestamp A and timestamp B, the controller can divide that time frame into different time intervals (i.e., sliding time windows), referred to as "delta" amounts of time. In some examples, the weights of all tasks for obtaining a particular data type can initially be set to 1 (i.e., equal weights). Then, the controller creates the following tasks and assigns them to collectors on a cyclic basis:
[0066]
[0067]
[0068] The controller can maintain a list of collectors sorted based on load (e.g., a doubly linked list). When a task is assigned to a collector, its load value increases, and the list is rearranged or re-sorted. When a task is completed and the appropriate collector notifies the controller, the load value decreases, and the list is rearranged / resorted. The collector also provides the controller with information about how much data was collected and the time required for retrieval. This information can be used to re-evaluate the weights associated with tasks for the corresponding data type.
[0069] If a collector crashes, the controller receives an event when the control connection between the collector and the controller is disconnected. This enables tasks assigned to the crashed collector to be assigned to other available collectors.
[0070] The controller can determine the number of collectors to increase / decrease. For example, the controller can periodically (e.g., once a day or based on configuration) check whether the collectors need to scale out / in depending on the load trend. An exemplary heuristic formula for determining the number of collectors to increase and / or decrease is outlined below:
[0071] W1 = the weights of all tasks in the queue at time t1
[0072] W2 = the weights of all tasks in the queue at time t2
[0073] Growth rate (RI) = (W2 - W1) / (t2 - t1)
[0074] Completion rate (RC) = (the weights of all tasks completed between t1 and t2) / (t2 - t1)
[0075] Number of collectors = c
[0076] Average efficiency (E) of the collector = RC / c
[0077] If RC is too small compared to RI, then
[0078] Number of additional collectors to be added = RI / E
[0079] Assume that for timestamps t1 and t2, t1 - t2 = 10, and at t1, W1 = 50, and at t2, W2 = 70.
[0080] Therefore, RI = 2
[0081] If the weight of all tasks completed between t1 and t2 = 10, then RC = 1
[0082] Number of collectors = 5
[0083] Average efficiency of the collector = 0.2
[0084] Number of additional collectors to be added = 2 / 0.2 = 10
[0085] The controller can also ensure fairness in a multi-tenant system to ensure that one tenant (with a larger load) does not starve another tenant. To ensure fairness, the controller can apply any one of a variety of different methods, including any one of the two methods described below.
[0086] In the first method, the controller 240 is configured to modify how each task 101 is added to the task list 258. In such an example, when the controller adds a new tenant task to the task list, it linearly scans the list to find other tasks of the same tenant associated with the new task. If the controller finds two consecutive tasks (or a configured number of consecutive tasks) belonging to the same tenant, the controller inserts the task between the two tasks. Otherwise, the controller inserts the new task at the end of the list.
[0087] In the second method, the controller can apply a formula-based algorithm that addresses priority-based allocation (facilitating real-time data) and tenant-based allocation (fairness). In this method, the controller initially assigns an equal constant weight (e.g., "W") to each task. At time t2, the constant W is added to any unfinished task that is equal to or earlier than t1. Thus, since the task remains unfinished, its weight increases over time.
[0088] Then, the controller initializes another value for each tenant, representing "work completed" ("J"). Each time the collector finishes a task, the value J is incremented by a constant value or the number of records processed. In this way, as more tasks are completed for a given tenant, the work completed for that tenant also increases. Additionally, the controller assigns a "priority value" ("P") to each tenant. In some examples, a tenant who pays a premium can be assigned a higher priority value (e.g., a multiple of P, such as 2P, 3P, or higher).
[0089] The controller calculates a final weight, which can be used to determine the order of task assignment. For example, in some examples, the final weight can be calculated as follows:
[0090] Final weight = Current task weight (W) + (Sum of weights of all outstanding tasks for the tenant at time T) * (1 / Work completed) * Tenant priority (P)
[0091] Since the sum of the weights of all outstanding tasks for the tenant is included in the calculation of the final weight, tenants with more work to be done are given a higher priority. Similarly, at any point in time, the priority associated with a tenant can be increased to enable faster processing of tasks. Next, the task with the highest final weight is assigned.
[0092] Figure 2 The modules shown (e.g., the receiving module 251, the calibration module 252, and the assignment module 253) and / or the modules shown or described elsewhere in this disclosure can perform the described operations using software, hardware, firmware, or a combination of hardware, software, and firmware residing in and / or executed at one or more computing devices. For example, a computing device can utilize multiple processors or multiple devices to execute one or more of such modules. A computing device can execute one or more such modules, such as a virtual machine executing on underlying hardware. One or more such modules can be executed as one or more services of an operating system or a computing platform. One or more such modules can be executed as one or more executable programs at the application layer of a computing platform. In other examples, the functionality provided by the modules can be implemented by dedicated hardware devices.
[0093] Although certain modules, data memories, components, programs, executable programs, data items, functional units, and / or other items included within one or more storage devices may be shown separately, one or more of such items may be combined and operate as a single module, component, program, executable program, data item, or functional unit. For example, one or more modules or data memories may be combined or partially combined such that they operate as a single module or provide a function. Additionally, one or more modules may interact with and / or cooperate with each other such that, for example, one module acts as a service or extension of another module. Further, each module, data memory, component, program, executable program, data item, functional unit, or other item shown within the storage device may include multiple components, sub-components, modules, sub-modules, data memories, and / or other components or modules or data memories not shown.
[0094] Additionally, each module, data memory, component, program, executable program, data item, functional unit, or other item shown within the storage device may be implemented in various ways. For example, each module, data memory, component, program, executable program, data item, functional unit, or other item shown within the storage device may be implemented as a downloadable or pre-installed application or “app”. In other examples, each module, data memory, component, program, executable program, data item, functional unit, or other item shown within the storage device may be implemented as part of an operating system executing on a computing device.
[0095] Figure 3 is a flowchart showing operations performed by an exemplary controller in accordance with one or more aspects of the present disclosure. The present disclosure is described herein in the context of Figure 1 controller 140 of Figure 3 . In other examples, Figure 3 the operations described in Figure 3 may be performed by one or more other components, modules, systems, or devices. Additionally, in other examples, the operations described in connection with
[0096] In Figure 3 the process shown, and in accordance with one or more aspects of the present disclosure, controller 140 may receive a task (301). For example, referring to Figure 1 , controller 140 may receive a first set of tasks 101, where each task in the first set has a task type.
[0097] Controller 140 may assign the task based on the task type (302). For example, Figure 1The controller 140 can assign each of the tasks 101 to a compute node or instance 113. In some examples, particularly before performing calibration of important task types, the controller 140 can assume that the expected processing volume associated with each task (e.g., the "weight" of each task) is equal. In such examples, the controller 140 assigns tasks based on task type, where the task types are effectively assumed to be the same.
[0098] The controller 140 can receive feedback (303). For example, the controller 140 receives feedback information from each of the instances 113 that execute one of the tasks 101. The feedback information provides details about the processing performed by the instance 113 in completing the task (e.g., the type of task, the time taken, the processing cycles used, the utilization of the instance 113 or server 112 that executes the task).
[0099] The controller 140 can calibrate the task type weights based on the feedback (304). For example, the controller 140 analyzes the feedback and uses the received feedback (and other feedback) to generate (or update) a model to predict the processing volume required to execute each task type (e.g., the "weight" of the task) based on task type.
[0100] The controller 140 can wait for additional tasks (305), and when additional tasks are received (the "yes" path from 305), the controller 140 can assign tasks based on task type (302). For example, Figure 1 the controller 140 receives a second set of tasks, each with a task type. The controller 140 predicts the weight of each task in the second set based on the feedback 102 received during the processing of the first set of tasks 101 (or based on the feedback 102 received during the processing of earlier tasks). The controller 140 assigns each of the tasks 101 in the second set to an instance 113 based on the predicted weight of each of the tasks 101 in the second set. The assignment can be performed in a manner that maximizes efficient processing. This calibration process based on feedback and the further assignment of tasks after calibration can continue when additional tasks or sets of tasks are received, enabling the controller 140 to further improve its ability to predict the weight of a given task based on the additional feedback 102 received for such additional tasks. Additionally, for a multi-tenant environment, the controller 140 can also assign tasks in a way that ensures that no single tenant consumes resources in a manner that could have an adverse impact on tasks being performed on behalf of other tenants.
[0101] For the processes, apparatuses, and other examples or illustrations described herein that are included in any flow chart or process diagram, certain operations, actions, steps, or events included in any of the techniques described herein may be performed in a different order, may be added, combined, or entirely omitted (e.g., not all described actions or events are necessary for practicing the technique). Additionally, in some examples, operations, actions, steps, or events may be performed in parallel, rather than sequentially, for example, by multithreading, interrupt processing, or multiple processors. Further, certain operations, actions, steps, or events may be performed automatically, even if not specifically identified as being performed automatically. Additionally, certain operations, actions, steps, or events described as being performed automatically may alternatively be performed non-automatically, but in some examples, such operations, actions, steps, or events may be performed in response to an input or another event.
[0102] The disclosures of all publications, patents, and patent applications mentioned herein are incorporated herein by reference. To the extent any material incorporated by reference conflicts with the present disclosure, the present disclosure controls.
[0103] For ease of illustration, a limited number of apparatuses (e.g., controller 140, server 112, collector 213, APM 225, and other apparatuses) are shown in the figures and / or other diagrams referred to herein. However, the techniques in accordance with one or more aspects of the present disclosure may be implemented using more such systems, components, apparatuses, modules, and / or other items, and a collective reference to such systems, components, apparatuses, modules, and / or other items may represent any number of such systems, components, apparatuses, modules, and / or other items.
[0104] The figures included herein each illustrate at least one exemplary implementation of an aspect of the present disclosure. However, the scope of the present disclosure is not limited to such implementations. Thus, other examples or alternative implementations of the systems, methods, or techniques described herein (other than those shown in the figures) may be suitable for other instances. Such implementations may include a subset of the apparatuses and / or components included in the figures and / or may include additional apparatuses and / or components not shown in the figures.
[0105] The detailed description set forth above is intended as a description of various configurations and is not intended to represent the only configurations in which the concepts described herein may be practiced. The detailed description includes specific details for the purpose of providing a thorough understanding of the various concepts. However, the concepts may be practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form in the reference figures to avoid obscuring such concepts.
[0106] Accordingly, although one or more implementations of various systems, devices, and / or components may be described with reference to specific figures, such systems, devices, and / or components may be implemented in many different ways. For example, one or more systems shown separately may alternatively be implemented as a single system; one or more components shown as separate components may alternatively be implemented as a single component. Additionally, in some examples, one or more systems shown as a single system in the figures herein may alternatively be implemented as multiple systems; one or more components shown as a single component may alternatively be implemented as multiple components. Each of such multiple systems and / or components may be directly coupled via wired or wireless communication and / or remotely coupled via one or more networks. Further, one or more systems or components shown in the various figures herein may alternatively be implemented as part of another system or component not shown in such figures. In this and other ways, some of the functions described herein may be performed via distributed processing by two or more systems or components.
[0107] Moreover, certain operations, techniques, features, and / or functions may be described herein as being performed by specific components, systems, and / or modules. In other examples, such operations, techniques, features, and / or functions may be performed by different components, systems, or modules. Thus, some operations, techniques, features, and / or functions that may be described herein as being attributable to one or more components, systems, or modules may in other examples be attributable to other components, systems, and / or modules, even if not specifically described herein in such a manner.
[0108] Although specific advantages have been identified in connection with the description of some examples, various other examples may include some, all, or none of the recited advantages. Other technical or other advantages may become apparent to those of ordinary skill in the art in accordance with this disclosure. Additionally, although specific examples have been disclosed herein, aspects of the disclosure may be implemented using any number of techniques, whether currently known or not, and thus, the disclosure is not limited to the examples specifically described and / or shown herein.
[0109] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored on and / or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates transfer of a computer program from one place to another (e.g., transferred according to a communication protocol). In this manner, the computer-readable medium generally may correspond to: (1) a non-transitory tangible computer-readable storage medium; or (2) a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0110] By way of example, and not limitation, such a computer-readable storage medium may include RAM, ROM, EEPROM, or optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Additionally, any connection may be properly termed a computer-readable medium. For example, if instructions are transmitted using a wired (e.g., coaxial cable, fiber optic cable, twisted pair) or wireless (e.g., infrared, radio, and microwave) connection from a website, server, or other remote source, the wired or wireless connection is included in the definition of the medium. However, it should be understood that the computer-readable storage medium and data storage medium do not include connections, carrier waves, signals, or other transient media, but rather refer to non-transitory tangible storage media.
[0111] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, as used herein, the terms “processor” or “processing circuit” may each refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some examples, the described functionality may be provided within dedicated hardware and / or software modules. Further, the techniques may be fully implemented in one or more circuits or logic elements.
[0112] The techniques of the present disclosure can be implemented in a wide variety of apparatuses or devices, including, to a suitable extent, wireless handsets, mobile or non-mobile computing devices, wearable or non-wearable computing devices, integrated circuits (ICs) or IC sets (e.g., chip sets). Various components, modules, or units are described in the present disclosure to emphasize functional aspects of apparatuses configured to perform the disclosed techniques, but need not necessarily be implemented by distinct hardware units. Instead, as described above, the various units can be combined in a hardware unit or provided by a collection of interoperating hardware units including one or more processors as described above in conjunction with suitable software and / or firmware.
Claims
1. A task assignment method, comprising: receiving, by a controller, a first set of tasks, wherein each task in the first set of tasks has a task type of one of a plurality of task types; assigning, by the controller, each task in the first set of tasks to a worker node for processing by the worker node; receiving, by the controller, feedback information regarding processing performed by the worker nodes for at least some of the first set of tasks; determining, by the controller, an expected throughput associated with each of the plurality of task types based on the feedback information; receiving, by the controller, a second set of tasks, wherein each task in the second set of tasks has a task type from among the plurality of task types; and Each task in the second set of tasks is assigned, by the controller, to the worker node for processing based on the expected processing volume associated with each task type.
2. The method according to claim 1, further comprising: Information about tasks assigned to each of the plurality of worker nodes is stored by the controller.
3. The method according to claim 2, wherein: Assigning each task in the second set of tasks comprises: Each task in the second set of tasks is assigned further based on information about the tasks assigned to each worker node in the plurality of worker nodes.
4. The method according to claim 2, wherein: Storing information about tasks assigned to each of the plurality of worker nodes includes: Information about a type associated with each task assigned to each of the plurality of worker nodes is stored.
5. The method according to claim 4, wherein: Assigning each task in the second set of tasks further comprises: Each task in the second set of tasks is further assigned based on information about a type associated with each task assigned to each worker node in the plurality of worker nodes.
6. The method according to any one of claims 1 to 5, wherein: The worker node is included in a multi-tenant computing environment, wherein the tasks in the first set of tasks are each associated with one of the tenants in the multi-tenant computing system, and wherein assigning each task in the second set of tasks further comprises: Each task in the second set of tasks is further assigned based on the information of the tenant associated with each task assigned to each of the plurality of worker nodes to ensure access to the worker nodes by each tenant in the multi-tenant computing environment.
7. The method according to any one of claims 1 to 5, further comprising: Determining, by the controller based on the feedback information and the second set of tasks, that instantiating an additional worker node will be able to more efficiently process the second set of tasks; as well as The additional working node is instantiated by the controller.
8. The method according to any one of claims 1 to 5, wherein: Each of the worker nodes is executed within a computing node, and wherein the method further comprises: Determining, by the controller based on the feedback information and the second group of tasks, a process that can efficiently execute the second group of tasks with fewer working nodes; and One of the working nodes is de-assigned by the controller.
9. The method according to any one of claims 1 to 5, wherein: Determining the expected processing volume associated with each task type includes: Determine the weight associated with each task type.
10. The method according to claim 1, wherein: Receiving the first set of tasks includes: A set of tasks associated with data collection is received from a plurality of application performance monitoring systems.
11. A task assignment computing system, comprising a processing circuit and a storage device, wherein: The processing circuit has access rights to the storage device and is configured to: receiving a first set of tasks, wherein each task in the first set of tasks has a task type of one of a plurality of task types; Assigning each task in the first set of tasks to a worker node for processing by the worker node; receiving feedback information regarding processing performed by the worker nodes for at least some of the first set of tasks; determining an expected throughput associated with each of the plurality of task types based on the feedback information; receiving a second set of tasks, wherein each task in the second set of tasks has a task type from among the plurality of task types; and Each task in the second set of tasks is assigned to the worker node for processing based on the expected processing volume associated with each task type.
12. The computing system of claim 11, wherein: The processing circuit is further configured to: Information about tasks assigned to each of the plurality of worker nodes is stored.
13. The computing system of claim 12, wherein: To assign each task in the second set of tasks, the processing circuit is further configured to: Each task in the second set of tasks is assigned further based on information about the tasks assigned to each worker node in the plurality of worker nodes.
14. The computing system of claim 12, wherein: In order to store information about tasks assigned to each of the plurality of worker nodes, the processing circuit is further configured to: Information about a type associated with each task assigned to each of the plurality of worker nodes is stored.
15. The computing system of claim 14, wherein: To assign each task in the second set of tasks, the processing circuit is further configured to: Each task in the second set of tasks is further assigned based on information about a type associated with each task assigned to each worker node in the plurality of worker nodes.
16. The computing system of any one of claims 11 to 15, wherein: The worker node is included in a multi-tenant computing environment, wherein the tasks in the first set of tasks are each associated with one of the tenants in the multi-tenant computing system, and wherein, to assign each task in the second set of tasks, the processing circuit is further configured to: Each task in the second set of tasks is further assigned based on information about the tenant associated with each task assigned to each worker node in the plurality of worker nodes.
17. The computing system of any one of claims 11 to 15, wherein: The processing circuit is further configured to: Determining, based on the feedback information and the second set of tasks, that instantiating an additional worker node will be able to more efficiently process the second set of tasks; as well as The additional working node is instantiated.
18. The computing system of any one of claims 11 to 15, wherein: The processing circuit is further configured to: Determining, based on the feedback information and the second set of tasks, a process that can efficiently execute the second set of tasks using fewer worker nodes; as well as One of the worker nodes executing within a compute node is de-dispatched.
19. The computing system of any one of claims 11 to 15, wherein: Determining an expected amount of processing associated with each task type, the processing circuitry is further configured to: Determine the weight associated with each task type.
20. A computer-readable storage medium encoded with instructions for causing one or more programmable processors to be configured to perform the method according to any one of claims 1 to 10 or to be configured as a tasking computing system according to any one of claims 11 to 19.