Orchestration of Containerized Applications
By optimizing the utilization of computing resources by orchestration engine, the problem of difficult real-time computing services in the prior art is solved, efficient utilization of resources and reduction of expensive hardware are achieved, and suitable for environments under remote or adverse conditions.
Patent Information
- Application Number
- CN201980079235.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-10-02
- Filing Date
- 2019-08-30
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2039-08-30
AI Technical Summary
The prior art is difficult to efficiently utilize computing resources when providing real-time computing services, and has high demands for expensive high-performance local computing hardware and maintenance, especially in remote or adverse environments.
The orchestration engine receives task requests from multiple process engines, selects computing instances with available computing resources, and determines scheduling and allocation plans based on the predicted runtime, task requirements and available resources to optimize the execution of computing tasks.
It enables efficient utilization of computing resources while meeting real-time task requirements, reduces dependence on expensive and high-performance local hardware, and simplifies maintenance in remote or difficult-to-access locations.
Smart Images

Figure CN113168347B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of U.S. Provisional Application Serial No. 62 / 740,034, filed Oct. 2, 2018, the disclosure of which is hereby incorporated by reference in its entirety. Technical Field
[0003] This application relates to computing, such as utility computing. More specifically, this application relates to systems and methods for orchestrating computing tasks on a computing platform. Background Art
[0004] Containerization can be used to deploy control software located at remote or hard - to - access locations. However, when such deployments are time - sensitive, such software is typically provided by on - site technical support. For example, a professional technician can be dispatched to a remote or restricted - access location, such as an offshore oil platform, a desert solar farm, etc. It has been recognized herein that providing on - site support at such locations is not only expensive, but it may also prevent immediate response to other critical events, such as events that can stop production or cause irreparable damage to infrastructure. In some cases, high - performance computing hardware can be implemented at a remote location instead of dispatching an expert to the remote location. However, it has been further recognized herein that, among other drawbacks, such hardware can be expensive and difficult to maintain. Additionally, maintenance can be particularly difficult in remote areas subject to adverse conditions.
[0005] In utility computing applications, a container can include a lightweight executable software package. A container can include everything needed to run the software. Software containers can enable streamlined application portability across different environments. Thus, among other things, containerizing an application is a common practice for exporting a service from a host system to a computing platform. Recent advances in container technology allow for the development of containers based on a real - time operating system (i.e., real - time containers). One of the challenges of real - time containers is to schedule container execution with real - time guarantees while using minimal computing resources. Summary of the Invention
[0006] Embodiments of the present invention solve and overcome one or more of the disadvantages described herein by providing a method, system, and apparatus for orchestrating containerized applications associated with real-time requirements. Aspects of the present invention include methods and systems for performing multiple computing tasks. An orchestration engine may receive task requests from multiple process engines over a network. The process engines may correspond to respective remote edge and / or field devices as compared to the orchestration engine. Each task request may indicate at least one task requirement for performing a corresponding computing task. A plurality of computing instances having available computing resources may be selected from a set of computing instances. A predicted runtime may be generated for each of the computing tasks. In one example, based on the predicted runtime, task requirements, and available computing resources, the orchestration engine determines a scheduling and allocation scheme. The scheduling and allocation scheme defines when each of the multiple computing tasks is to be executed and which of the multiple selected computing instances executes each of the multiple computing tasks. The selected computing instances execute the multiple computing tasks according to the scheduling and allocation scheme. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The foregoing and other aspects of the present invention are best understood from the following detailed description when read in conjunction with the accompanying drawings. For purposes of illustrating the present invention, there are shown in the drawings presently preferred embodiments, it being understood, however, that the invention is not limited to the specific means disclosed. The following drawings are included:
[0008] Figure 1 is a block diagram of an exemplary architecture including an orchestration engine according to an embodiment of the present disclosure.
[0009] Figure 2 is a process diagram depicting an exemplary process performed by an orchestration engine and other nodes within an exemplary computing platform according to an embodiment of the present disclosure.
[0010] Figure 3 is a block diagram of an exemplary orchestration engine according to an embodiment of the present disclosure.
[0011] Figure 4 is a flowchart of an exemplary process for performing computing tasks according to an embodiment of the present disclosure.
[0012] Figure 5 shows an example of a computing environment in which embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION
[0013] Methods and systems are disclosed for orchestrating containerization of application programs from edge and / or field devices with real-time requirements to a computing platform. It is recognized herein that current methods for containerization solutions for real-time applications lack capabilities and efficiency. For example, real-time performance is often not guaranteed because each application program is granted a dedicated computing instance for its execution, resulting in real-time behavior, but there is no proper mechanism to enforce it. It is further recognized herein that having a dedicated computing instance for each application program is inefficient and cannot scale to a large number of application programs. Furthermore, network transmission latency is not considered in current methods.
[0014] In an example aspect, an orchestration engine receives task requests from multiple process engines over a network. The process engines may correspond to respective edge and / or field devices located remotely compared to the orchestration engine or compared to the computing platform that executes the tasks. Each task request may indicate at least one task requirement for performing a corresponding computing task. For example, some computing tasks may require real-time execution. In some cases, real-time execution of a computing task may mean that data from the task can be used as feedback within a specific time requirement. Some computing tasks are executed with real-time output data that can be used as feedback almost immediately (e.g., within a few milliseconds). In the example, based on the predicted runtime of the tasks, the task requirements, and the available computing resources, the orchestration engine determines a scheduling and allocation scheme. The scheduling and allocation scheme defines when each of the multiple computing tasks is to be executed and where (i.e., which computing instance) each of the computing tasks is to be executed.
[0015] The disclosed methods and systems improve the functionality of a computer for performing such computer-based tasks. Additionally, the disclosed methods and systems can reduce the need for expensive high-performance local computing hardware by extracting real-time control software located in remote or difficult-to-access locations. Such local hardware may be difficult to maintain in practice. In various examples, the real-time software can be extracted to a computing platform where a large amount of computing resources are available and technical support can be provided instantaneously. In some cases, according to various exemplary embodiments, even though 5G access and core networks can provide low latency and high communication reliability, network impacts are considered when containerizing time-critical application programs for computing on an external platform.
[0016] First, refer to Figure 1, which shows an exemplary architecture or system 100 for performing computing tasks according to an embodiment of the present disclosure. System 100 includes a computing platform 102, which includes an orchestration engine 104 and a plurality of computing instances 106 that communicate with the orchestration engine 104. In particular, the orchestration engine 104 can instruct the computing instances 106 to execute corresponding tasks at a specific time, thereby saving the computing resources of the computing instances 106 while meeting task requirements (such as real-time task requirements). The computing instances 106 can include virtual central processing units (CPUs), virtual servers, edge bare-metal resources, etc. In some examples, the computing instances 106 can vary based on the computing platform in which they are implemented. The computing instances 106 can have the same or different processing capabilities relative to each other. The computing platform 102 can be implemented as a cloud computing platform such that the orchestration engine 104 and the computing instances 106 run in the cloud. Additionally, each of the computing instances 106 and the orchestration engine 104 can run on separate, different resources. Still alternatively, the computing instances 106 and the orchestration engine 104 can run on a combination of the cloud and separate resources.
[0017] It should be understood that for illustrative purposes, the exemplary system 100 is simplified, and other systems can be used to perform the computing task orchestration herein, and all such systems are considered to be within the scope of the present disclosure. For example, the computing platform 102 can include any number of computing instances 106 as needed. In addition, the orchestration engine 104 and the computing instances 106 can be hosted on any number of devices at any number of physical locations. In an example, the orchestration engine 104 is hosted on a single device that is located at the same location as the computing instances 106, although it should be understood that the embodiments herein are not limited thereto.
[0018] The exemplary system 100 also includes a plurality of tasks 108. The tasks 108 can be sent to the orchestration engine 104 in a container. As used herein, unless otherwise specified, a container refers to an executable package of software. For example, the orchestration engine 104 can receive a real-time task or a containerized real-time application in a container. Additionally or alternatively, the tasks 108 can be sent to the orchestration engine 104 as data. For example, the tasks 108 can be sent to the orchestration engine 104 as input data for software that is already available on the computing platform 102. Thus, the tasks 108 can be provided to the orchestration engine 106 as containers and / or data.
[0019] Referring again to Figure 2 , which shows an exemplary process that can be executed within the computing platform 102. Figure 2An exemplary system 200 is shown that includes a computing platform 102 in communication with a plurality of process engines 110 via a network 112. In some cases, the process engines 110 correspond to respective machines or edge or field devices that are remotely located compared to the orchestration engine 104 and thus the computing platform 102. The process engines 110 may send tasks 108 to the orchestration engine 104. For example, the tasks 108 may be associated with the operation of the machines or devices corresponding to the process engines 110 such that the computing platform 102 may provide real-time control of the machines or devices. The computing platform 102 may be remote compared to the process engines 110. Thus, the computing platform 102 may be remote compared to the devices and machines corresponding to the process engines 110. In such a configuration, the computing platform 102 may remotely provide services for real-time operation. For example, the process engines 110 may correspond to the operation of gas turbines on a remote oil platform such that the computing platform 102 may control the gas turbines from a remote location compared to the gas turbines. Additionally, the computing platform 102 may be co-located with the process engines 110.
[0020] It should be appreciated that the orchestration engine 104 and thus the computing platform 102 may be configured to perform the operations of any of the process engines 110 as needed. Additionally, the process engines 110 may be co-located relative to each other to serve common operations, or the process engines 110 may be distributed at different locations to serve operations independent of each other. The process engines 110 and the orchestration engine 104 may communicate with each other via the network 112, which may be any network or system commonly known in the art, including the Internet, intranet, local area network (LAN), wide area network (WAN), metropolitan area network (MAN), direct connection or series of connections, cellular telephone network, or any other network or medium capable of facilitating communication between the process engines 110 and the computing platform 102, particularly the orchestration engine 104. The network 112 may be wired, wireless, or a combination thereof. Additionally, several networks may work alone or communicate with each other to facilitate communication in the network 112.
[0021] Continuing reference Figure 2, at 201, a task request is sent by the process engine 110 and received by the orchestration engine 104. The task request may correspond to multiple computing tasks 108. The task request may be received from the process engine 110 via the network 112, and each task request may indicate at least one task requirement for performing the corresponding computing task 108. At 202, the orchestration engine 104 processes the task request. In particular, the orchestration engine 104 may generate a priority for completing each task request. The tasks 108 may be completed in the order defined by the corresponding priorities. In some cases, tasks with higher priorities are executed before tasks with lower priorities compared to other tasks. In various examples, the priorities are determined and generated based on the task requirements. For example, in some cases, each container including one or more tasks 108 is associated with a list of task requirements. Additionally or alternatively, each task within the container may be associated with its own task requirements. By way of example, if there is a task requirement to complete a given computing task in real time, the orchestration engine 104 assigns a priority to that task that is higher than the priority of a different task that is required to be completed within, for example, one hour or a certain time greater than immediately.
[0022] In some examples, the task request or the task 108 may indicate a specific deadline for completing the corresponding task or task request. Thus, the task requirements may include a specific deadline and priorities may be assigned based on the deadline. Additionally or alternatively, exemplary task requirements may indicate the task precedence for completing multiple tasks. For example, it may be necessary to execute a first task before a second task can be executed, so the first task may indicate a task precedence over the second task. Continuing with this example, a higher priority may be assigned to the first task than to the second task. Further, task priorities may change over time. For example, the priority of a given task may increase as its deadline approaches. Referring again to the task precedence example, in some cases, the second task is not scheduled until the first task with precedence has been scheduled. Thus, the priority of the second task may depend on the deadline of the first task, and thus the priority of the second task may increase as the deadline of the first task approaches. In addition to task deadlines, task precedence, etc., the criteria for determining task priorities may also include network conditions associated with processing the task and sending / receiving task-related input / output, as further described below.
[0023] The orchestration engine 104 may monitor the compute instances 106. The orchestration engine 104 may monitor the compute instances 106 continuously or periodically in order to continuously or periodically determine the state of the compute resources of the compute instances 106. The orchestration engine 104 may monitor the utilization of each compute instance 106. Compared to other compute instances 106, a given compute instance 106 may have different virtual cores with different processing capabilities. By monitoring the utilization of each compute instance 106, the orchestration engine 104 may identify a quantifiable amount of the compute resources of the compute instances 106 available at any given time. Thus, in one example, after receiving a task request over the network 112, the orchestration engine 104 selects a compute instance 106 from a group of compute instances, where the selected compute instance 106 has available compute resources.
[0024] The orchestration engine 104 may also monitor the network 112 continuously or periodically. In one example, the orchestration engine 104 monitors the network impact (such as latency or jitter) of the tasks 108 (e.g., each task 108). Additionally, the orchestration engine 104 may identify the corresponding destination associated with each completed compute task. For example, a task request may indicate from which process engine 110 the associated task was initiated, which may correspond to the destination of the task output. Alternatively or additionally, the task request may indicate to which process engine 110 the associated task output should be sent. At 202, the orchestration engine 104 may also generate a task priority based on the network impact associated with the task. Thus, the orchestration engine 104 may generate a priority order for completing a task request based on the task requirements and performance of the network 112. By way of example, if a first exemplary task and a second exemplary task have the same task requirements, but the destination of the output of the second task lies in the process engine 110 associated with network latency that the orchestration engine 104 has observed, then the orchestration engine 104 may assign a higher priority to the second task than to the first task based on the network impact. Thus, in some cases, due to the network state associated with a given task, even if the tasks may have the same task requirements, the orchestration engine 104 may assign a higher (or lower) priority to a given task than to another task. For various reasons, the network states associated with different tasks may vary. For example but not limited to, the network connection for a given task may exhibit latency or jitter, the process engine for a given task may be geographically further away from the process engine of another task, or the network connection associated with the processing of a specific task may be congested compared to the network connection associated with another task. Thus, compared to the same task not associated with network latency, the orchestration engine 104 may assign a higher priority to a task associated with network latency (for various reasons).
[0025] Still referring to Figure 2, at 204, the orchestration engine 104 can predict the running time of a given task. That is, in some cases, the orchestration engine 104 can predict how long it will take for a given task to complete. In some examples, at 203, the orchestration engine 104 monitors the task 108 to obtain metrics associated with past task performance. Thus, the orchestration engine 104 can obtain historical performance data associated with the computing task and / or the computing instance 106. The orchestration engine 104 can apply machine learning to model the running times of various tasks based on the past performance of the tasks. Thus, when a task request is received, at 204, the orchestration engine 104 can predict its running time. In particular, the orchestration engine 104 can generate a predicted running time for the computing task 108 based on the historical performance data associated with the computing task 108 or the computing instance 106. Additionally or alternatively, the orchestration engine 104 can collect performance data associated with the network 112 and can generate a predicted running time for the computing task 108 based on the performance data associated with the network 112.
[0026] At 204, the predicted running time of a given task can also be compared with a predetermined threshold associated with that task. In an example, when the predicted running time is greater than the predefined threshold, an alert indication is triggered at 205. The predefined threshold can represent a critical duration during which the associated task is expected to complete. In an example, if a given task is not executed within its critical duration, it may result in significant consequences, such as a delay or shutdown of operations in an industrial environment. Thus, to avoid or mitigate such consequences, various actions can be taken in response to the triggering of the alert indication when the predicted running time is greater than its corresponding predetermined threshold. For example, when the predicted running time is greater than its corresponding predetermined threshold, at 207, the orchestration engine 104 can determine the process engine 110 associated with the predicted running time, and the orchestration engine 104 can send the alert to the determined process engine 110. By doing so, the process engine 110 can take actions to mitigate the task running time predicted to be greater than its threshold. For example, the task request can indicate the process engine 110 that initiated the task request.
[0027] In another example, when at least one of the predicted running times is greater than its predetermined threshold, the orchestration engine 104 can identify additional computing instances 106 with available computing resources. In this case, the orchestration engine 104 can speed up tasks with a predicted running time that is too long, for example, by identifying other computing resources and using those other computing resources to execute the sped-up tasks. Additionally or additionally, the orchestration engine 104 can adjust (e.g., increase) the priority of a given task with a predicted running time greater than the predefined duration. In this case, the orchestration engine 104 can speed up the task according to the corresponding predicted running time of the task.
[0028] Continue to refer toFigure 2 , the orchestration engine 104 determines a scheduling and allocation scheme for the tasks 108, e.g., the execution of each task 108 received by the orchestration engine 104. The scheduling and allocation scheme may define when each computational task 108 is to be executed and which of the multiple selected computing instances 110 is to execute each of the multiple computational tasks 110. The orchestration engine 104 may determine the scheduling and allocation scheme for the tasks 108 based on the corresponding predicted runtimes, task requirements, and available computing resources. In particular, at 206, the orchestration engine 104 generates an optimization problem, and at 208, solves the optimization problem in order to generate the scheduling and allocation scheme. The optimization problem may be generated based on the priority order of the task requests and the available computing resources.
[0029] In some cases, the schedule is updated based on events. For example, an event may trigger the orchestration engine 104 to generate a new or updated scheduling and allocation scheme. Exemplary events include receiving a new task, the execution of a task being completed, etc. A subset of the events may also trigger an update to the scheduling and allocation scheme. Exemplary subsets include, but are not limited to, violating a task deadline or identifying a problem with a computing instance 106. Additionally or alternatively, the orchestration engine 104 may generate the schedule periodically.
[0030] In the example, after generating the predicted runtime at 204, it may define the input to the optimization problem generated at 206 at 209. In some cases, when the predicted runtime for a given task is less than a predetermined threshold, the input is triggered at 209. The input may include the predicted runtime for a specific task or group of tasks. As described above, the predetermined threshold may represent a critical duration during which the associated task or group of tasks is expected to be completed. Alternatively or additionally, the predetermined threshold may represent a time of day at which the associated task or group of tasks needs to be completed. The predicted runtime may indicate a range of durations during which it is predicted that the associated task or group of tasks will be completed. Alternatively or additionally, the predicted runtime may indicate a specific duration for which it is predicted that the associated task or group of tasks will be completed. The predicted runtime may also or alternatively indicate a time of day or a range of times of day at which it is predicted that the associated task or group of tasks will be completed.
[0031] Various task information and resource information can also limit the input to the optimization problem generated at 206 at 211. For example, after processing tasks at 202 to determine their respective priorities, those task priorities can be input into the optimization problem. More generally, task 108 can indicate its corresponding task requirements, and those task requirements can be input into the optimization problem generated at 206 at 211. In addition, as described above, the orchestration engine 104 can monitor the computing instances 106 to select computing instances with available computing resources. At 211, the available computing resources can be input into the optimization problem. Thus, in the example, the optimization includes predicted runtimes, available computing resources, task requirements, and / or task priorities as inputs. In some cases, the optimization can also include the state of network 112 or the predicted effect of network 112 as an input. In one example, an optimization is generated each time the schedule is updated. Regarding available computing resources, the optimization specifically considers the current load of the available computing instances. As described further below, the output of the optimization can be processed at 202 to produce an allocation and scheduling scheme.
[0032] After generating the optimization problem at 206, the problem can be provided to a solver at 213. At 208, the optimization problem is solved to generate an output. The output generated at 208 can include a scheduling and allocation scheme. In the example, at 208, discrete stochastic optimization is performed to generate a scheduling and allocation scheme. Thus, discrete stochastic optimization can be performed using predicted runtimes, task requirements, and available computing resources as inputs to the discrete stochastic optimization. In one example, discrete stochastic optimization is performed when the optimization engine 104 receives each task 108. That is, at 208, one or more task requests (e.g., a predetermined number of task requests) received by the optimization engine 104 can trigger the optimization, especially an iteration of the optimization. Alternatively or additionally, the optimization engine 104 can perform the optimization periodically to generate updated scheduling and allocation schemes. Thus, it will be understood that the scheduling and allocation schemes can be generated at various times, or in response to various triggers as needed.
[0033] Still referring to Figure 2 , as a result of the optimization performed at 206 and 208, the scheduling and allocation scheme can be output at 215. At 202, the orchestration engine 104 can process the received tasks 108 according to the scheduling and allocation scheme. In particular, at 217, the orchestration engine 104 can send instructions to the computing instances 106 selected to execute the computing tasks 108. The instructions can identify one or more computing instances 106 selected to execute the corresponding computing tasks or task groups according to the allocation scheme. The instructions can also indicate the time or order associated with the time at which each task should be executed according to the schedule generated at 208. In response to the instructions at 217, multiple computing tasks 108 are executed by multiple selected computing instances 106.
[0034] The orchestration engine 104 can monitor the execution of tasks by the computing instance 106. In some cases, if the execution of a particular task encounters a delay, such as causing the task to miss its deadline, the orchestration engine 104 detects the delay in task execution. In response to detecting the delay, the orchestration engine 104 can increase or raise the priority associated with the delayed task when it next receives the delayed task at 201. Similarly, the orchestration engine 104 can monitor the network 112 and can detect whether the network connection associated with a particular task or group of tasks has deteriorated. In response to detecting that the network connection associated with a specific task or group of tasks has deteriorated, the orchestration engine 104 can increase or raise the priority of the current and / or future tasks associated with the deteriorated network connection. In this case, the orchestration engine can ensure that tasks are completed according to their respective requirements (e.g., real-time requirements). Also, if the degraded network connection is repaired or otherwise returns to its expected operational capacity, the orchestration engine 104 can decrease or lower the priority of the current and / or future tasks associated with the repaired network connection.
[0035] The computing task 106 can execute the computing task 106 according to the scheduling and allocation scheme generated by the orchestration engine 104, thereby generating a result. In some examples, at 219, the result of the executed task is returned to the orchestration engine 104. The result can include data and / or operation instructions for the process engine 110. At 207, the orchestration engine 104 can send the result of the executed task to the corresponding process engine 110 that initiated the corresponding task request via the network 112. Alternatively or additionally, the computing instance 106 can directly send the result of the executed task to the corresponding process engine that initiated the corresponding task request.
[0036] Now referring to Figure 3 , the orchestration engine 104 can include one or more processors 311 and a memory 321, in which application programs, agents, and computer program modules for implementing the embodiments of the present disclosure are stored, including a data module 322, an orchestrator module 331, and an artificial intelligence (AI) module 341. The orchestrator module 331 can include an optimization module 332 and a solver module 333. A module can refer to a software component that executes one or more functions. Each module can be a discrete unit, or the functions of multiple modules can be combined into one or more units that form parts of a large program. In Figure 3 the depicted example, the data module 322, the AI module 323, and the orchestrator module 331 are organized to form a program for orchestrating the execution of computing tasks.
[0037] The data module 322 can analyze the data used in the orchestration of executing computational tasks. For example, the data module 322 can analyze task 108 or a task request to determine the requirements of the task and / or assign priorities to the task. The data module 322 can also analyze the alert indications associated with a task or a task group, and thus adjust the task priorities based on the alert indications. The generation of a given scheduling and allocation scheme can be triggered at the data module 322. In particular, in some cases, the task or container orchestration is triggered and iterated at the data module 322. The data module 322 can be used as a backbone to communicate with different analysis components (such as the AI module 323 and the orchestrator 331). The data module 322 can also communicate with nodes external to the orchestration engine 104, such as the computing instance 106 or the process engine 110. For example, the data module 322 can send instructions to the computing instance 106 based on the scheduling and allocation scheme. The data module can also receive task requests from the process engine 110 or send task results to the process engine. For example, in some embodiments, the data module 322 includes one or more application programming interfaces (APIs) that allow the analysis components to access and update the analysis data. In a network-based architecture, the Representational State Transfer (REST) design can be used to access and manipulate the analysis data at the data module 322.
[0038] The AI module 323 can monitor the network to which the orchestration engine 104 is connected. In particular, the AI module can monitor the network connection between the orchestration engine 104 and the corresponding process engine 110. The AI module 323 can also monitor the computing resources available for executing computational tasks, such as the computing instance 106. Through monitoring, the AI module 323 can obtain historical data related to the running time and / or network latency of a task or the impact on various task executions. The AI module 323 can use the historical data to make various predictions. For example, the AI module 323 can apply statistics and machine learning to the data to predict the running time of a given task or task group. Additionally or alternatively, the AI module 323 can apply statistics and machine learning to the data to predict the network impact (i.e., latency) of a given task or task group. In some cases, for example, when the predicted running time is greater than a predetermined threshold, the AI module 323 can generate an alert indication. The predictions or output data of the AI module 323 can be provided to the orchestrator module 331 and / or the data module 322. In an example, the predicted running time and / or the predicted network impact are provided to the orchestrator module 331 so that they can define the input for optimization. In some cases, the alert indication can be provided to the data module 322 so that the task can be accelerated.
[0039] The orchestrator module 331 can receive various data or predictions from the data module 322 and the AI module 323. In particular, the optimization module 332 can receive data associated with task priorities, task requirements, network status, available computing resources, predicted runtimes of tasks, and / or predicted network impacts. The optimization module 332 can generate an optimization problem based on the input to the orchestrator 331. The solver module 333, which can include one or more solvers, can receive the optimization problem as input and can solve the given optimization problem to generate a scheduling and allocation scheme as output. The scheduling and allocation scheme can be obtained by the data module 322 such that tasks can be executed according to the scheduling and allocation scheme.
[0040] Figure 4 A flowchart of an exemplary process 400 for performing multiple computing tasks is shown. At 402, multiple computing resources are monitored. For example, the orchestration engine 104 can monitor various computing instances 106 having various computing resources. By monitoring the computing instances 106, the orchestration engine 202 can obtain historical performance data associated with the computing tasks and / or the computing instances 106. At 404, one or more network connections are monitored. For example, the orchestration engine 404 can monitor the network connections between the orchestration engine 104 and multiple process engines 110. By monitoring the network connections, the orchestration engine 104 can collect performance data associated with the network, and / or the orchestration engine 104 can identify any problems or delays associated with the network connections to the process engines 110.
[0041] Continuing to refer Figure 4 , at 406, the orchestration engine 104 can receive one or more task requests. The task requests can be received as containers that indicate the computing tasks and the task requirements for performing the computing tasks. The task requests can be received from a process engine or other device that is remote from the orchestration engine. By way of example, the requirements can indicate the time or duration by which the corresponding task needs to be completed. At 408, the orchestration engine can assign or determine priorities for each of the tasks or task groups. Higher-priority tasks can be executed before lower-priority tasks. At 410, based on the monitoring performed at 402, the orchestration engine 104 can select computing instances having available computing resources. At 412, the orchestration engine 104 can predict the runtime associated with the tasks received at 406 according to an embodiment of the present disclosure. The prediction at 412 can be based on the available computing resources and the task priorities and, in turn, on the task requirements according to which the task priorities are generated. The prediction can also be based on data obtained from the monitoring at 402 and / or data obtained from the monitoring at 404.
[0042] Still referring Figure 4, according to the example shown, at 414, the orchestration engine 104 determines whether the predicted runtime is greater than a predetermined threshold. In particular, for example, the orchestration engine 104 can determine whether the execution of a predicted task or task group takes too much time. If it is determined that the predicted runtime of the task or task group is greater than the predetermined threshold, the process can proceed to 416, where the task with the too-long predicted runtime is accelerated, as described herein. For example, an alert indication can be generated, and the orchestration engine can identify other available computing resources for executing the task (and the associated trade-off costs for utilizing the other available computing resources) to accelerate the task. As other examples, at 416, the priority of the task can be increased to accelerate the task with a predicted runtime greater than the corresponding predetermined threshold. If it is determined that the predicted runtime of the task or task group is less than the predetermined threshold, the process can proceed to 418, where a scheduling and allocation scheme is generated. The scheduling and allocation scheme can be determined based on the predicted runtime from 412 and the task requirements that can be received at 406. The scheduling and allocation scheme can also be or alternatively be based on the priorities assigned at 408. Further, the scheduling and allocation scheme can be based on the available computing resources of the computing instances selected at 410 and / or the network data obtained at 404. The scheduling and allocation scheme can define when each of the computational tasks or task groups is to be executed. The scheduling and allocation scheme can also define which of the selected computing instances executes each of the multiple computational tasks. At 420, according to the example shown, the selected computing instances execute the computational tasks according to the scheduling and allocation scheme.
[0043] Figure 5 An example of a computing environment in which embodiments of the present disclosure can be implemented is shown. The computing environment 500 includes a computer system 510, which can include a communication mechanism such as a system bus 521 or other communication mechanisms for transferring information within the computer system 510. The computer system 510 also includes one or more processors 520 coupled to the system bus 521 for processing information.
[0044] The processor 520 may include one or more central processing units (CPUs), graphics processing units (GPUs), or any other processors known in the art. More generally, a processor herein is a device for executing machine-readable instructions stored on a computer-readable medium to perform tasks, and may include either or a combination of hardware and firmware. The processor may also include a memory that stores machine-readable instructions executable for performing tasks. The processor operates on information by manipulating, analyzing, modifying, transforming, or transmitting it for use by an executable procedure or information device, and / or by routing the information to an output device. For example, the processor may use or include the capabilities of a computer, a controller, or a microprocessor, and may use executable instructions to effectuate regulation to perform special functions not performed by a general-purpose computer. The processor may include any type of suitable processing unit, including but not limited to a central processing unit, a microprocessor, a reduced instruction set computer (RISC) microprocessor, a complex instruction set computer (CISC) microprocessor, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a system on a chip (SoC), a digital signal processor (DSP), etc. Additionally, the processor 520 may have any suitable microarchitecture design, including any number of constituent components such as registers, multiplexers, arithmetic logic units, cache controllers for controlling read / write operations to cache memories, branch predictors, etc. The microarchitecture design of the processor may be capable of supporting any one of a plurality of instruction sets. The processor may be coupled (electrically and / or including executable components) to any other processor, enabling interaction and / or communication therebetween. A user interface processor or generator is a known element that includes an electronic circuit or software or a combination of both for generating a display image or a portion thereof. The user interface includes one or more display images that enable a user to interact with the processor or other devices.
[0045] The system bus 521 may include at least one of a system bus, a memory bus, an address bus, or a message bus, and may permit the exchange of information (e.g., data (including computer-executable code), signaling, etc.) among the various components of the computer system 510. The system bus 521 may include but is not limited to a memory bus or memory controller, a peripheral bus, an accelerated graphics port, etc. The system bus 521 may be associated with any suitable bus architecture, including but not limited to, an Industry Standard Architecture (ISA), a Micro Channel Architecture (MCA), an Enhanced ISA (EISA), a Video Electronics Standards Association (VESA) architecture, an Accelerated Graphics Port (AGP) architecture, a Peripheral Component Interconnect (PCI) architecture, a PCI-Express architecture, a Personal Computer Memory Card International Association (PCMCIA) architecture, a Universal Serial Bus (USB) architecture, etc.
[0046] Continue to refer toFigure 5 The computer system 510 may also include a system memory 530 coupled to the system bus 521 for storing information and instructions to be executed by the processor 520. The system memory 530 may include computer-readable storage media in the form of volatile and / or non-volatile memory, such as read-only memory (ROM) 531 and / or random access memory (RAM) 532. The RAM 532 may include other dynamic storage devices (e.g., dynamic RAM, static RAM, and synchronous DRAM). The ROM 531 may include other static storage devices (e.g., programmable ROM, erasable PROM, and electrically erasable PROM). In addition, the system memory 530 may be used to store temporary variables or other intermediate information during the execution of instructions by the processor 520. The basic input / output system 533 (BIOS) may be stored in the ROM 531, which contains basic routines that help transfer information between elements within the computer system 510 (such as during startup). The RAM 532 may contain data and / or program modules that can be immediately accessed by the processor 520 and / or are currently operating on the processor. The system memory 530 may additionally include, for example, an operating system 534, application programs 535, and other program modules 536. The application programs 535 may also include a user portal for developing application programs, allowing input and modification of input parameters as needed.
[0047] The operating system 534 may be loaded into the memory 530 and may provide an interface between other application software executed on the computer system 510 and the hardware resources of the computer system 510. More particularly, the operating system 534 may include a set of computer-executable instructions for managing the hardware resources of the computer system 510 and for providing common services to other application programs (e.g., managing memory allocation between various application programs). In certain exemplary embodiments, the operating system 534 may control the execution of one or more program modules depicted as stored in the data storage device 540. The operating system 534 may include any operating system known now or developed in the future, including but not limited to any server operating system, any mainframe operating system, or any other proprietary or non-proprietary operating system.
[0048] The computer system 510 may also include a disk / media controller 543 coupled to the system bus 521 to control one or more storage devices for storing information and instructions, such as a magnetic hard disk 541 and / or a removable media drive 542 (e.g., a floppy disk drive, a compact disc drive, a tape drive, a flash drive, and / or a solid state drive). The storage devices 540 may be added to the computer system 510 using an appropriate device interface (e.g., Small Computer System Interface (SCSI), Integrated Device Electronics (IDE), Universal Serial Bus (USB), or FireWire). The storage devices 541, 542 may be external to the computer system 510.
[0049] The computer system 510 may also include a field device interface 565 coupled to the system bus 521 to control field devices 566 (such as devices used in a production line). The computer system 510 may include a user input interface or GUI 561, which may include one or more input devices, such as a keyboard, a touch screen, a tablet computer, and / or a pointing device, for interacting with a computer user and providing information to the processor 520.
[0050] The computer system 510 may perform some or all of the processing steps of the embodiments of the present invention in response to one or more sequences of one or more instructions contained in a memory (such as the system memory 530) executed by the processor 520. Such instructions may be read into the system memory 530 from another computer-readable medium of the storage device 540, such as the magnetic hard disk 541 or the removable media drive 542. The magnetic hard disk 541 and / or the removable media drive 542 may contain data storage devices and data files used by one or more embodiments of the present disclosure. The data storage device 540 may include, but is not limited to, a database (e.g., relational, object-oriented, etc.), a file system, a flat file, a distributed data storage where data is stored on more than one node of a computer network, a peer-to-peer network data storage, etc. The data storage device may store various types of data, such as skill data, sensor data, or any other data generated according to the embodiments of the present disclosure. The content of the data storage device and the data files may be encrypted to improve security. The processor 520 may also be employed in a multiprocessing device to execute one or more sequences of instructions contained in the system memory 530. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions. Accordingly, the embodiments are not limited to any specific combination of hardware circuitry and software.
[0051] As described above, computer system 510 may include at least one computer-readable medium or memory for storing instructions programmed according to embodiments of the present invention and for containing data structures, tables, records, or other data herein. As used herein, the term "computer-readable medium" refers to any medium that participates in providing instructions to processor 520 for execution. Computer-readable media can take many forms, including but not limited to non-transitory, non-volatile media, volatile media, and transmission media. Non-limiting examples of non-volatile media include optical discs, solid state drives, magnetic disks, and magneto-optical discs, such as magnetic hard disk 541 or removable media drive 542. Non-limiting examples of volatile media include dynamic memory, such as system memory 530. Non-limiting examples of transmission media include coaxial cables, copper wire, and fiber optics, including the wires that make up system bus 521. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.
[0052] Computer-readable medium instructions for performing the operations of the present disclosure may be assembly instructions, instruction set-architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages (such as Smalltalk, C++, etc.) and conventional procedural programming languages (such as the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may establish a connection to an external computer (e.g., using an Internet service provider through the Internet). In some embodiments, electronic circuits, including, for example, programmable logic circuits, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA), may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuits in order to perform aspects of the present disclosure.
[0053] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable medium instructions.
[0054] The computing environment 500 may also include a computer system 510 that operates in a networked environment using a logical connection to one or more remote computers, such as remote computing device 580. The network interface 570 may enable communication, for example, with other remote devices 580 or systems and / or storage devices 541, 542 via network 571. The remote computing device 580 may be a personal computer (laptop or desktop), a mobile device, a server, a router, a network PC, a peer device, or other common network nodes, and generally includes many or all of the elements described above with respect to computer system 510. When used in a networked environment, computer system 510 may include a modem 572 for establishing communication over a network 571 such as the Internet. The modem 572 may be connected to the system bus 521 via the user network interface 570 or via another suitable mechanism.
[0055] Network 571 may be any network or system commonly known in the art, including the Internet, an Intranet, a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a direct connection, or a series of connections, a cellular telephone network, or any other network or medium capable of facilitating communication between computer system 510 and other computers (e.g., remote computing device 580). Network 571 may be wired, wireless, or a combination thereof. A wired connection may be implemented using Ethernet, Universal Serial Bus (USB), RJ-6, or any other wired connection commonly known in the art. A wireless connection may be implemented using Wi-Fi, WMAX, and Bluetooth, infrared, cellular networks, satellites, or any other wireless connection method commonly known in the art. Additionally, several networks may work independently or communicate with each other to facilitate communication in network 571.
[0056] It should be understood that the program modules, applications, computer-executable instructions, code, etc. depicted in Figure 5 system memory 530 are illustrative only and not exhaustive, and the processing described as being supported by any specific module may alternatively be distributed across multiple modules or performed by different modules. Additionally, various program modules, scripts, plugins, application programming interfaces (APIs), or any other suitable computer-executable code hosted locally on computer system 510 / remote device 580 and / or on other computing devices accessible via one or more networks 571 may be provided to support the functions and / or additional or alternative functions provided by the Figure 5 program modules, applications, or computer-executable code shown. Additionally, the functionality may be modularized differently such that the processing described as being performed by Figure 5The processing supported jointly by the set of program modules shown can be performed by fewer or more modules, or the functionality described as being supported by any specific module can be at least partially supported by another module. Additionally, according to any suitable computing model (e.g., client-server model / peer-to-peer model, etc.), the program modules that support the functionality described herein can form part of one or more applications executable on any number of systems or devices. Further, any functionality described as being supported by any program module depicted in Figure 5 can be implemented at least partially in hardware and / or firmware on any number of devices.
[0057] It should be further understood that, without departing from the scope of the present disclosure, computer system 510 can include alternative and / or additional hardware, software, or firmware components in addition to those described or depicted. More particularly, it should be understood that the software, firmware, or hardware components depicted as forming part of computer system 510 are illustrative only, and in various embodiments, certain components may be absent or additional components may be provided. Although various illustrative program modules have been depicted and described as software modules stored in system memory 530, it should be understood that the functionality described as being supported by the program modules can be implemented by any combination of hardware, software, and / or firmware. It should also be understood that in various embodiments, each of the above modules can represent a logical partitioning of the supported functionality. This logical partitioning is depicted for ease of explaining the functionality, and it may not represent the structure of the software, hardware, and / or firmware for implementing the functionality. Thus, it should be understood that in various embodiments, the functionality described as being provided by a specific module can be at least partially provided by one or more other modules. Additionally, in certain embodiments, one or more of the depicted modules may be absent, and in other embodiments, additional modules not depicted may be present and can support at least a portion of the described functionality and / or additional functionality. Further, although certain modules may be depicted and described as sub-modules of another module, in certain embodiments, these modules can be provided as independent modules or sub-modules of other modules.
[0058] Although specific embodiments of the present disclosure have been described, those of ordinary skill in the art will recognize that many other modifications and alternative embodiments are within the scope of the present disclosure. For example, any functionality and / or processing capabilities described relative to a specific device or component can be performed by any other device or component. Additionally, although various illustrative implementations and architectures have been described in accordance with embodiments of the present disclosure, those of ordinary skill in the art will recognize that many other modifications to the illustrative implementations and architectures described herein are within the scope of the present disclosure. Further, it should be understood that any operation, element, component, data, etc. described herein as being based on another operation, element, component, data, etc. can additionally be based on one or more other operations, elements, components, data, etc. Accordingly, the phrase "based on" or variations thereof should be interpreted as "at least partially based on".
[0059] Although embodiments have been described in language specific to structural features and / or methodological acts, it should be understood that the present disclosure is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as illustrative forms of implementing the embodiments. Conditional language, such as "can", "could", "might", or "may", unless specifically stated otherwise or otherwise understood within the context in which it is used, is generally intended to convey that certain embodiments can include, while other embodiments do not include, certain features, elements, and / or steps. Thus, such conditional language is generally not intended to imply that the features, elements, and / or steps are in any way required for one or more embodiments, or that one or more embodiments must include logic for determining, with or without user input or prompting, whether these features, elements, and / or steps are included or are to be performed in any particular embodiment.
[0060] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may not occur in the order noted in the figures. For example, depending on the functionality involved, two blocks shown in succession may in fact be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a special purpose hardware-based system that performs the specified functions or acts, or combinations of special purpose hardware and computer instructions.
Claims
1. A method for performing multiple computing tasks, the method comprises: receiving, by an orchestration engine via a network, task requests as containers from a plurality of process engines corresponding to respective machines, each task request indicating at least one real-time task requirement for performing a corresponding computing task among the multiple computing tasks, the orchestration engine being included in a computing platform, the computing platform being remote compared to the process engines; selecting a plurality of computing instances from a set of computing instances, the selected computing instances respectively having available computing resources, the computing instances being included in the computing platform and the computing instances communicating with the orchestration engine; collecting performance data associated with the network; generating a predicted running time for each of the computing tasks based on the performance data associated with the network; determining a scheduling and allocation scheme based on the predicted running time, the task requirements, and the available computing resources, the scheduling and allocation scheme defining when to execute each of the multiple computing tasks and which of the multiple selected computing instances executes each of the multiple computing tasks; and executing the multiple computing tasks by the multiple selected computing instances according to the scheduling and allocation scheme.
2. The method according to claim 1, the method further comprises: obtaining historical performance data associated with the computing tasks or the multiple selected computing instances; and generating the predicted running time for the computing tasks based on the historical performance data.
3. The method according to claim 1 or 2, wherein generating the scheduling and allocation scheme includes performing the discrete stochastic optimization using the predicted running time, task requirements, and available computing resources as inputs to the discrete stochastic optimization.
4. The method according to claim 3, the method further comprises: performing the discrete stochastic optimization upon receiving each task.
5. The method according to claim 1, the method further comprises: comparing the predicted running time with a corresponding predetermined threshold; when the predicted running time is greater than the corresponding predetermined threshold, determining the process engine that initiated the task request associated with the predicted running time greater than the corresponding predetermined threshold; and sending an alert to the determined process engine.
6. The method according to claim 1, the method further comprises: comparing the predicted running time with a corresponding predetermined threshold; and identifying additional computing resources when at least one of the predicted running times is greater than the predetermined threshold.
7. The method according to claim 1, wherein determining the scheduling and allocation scheme further comprises: generating a priority order for completing the task requests based on the task requirements and performance of the network; generating an optimization problem based on the priority order and the available computing resources; and solving the optimization problem to generate the scheduling and allocation scheme due to the optimization problem.
8. A system for performing computing tasks, the system comprises: a process for an execution module; and a memory for storing the module, comprising: A data module, configured to receive, via a network, task requests as containers from a plurality of process engines corresponding to respective machines, each task request indicating at least one real-time task requirement for performing a corresponding one of the plurality of computing tasks, wherein an orchestration engine including the data module is included in a computing platform, and the computing platform is remote compared to the process engines; An artificial intelligence module, configured to collect performance data associated with the network and generate a predicted run time for each of the computing tasks based on the performance data associated with the network; An orchestrator module, configured to select a plurality of computing instances from a set of computing instances, the selected computing instances respectively having available computing resources, the computing instances being included in the computing platform and the computing instances communicating with the orchestration engine, wherein the orchestrator module is further configured to determine a scheduling and allocation scheme based on the predicted run time, the task requirements, and the available computing resources, the scheduling and allocation scheme defining when to execute each of the plurality of computing tasks and which of the plurality of selected computing instances executes each of the plurality of computing tasks; and the system further includes the plurality of selected computing instances configured to execute the plurality of computing tasks according to the scheduling and allocation scheme.
9. The system according to claim 8, wherein the artificial intelligence module is further configured to: obtain historical performance data associated with the computing tasks or the plurality of selected computing instances; and generate the predicted run time for the computing tasks based on the historical performance data.
10. The system according to claim 8, wherein the orchestrator module is further configured to perform discrete stochastic optimization using the predicted run time, the task requirements, and the available computing resources as inputs to the discrete stochastic optimization to generate the scheduling and allocation scheme.
11. The system according to claim 8, wherein the artificial intelligence module is further configured to: compare the predicted run time with a corresponding predetermined threshold; when the predicted run time is greater than the corresponding predetermined threshold, determine the process engine that initiated the task request associated with the predicted run time greater than the corresponding predetermined threshold; and send an alert to the determined process engine.
12. The system according to claim 8, wherein the orchestrator module is further configured to: generate a priority order for completing the task request based on the task requirements and performance of the network; generate an optimization problem based on the priority order and the available computing resources; and solve the optimization problem to generate the scheduling and allocation scheme due to the optimization problem.
Citation Information
Patent Citations
Cost-minimizing task scheduler
CN105074664A
System and method for dynamic allocation of resources in a computing grid
US20070240161A1
Minimizing execution time of a compute workload based on adaptive complexity estimation
US20180103088A1