A computer implementation, computer system, and computer program for managing multiple jobs using multiple job processing pools.
Patent Information
- Application Number
- JP2022168723
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-10-21
- Filing Date
- 2022-10-20
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2042-10-20
AI Technical Summary
【0010】 例示的な実施態様は、時間の閾値量を超えてアイドル状態になっているアイドル状態の複数のジョブプロセッサを前記コンピュータシステムによって許容的に取り除くことができる。結果として、例示的な実施態様は、ジョブ処理の全体的なスループットのうちの少なくとも1つが増加されるところのコンピュータシステムにおいて複数のジョブを実行する際の性能を向上させること、複数のジョブを実行する為の待機時間が短縮されること、又はジョブプロセッサによって保持されるリソースが、アイドル状態の複数のジョブプロセッサを選択的に取り除くことによっていかなる時でも削減されることの技術的効果を提供することができる。例示的な実施態様は、複数の前記ジョブの為の時間ベースのジョブ予測を前記コンピュータシステムによって決定し;及び、複数の前記ジョブの為の前記時間ベースのジョブ予測を使用して、前記複数のジョブ処理プールのサイズを前記コンピュータシステムによって許容的に調和させることができる。結果として、該例示的な実施態様は、ジョブ処理の全体的なスループットのうちの少なくとも1つが増加されるところのコンピュータシステムにおいて複数のジョブを実行する際の性能を向上させること、複数のジョブを実行する為の待機時間が短縮されること、又はジョブプロセッサによって保持されるリソースが、将来のジョブの予測を使用してジョブ処理プールを管理することを通じて、いかなる時でも削減されることの技術的効果を提供することができる。複数の前記ジョブの為の前記時間ベースのジョブ予測を使用して、前記複数のジョブ処理プールのサイズを前記コンピュータシステムによって調和させる際に、例示的な実施態様は、前記時間ベースのジョブ予測に基づいてサイズの縮小を要求する前記複数のジョブ処理プールから複数の現在のジョブプロセッサを前記コンピュータシステムによって取り除き;及び、前記時間ベースのジョブ予測と複数のリソース制限の1セットとに基づいて、新しい複数のジョブプロセッサを前記複数のジョブ処理プールに、前記コンピュータシステムによって許容的に追加することができる。結果として、該例示的な実施態様は、ジョブ処理の全体的なスループットのうちの少なくとも1つが増加されるところのコンピュータシステムにおいて複数のジョブを実行する際の性能を向上させること、複数のジョブを実行する為の待機時間が短縮されること、又はジョブプロセッサによって保持されるリソースが、ジョブの為に、時間ベースのジョブ予測に基づいて、ジョブプロセッサを前記複数のジョブ処理プールからを選択的に追加且つ削除することによって、いかなる時でも削減されることの技術的効果を提供することができることの技術的効果を提供することができる。複数のリソース制限の前記1セットに基づいて、前記新しい複数のジョブプロセッサを前記複数のジョブ処理プールに前記コンピュータシステムによって追加する際に、例示的な実施態様は、前記時間ベースのジョブ予測に基づいて、サイズの増加を要求する前記複数のジョブ処理プールの為の1つの追加リソース要件を前記コンピュータシステムによって決定し;前記新しい複数のジョブプロセッサを前記複数のジョブ処理プールに追加することが複数のリソース制限の前記1セットを超えないという決定に応答して、前記新しい複数のジョブプロセッサを前記コンピュータシステムによって追加すること;前記新しい複数のジョブプロセッサを前記複数のジョブ処理プールに追加することが複数のリソース制限の前記1セットを超えるという決定に応答して、前記複数のジョブ処理プールの為の集約ジョブ優先度と、前記ジョブの為の時間ベースのジョブ予測に基づく前記複数のジョブ処理プール内の前記1つのジョブプロセッサの為の複数のリソース要件と、に基づく順序に、前記複数のジョブ処理プールを前記コンピュータシステムによって並べ替えること;及び、複数のリソース制限の前記1セットを超えることなしに、前記順序に基づいて前記新しい複数のジョブプロセッサを前記複数のジョブ処理プールに前記コンピュータシステムによって許容的に追加することができ、ここで、複数のリソース制限の前記1セットは、グローバルリソース制限、コンシューマリソース制限、又はコンテナタイプごとの制限のうちの1つから選択される。結果として、該例示的な実施態様は、ジョブ処理の全体的なスループットのうちの少なくとも1つが増加されるところのコンピュータシステムにおいて複数のジョブを実行する際の性能を向上させること、複数のジョブを実行する為の待機時間が短縮されること、又はジョブプロセッサによって保持されるリソースが、ジョブの為に時間ベースのジョブ予測に基づいてジョブプロセッサを追加するときにリソースの制限を考慮することによって、いかなる時でも削減されることの技術的効果を提供することができる。例示的な実施態様はまた、実行時間閾値を超える期間で現在のジョブが実行していることに応答して、前記1つのジョブ処理プール内の前記1つのジョブプロセッサから前記複数のジョブ処理プール以外の専用の1つのジョブプロセッサに前記現在の1つのジョブを前記コンピュータシステムによって許容的に移動することができる。結果として、該例示的な実施態様は、ジョブ処理の全体的なスループットのうちの少なくとも1つが増加されるところのコンピュータシステムにおいて複数のジョブを実行する際の性能を向上させること、複数のジョブを実行する為の待機時間が短縮されること、又はジョブプロセッサによって保持されるリソースが、そのジョブの為に別のジョブプロセッサ内のジョブ処理プール内でジョブを実行する必要な閾値よりも長い実行時間を有するジョブを実行することによって、いかなる時でも削減されることの技術的効果を提供することができる。
Smart Images

Figure 0007909356000001 
Figure 0007909356000002 
Figure 0007909356000003
Abstract
Description
Technical Field
[0002]
[0001] The present invention generally relates to an improved computer system, particularly to container pool management, and more specifically to managing workloads within a container pool.
Background Art
[0002] In the training of a machine learning model, model training jobs are executed to train the machine learning model. These model training jobs can be sent to a training service. The training service can receive model training jobs for execution in a short period and with high frequency. These jobs can complete the training of the machine learning model in a short time. It is necessary to make the resources for processing these training jobs available in a short period or immediately.
[0003] Currently, each model training job waits until the resources in the cluster become available for executing the model training job. The cluster can be, for example, a Kubernetes cluster composed of nodes that execute containerized applications. These containerized applications are designed to execute model training jobs. When resource availability exists, the model training job is scheduled to execute. The scheduling includes the time overhead in creating and initializing the containers required for executing the model training job. Further, the scheduling assumes fixed processing resources, where the calculation focuses on selecting a job and allocating the selected job to these resources.
[0004] This type of process may not reduce the waiting time required to run model training jobs. Because the number of jobs increases and their frequency changes as needed, this type of process may not reduce the waiting time required to run model training jobs.
[0005] Therefore, it is desirable to have methods and apparatus that take into account at least some of the problems discussed above, as well as other possible problems. For example, it would be desirable to have methods and apparatus that overcome technical problems by increasing the efficiency of performing multiple jobs, such as model training jobs. [Overview of the project] [Problems that the invention aims to solve]
[0006] The objective is to provide a computer-implemented method, computer system, and computer program product or computer program for managing multiple jobs using multiple job processing pools. [Means for solving the problem]
[0007] According to one exemplary embodiment, a computer-implemented method is provided for managing multiple jobs using multiple job processing pools. A computer system receives one job having one job type. The computer system identifies one job processing pool in the multiple job processing pools to execute multiple jobs of the job type, wherein the one job processing pool comprises multiple job processors for executing multiple jobs of the job type, and another one job processing pool in the multiple job processing pools comprises other multiple job processors for executing multiple jobs of different job types. The computer system executes the job having the job type using one job processor of the job type in the one job processing pool for the job type.
[0008] In other exemplary embodiments, a computer system and computer program product or computer program is provided for managing multiple jobs using multiple job processing pools. As a result, the exemplary embodiments can provide a technical effect of improving performance when running multiple jobs in a computer system, where at least one of the overall job processing throughputs is increased, waiting times for running multiple jobs are reduced, or the resources held by the job processor are reduced at any given time.
[0009] When the computer system executes the job using the one job processor of the job type in the one job processing pool for the job type, the exemplary embodiment allows the computer system to execute the job to some extent using the one job processor of the job type in the one job processing pool for the job type, in response to the fact that the one job processor of the job type in the one job processing pool for the job type is available to execute the job. As a result, the exemplary embodiment can provide technical effects such as improved performance when executing multiple jobs in a computer system where at least one of the overall throughput of job processing is increased, reduced waiting time for executing multiple jobs, or reduced at any time by using multiple available job processors having the same processing type as the job. When the computer system executes the job using the one job processor of the job type in the one job processing pool for the job type, an exemplary embodiment allows the computer system to determine whether the one job processor of the job type can be added to the one job processing pool to execute the job having the job type; adds the one job processor of the job type to the one job processing pool in response to the determination that the one job processor of the job type can be added to the one job processing pool to execute the job having the job type, and the determination that adding the one job processor of the job type to the one job processing pool does not exceed one set of multiple resource limits; and allows the computer system to execute the job using the job processor of the job type.As a result, exemplary embodiments can provide technical effects such as improved performance when running multiple jobs in a computer system where at least one of the overall job processing throughputs is increased, reduced waiting times for running multiple jobs, or reduced resources held by job processors at any time by taking into account a set of multiple resource limits. When the computer system adds the one job processor of a job type in response to a decision that the one job processor of the job type can be added to run the job having the job type, the exemplary embodiment allows the computer system to remove a set of idle job processors to free up resources to add the one job processor; and the computer system adds the one job processor of the job type to run the job having the job type. As a result, exemplary embodiments can provide technical effects such as improved performance when running multiple jobs in a computer system where at least one of the overall job processing throughputs is increased, reduced waiting times for running multiple jobs, or reduced resources held by job processors at any time by freeing up resources.In response to a decision that the one job processor of the job type can be added to execute the job having the job type, an exemplary embodiment of adding the one job processor of the job type by the computer system involves the computer system allowing the plurality of job processing pools to be sorted in an order based on aggregate job priorities for the plurality of job processing pools and resource requirements for the one job processor in the plurality of job processing pools; the computer system removing multiple idle job processors from the plurality of job processing pools having higher ranks so that sufficient resources are available to add the one job processor; and the computer system adding the one job processor of the job type to the one job processing pool. As a result, exemplary embodiments can provide technical benefits such as improved performance when running multiple jobs in a computer system where at least one of the overall job processing throughputs is increased, reduced waiting times for running multiple jobs, or reduced at any time by removing idle jobs based on pool ranking using aggregate priority, freeing up resources, and adding new multiple job processors.In response to a decision that the one job processor of the job type can be added to execute the job having the job type, when the computer system adds the one job processor of the job type, an exemplary embodiment allows the computer system to identify a set of multiple job candidates having a lower priority than the job, where the resources freed by moving the set of multiple job candidates are sufficient to add the one job processor of the job type; the computer system places the set of multiple job candidates in a waiting queue; and the computer system adds the one job processor of the job type to the one job processing pool. As a result, the exemplary embodiment can provide technical effects such as improved performance when executing multiple jobs in a computer system where at least one of the overall throughput of job processing is increased, reduced waiting times for executing multiple jobs, or reduced resources held by the job processor at any time by moving lower priority jobs to a waiting queue and freeing up resources for higher priority jobs.
[0010] An exemplary embodiment allows the computer system to selectively remove multiple idle job processors that have been idle for longer than a time threshold. As a result, the exemplary embodiment can provide technical effects such as improved performance when running multiple jobs in a computer system where at least one of the overall job processing throughputs is increased, reduced waiting times for running multiple jobs, or a reduction in resources held by job processors at any time by selectively removing multiple idle job processors. The exemplary embodiment allows the computer system to determine time-based job forecasts for the multiple jobs; and allows the computer system to selectively harmonize the size of the multiple job processing pools using the time-based job forecasts for the multiple jobs. As a result, the exemplary embodiment can provide technical effects such as improved performance when running multiple jobs in a computer system where at least one of the overall job processing throughputs is increased, reduced waiting times for running multiple jobs, or a reduction in resources held by job processors at any time by managing the job processing pools using future job forecasts. When the computer system harmonizes the size of the multiple job processing pools using the time-based job forecast for the multiple jobs, an exemplary embodiment allows the computer system to remove multiple current job processors from the multiple job processing pools that are requested to be reduced in size based on the time-based job forecast; and allows the computer system to permissibly add multiple new job processors to the multiple job processing pools based on the time-based job forecast and a set of multiple resource limits.As a result, the exemplary embodiment can provide the technical effect of improving performance when running multiple jobs in a computer system where at least one of the overall throughputs of job processing is increased, reducing waiting times for running multiple jobs, or reducing the resources held by job processors at any time by selectively adding and removing job processors from the multiple job processing pools based on time-based job forecasts for jobs. When the computer system adds the new multiple job processors to the multiple job processing pools based on the set of multiple resource limits, the exemplary embodiment determines, based on the time-based job forecast, one additional resource requirement for the multiple job processing pools that requests an increase in size; the computer system adds the new multiple job processors in response to the determination that adding the new multiple job processors to the multiple job processing pools does not exceed the set of multiple resource limits; In response to a decision that exceeds a certain limit, the computer system may reorder the multiple job processing pools in an order based on aggregated job priorities for the multiple job processing pools and multiple resource requirements for one of the job processors in the multiple job processing pools based on time-based job forecasts for the jobs; and the computer system may permissibly add the new multiple job processors to the multiple job processing pools based on the order without exceeding the one set of multiple resource limits, where the one set of multiple resource limits is selected from one of global resource limits, consumer resource limits, or limits per container type.As a result, the exemplary embodiment can provide technical effects such as improved performance when running multiple jobs in a computer system where at least one of the overall job processing throughputs is increased, reduced waiting times for running multiple jobs, or a reduction in resources held by job processors at any time by considering resource constraints when adding job processors for jobs based on time-based job forecasts. The exemplary embodiment can also allow the computer system to move the current job from the one job processor in the one job processing pool to a dedicated job processor outside of the multiple job processing pools in response to the current job running for a period exceeding an execution time threshold. As a result, the exemplary embodiment can provide technical effects such as improved performance when running multiple jobs in a computer system where at least one of the overall job processing throughputs is increased, reduced waiting times for running multiple jobs, or a reduction in resources held by job processors at any time by running jobs that have an execution time longer than the threshold required to run the job in a job processing pool in another job processor for that job.
[0011] According to further exemplary embodiments, the present invention is a computer program product for managing multiple jobs using multiple job processing pools, comprising a computer-readable recording medium having program instructions embedded in a computer-readable medium, wherein the program instructions, which are executable by a computer system, are provided to the computer system. To receive one job having one job type through the computer system, To execute multiple jobs of the job type, the computer system identifies one job processing pool within the multiple job processing pools, wherein the one job processing pool comprises multiple job processors for executing multiple jobs of the job type, and another job processing pool within the multiple job processing pools comprises other multiple job processors for executing multiple jobs of different job types, and Executing the job having the job type by the computer system using one job processor in the one job processing pool for the job type. The present invention provides a computer program product that enables the execution of the method described above. [Brief explanation of the drawing]
[0012] [Figure 1] Figure 1 shows a cloud computing environment in which an exemplary embodiment may be implemented. [Figure 2] Figure 2 shows an abstraction model layer according to an exemplary embodiment. [Figure 3] Figure 3 is a schematic representation of the network of a data processing system in which an exemplary embodiment may be implemented. [Figure 4] Figure 4 is a block diagram of a job processing environment according to an exemplary embodiment. [Figure 5] Figure 5 shows a container pool for jobs having various job types, according to an exemplary embodiment. [Figure 6] Figure 6 is a diagram of resource limitations according to an exemplary embodiment. [Figure 7] Figure 7 is a flowchart of a method for managing multiple jobs according to an exemplary embodiment. [Figure 8] Figure 8 is a flowchart of a method for performing a job according to an exemplary embodiment. [Figure 9]FIG. 9 is a flowchart of a method for executing one job according to an exemplary embodiment. [Figure 10] FIG. 10 is a flowchart of a method for adding one job processor according to an exemplary embodiment. [Figure 11] FIG. 11 is a flowchart of a method for adding one job processor according to an exemplary embodiment. [Figure 12] FIG. 12 is a flowchart of a method for adding one job processor according to an exemplary embodiment. [Figure 13] FIG. 13 is a flowchart of a method for managing a plurality of resources for processing a pool according to an illustrated example. [Figure 14] FIG. 14 is a flowchart of a method for managing one job using time-based job prediction managed according to an illustrated example. [Figure 15] FIG. 15 is a flowchart of a method for coordinating a plurality of job processing pools according to an exemplary embodiment. [[ID=二十]] [Figure 16] FIG. 16 is a flowchart of a method for adding a new plurality of job processors according to an exemplary embodiment. [Figure 17] FIG. 17 is a flowchart of a method for managing a plurality of jobs according to an exemplary embodiment. [Figure 18] FIG. 18 is a flowchart of a method for performing time-based job prediction according to an exemplary embodiment. [Figure 19] FIG. 19 is a flowchart of a method for coordinating a plurality of container pools according to an exemplary embodiment. [Figure 20] FIG. 20 is a flowchart of a method for removing a plurality of idle containers according to an exemplary embodiment. [Figure 21] FIG. 21 is a flowchart of a process for performing job preemption according to an exemplary embodiment. [Figure 22]FIG. 22 is a block diagram of a data processing system according to an exemplary embodiment. **DETAILED DESCRIPTION OF THE INVENTION**
[0013] The present invention can be a system, a method, a computer program, a computer program product, or a combination thereof at any possible technical detail level of integration. The computer program product can include one or more computer-readable storage media having computer-readable program instructions for causing a processor to execute aspects of the present invention.
[0014] The computer-readable storage medium may be a tangible device capable of holding and storing instructions for use by an instruction execution device. The computer-readable storage medium may, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of the computer-readable storage medium includes: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory sticks, floppy disks, mechanically encoded devices such as punch cards or grooved structures on which instructions are recorded, or any suitable combination thereof. As used herein, a computer-readable storage medium should not be interpreted as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through optical fiber cables), or electrical signals transmitted via wires.
[0015] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to individual computing devices / processors, or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may consist of copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface in each computing device / processor receives computer-readable program instructions from the network and transmits them for storage in a computer-readable storage medium within the individual computing device / processor.
[0016] The computer-readable program instructions for performing the operation of the present invention may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, such as object-oriented programming languages (e.g., SMALLTALK®, C++, etc.) or conventional procedural programming languages (e.g., the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, partially as a standalone software package on the user's computer, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, such as a local area network (LAN) or a wide area network (WAN), or such connection may be made to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuits, such as programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), can execute computer-readable program instructions by personalizing the electronic circuit using state information of computer-readable program instructions in order to carry out the aspects of the present invention.
[0017] The aspects of the present invention are described herein with reference to flowcharts, block diagrams, or combinations thereof, of methods, apparatus (systems), and computer programs or computer program products according to embodiments of the present invention. It will be understood that each block in the flowcharts, block diagrams, or combinations thereof, and of combinations of blocks in the flowcharts, block diagrams, or combinations thereof, can be implemented by computer-readable program instructions.
[0018] These computer-readable program instructions can be provided to a computer processor or other programmable data processing device to create a machine, such that the instructions, executed via the processor of the computer or other programmable data processing device, generate means for implementing functions / operations specified in one or more blocks of the flowchart or block diagram or a combination thereof. These computer-readable program instructions can also be stored in a computer-readable storage medium that can instruct a computer-programmable data processing device or other device or a combination thereof to function in a particular manner, such that the stored instructions include a product containing instructions that implement the functional / operational aspects specified in one or more blocks of the flowchart or block diagram or a combination thereof.
[0019] The computer-readable program instructions may also be loaded onto the computer, other programmable data processing device, or other device such that instructions executed on the computer, other programmable data processing device, or other device implement the functions / operations specified in one or more blocks of the flowchart or block diagram or combination thereof, causing the computer, other programmable device, or other device to execute a series of operational steps to generate a computer-implemented process.
[0020] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer programs or computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or part thereof of instructions, which includes one or more executable instructions for implementing one or more specified logical functions. In some alternative implementations, the functions shown in the blocks may occur in a different order than that shown in the drawings. For example, two consecutively shown blocks may actually be achieved as a single step executed simultaneously, substantially simultaneously, partially or entirely in a temporally overlapping manner, depending on the functions involved, or the blocks may be executed in reverse order. It should also be noted that each block in the block diagram or flowchart or a combination thereof, and any combination of blocks in the block diagram or flowchart or a combination thereof, may be implemented by a special-purpose hardware-based system that performs the specified functions or operations, or by a combination of special-purpose hardware and computer instructions.
[0021] While this disclosure includes a detailed description of cloud computing, it should be understood that the implementation of the teachings enumerated herein is not limited to cloud computing environments. Rather, embodiments of this disclosure can be implemented in combination with any other type of computing environment currently known or to be developed later.
[0022] Cloud computing is a service delivery model that enables convenient on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with service providers. This cloud model may include at least five features, at least three service models, and at least four deployment models.
[0023] The features are as follows:
[0024] On-demand self-service: Cloud consumers can unilaterally provision computing functions, such as server time and network storage, as needed, without requiring human interaction with the service provider.
[0025] Broad network access: The functionality is available over a network and accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).
[0026] Resource pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, and various physical and virtual resources are dynamically allocated and reallocated according to demand. Consumers generally do not have control or knowledge of the exact location of the resources provided, but can identify the location at a higher level of abstraction (e.g., country, state, or data center), making it location-independent.
[0027] Rapid Adaptability: Features can be provisioned quickly and flexibly, and in some cases automatically, they can scale out quickly, be released quickly, and scale in quickly. For consumers, the features available for provisioning are often unlimited and can be purchased at any amount at any time.
[0028] Service Measurement: Cloud systems automatically control and optimize resource usage by employing metric functions at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both service providers and consumers.
[0029] The service model is as follows:
[0030] Software as a Service (SaaS): This refers to the functionality provided to consumers for using a provider's applications running on a cloud infrastructure. These applications are accessible from various client devices through a thin client interface, such as a web browser (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, such as the network, servers, operating system, storage, or even the underlying cloud infrastructure encompassing individual application functions, with the possible exception of limited, user-specific application configuration settings.
[0031] Platform as a Service (PaaS): A feature provided to a consumer for deploying applications they have created or acquired, generated using programming languages and tools supported by the provider, onto a cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, such as the network, servers, operating system, or storage, but has control over the deployed applications and, in some cases, the application hosting environment configuration.
[0032] Infrastructure as a Service (IaaS): This is a service provided to a consumer to provision processing, storage, networking, and other basic computing resources, enabling the consumer to deploy and run any software, including operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but has limited control over the operating system, storage, deployed applications, and, in some cases, network components (e.g., the host's firewall).
[0033] The deployment model is as follows:
[0034] Private Cloud: A cloud infrastructure is operated solely for a specific organization. This cloud infrastructure may be managed by that organization or a third party, and may reside on-premises or off-premises.
[0035] Community Cloud: Cloud infrastructure is shared by several organizations and supports a specific community that shares common interests (e.g., mission, security requirements, policies, and compliance considerations). The cloud infrastructure may be managed by the organization or a third party and may reside on-premises or off-premises.
[0036] Public cloud: Cloud infrastructure is available to the general public or large industry groups and is owned by organizations that sell cloud services.
[0037] Hybrid Cloud: Cloud infrastructure is a hybrid of two or more clouds (private, community, or public) that remain separate entities but are brought together by standardized or proprietary technologies (e.g., cloud bursting for load balancing between clouds) that enable data and application migration.
[0038] Cloud computing environments are oriented services that focus on statelessness, low coupling, modularity, and semantic interoperability. The heart of cloud computing is the infrastructure, which includes a network of interconnected nodes.
[0039] Referring now to Figure 1, an exemplary embodiment of a cloud computing environment in which such an environment may be implemented is illustrated. In this specific example, the cloud computing environment 100 comprises one or more cloud computing nodes 110, and local computing devices used by cloud consumers, such as a personal digital assistant or smartphone 120A, a desktop computer 120B, a laptop computer 120C, or an automotive computer system 120N, or a combination thereof, can communicate with the cloud computing nodes 110.
[0040] The cloud computing nodes 110 can communicate with each other and can be physically or virtually grouped into one or more networks, for example, private clouds, community clouds, public clouds, or hybrid clouds as described herein, or a combination thereof. This allows the cloud computing environment 100 to provide infrastructure, platforms, or software, or a combination thereof, as a service that does not require cloud consumers to maintain resources on local computing devices, for example, local computing devices 120A to 120N. It is understood that the types of local computing devices 120A to 120N are intended to be illustrative only, and that the cloud computing nodes 110 and the cloud computing environment 100 can communicate with any type of computerized device via any type of network or network addressable connection or a combination thereof (for example, using a web browser).
[0041] Referring now to Figure 2, a diagram is depicted showing an abstraction model layer according to an exemplary embodiment. The set of functional abstraction layers shown in this specific example may be provided by a cloud computing environment, e.g., cloud computing environment 100 in Figure 1. It should be understood that the components, layers, and functions shown in Figure 2 are intended to be illustrative only, and that embodiments of the present invention are not limited thereto. The following layers and corresponding functions are provided as illustrated:
[0042] The abstraction layer 200 of the cloud computing environment includes the hardware and software layer 202, the virtualization layer 204, the management layer 206, and the workload layer 208. The hardware and software layer 202 includes the hardware and software components of the cloud computing environment. Hardware components may include, for example, a mainframe 210, RISC (Reduced Instruction Set Computer) architecture-based servers 212, 214, and blade servers 216, as well as storage devices 218 and network and networking components 220. In some embodiments, software components may include, for example, network application server software 222 and database software 224.
[0043] The virtualization layer 204 provides an abstraction layer from which the following examples of virtual entities are provided: namely, a virtual server 226, virtual storage 228, a virtual network 230, a virtual network 230 encompassing, for example, a virtual private network, a virtual application and operating system 232, and a virtual client 234.
[0044] In one example, the management layer 206 may provide the functions described below. Resource provisioning 236 provides the dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment. Measurement and pricing 238 provides cost tracking when resources are used within the cloud computing environment and billing or invoices for the consumption of these resources. In one example, these resources may include application software licenses. Security provides identification and verification of cloud consumers and tasks, and protection of data and other resources. The user portal 240 provides access to the cloud computing environment for consumers and system administrators. Service level management 242 provides the allocation and management of cloud computing resources to ensure that the required service levels are met. Service Level Agreement (SLA) planning and execution 244 provides the pre-placement and procurement of cloud computing resources for which future requirements are anticipated in accordance with the SLA.
[0045] Workload layer 208 provides examples of functions that can be utilized in a cloud computing environment. Examples of workloads and functions that can be provided by workload layer 208 include mapping and navigation 246, software development and lifecycle management 248, virtual classroom education delivery 250, data analysis processing 252, transaction processing 254, and container orchestration 256.
[0046] In this example, container orchestration 256 provides a service for managing the deployment of applications within the cloud computing environment 100 in Figure 1 or within the network at physical locations accessing the cloud computing environment 100 in Figure 1. In a specific example, container orchestration 256 can provide automated scaling and management of application deployments. For example, container orchestration 256 can automate the deployment, scaling, and operation of containers across various clusters of hosts. In this example, a container image is an executable software package containing all the components necessary to run an application. A container is an executable instance of an image.
[0047] These containers can be grouped into a pool by container orchestration 256 and run jobs. For example, container orchestration 256 can act as a training service where training jobs are used to train learning models.
[0048] Referring now to Figure 3, a graphical representation of a network of a data processing system in which an exemplary embodiment may be implemented is shown. The network data processing system 300 is a network of computers in which an exemplary embodiment may be implemented. The network data processing system 300 includes a network 302, which is a medium used to provide communication links between various devices connected together within the network data processing system 300 and computers. The network 302 may include connections such as wired communication links, wireless communication links, or fiber optic cables.
[0049] In the illustrated example, server computers 304 and 306 connect to network 302 along with storage unit 308. In addition, client device 310 connects to network 302. As shown, client device 310 includes client computers 312, 314, and 316. Client device 310 can be, for example, a computer, a workstation, or a network computer. In the illustrated example, server computer 304 provides information, such as boot files, operating system images, and applications, to client device 310. Furthermore, client device 310 can also include other types of client devices, such as mobile phones 318, tablet computers 320, and smart glasses 322. In this specific example, server computers 304, 306, storage unit 308, and client device 310 are network devices that connect to network 302, which is the communication medium for these network devices. Some or all of the client devices 310 can form the Internet of Things (IoT), where these physical devices can connect to the network 302 and exchange information with each other via the network 302.
[0050] In this example, client device 310 is a client to server computer 304. The network data processing system 300 may include additional server computers, client computers, and other devices not shown. Client device 310 connects to network 302 using at least one of wired, fiber optic, or wireless connections.
[0051] Program code located in the network data processing system 300 can be stored on a computer-recordable storage medium and downloaded to the data processing system or other devices for use. For example, program code can be stored on a computer-recordable storage medium in the server computer 304 and downloaded to the client device 310 via the network 302 for use by the client device 310.
[0052] In the illustrated example, the network data processing system 300 is the Internet, which has a network 302 representing a global collection of networks and gateways that use the Transmission Control Protocol / Internet Protocol (TCP / IP) suite to communicate with one another. At the heart of the Internet is a backbone of high-speed data communication lines between major node computers or host computers, which consist of thousands of commercial, government, educational, and other computer systems that route data and messages. Of course, the network data processing system 300 can also be implemented using several different types of networks. For example, network 302 can consist of at least one of the following: the Internet, an intranet, a local area network (LAN), a metropolitan area network (MAN), or a wide area network (WAN). Figure 3 is intended as an example and not as an architectural limitation for various exemplary embodiments.
[0053] As used herein, "several" when used in reference to an item means one or more items. For example, "several different types of networks" means one or more different types of networks.
[0054] Furthermore, when used in a list of items, the phrase "at least one" means that various combinations of one or more of the listed items can be used, and that only one of each item in the list may be required. In other words, "at least one" means that any combination of items and number of items in the list can be used, but not all items in the list are required. The item can be a specific object, thing, or category.
[0055] For example, “at least one of item A, item B, or item C” could include item A, item A and item B, or item B. This example could also include item A, item B and item C, or item B and item C. Of course, any combination of these items is possible. In some specific examples, “at least one of” could be, for example, two items A; one item B; and ten items C; four items B and seven items C; or any other suitable combination.
[0056] The container orchestration platform 330 could be, for example, a Kubernetes® architecture or environment. However, it should be understood that any specific example using Kubernetes refers only to an exemplary architecture and does not limit itself to exemplary embodiments. The container orchestration platform 330 can also be referred to as a container orchestration system.
[0057] The container orchestration platform 330 provides a platform for automating the deployment, scaling, and operation of customer applications, such as a job manager 332. The job manager 332 executes multiple jobs 334 received from one or more client devices 310. Multiple jobs 334 can be received dynamically. Furthermore, the frequency at which multiple jobs 334, each having different job types, are received can vary.
[0058] The container orchestration platform 330 provides automated deployment, scaling, and operation of multiple container pools 336 for use when executing multiple jobs 334. Each container pool within the multiple container pools 336 contains several containers 338 for executing multiple jobs 334 received by the job manager 332. The multiple containers 338 within a single container pool are for processing multiple jobs 334. In other words, the multiple containers 338 within a single container pool are for processing a specific type of job.
[0059] A container is an executable instance of a container image. A container image contains an executable package of software components necessary to run an application. A container is a standard unit of software for an application that packages program instructions and all their dependencies, and therefore the application can run in multiple computing environments. A container isolates the software from the environment in which it runs and ensures that the container behaves uniformly across different environments. A container for an application can share the operating system kernel on the machine with other containers for other applications. As a result, an operating system is not required for each container running on the machine.
[0060] In this specific example, the job manager 332 sends multiple jobs 334 to different container pools in order to execute multiple jobs 334 based on their job types. By sending multiple jobs 334 to specific container pools based on a single job type, the job manager 332 can improve efficiency when executing multiple jobs. When the containers that execute the jobs are designed or configured to execute jobs of that job type, and this is more efficient than executing different jobs of different job types, organizing multiple containers 338 into multiple container pools 336 based on job type can improve job processing throughput and reduce waiting times until resources become available to execute multiple jobs 334.
[0061] In this specific example, the job manager 332 can harmonize the sizes of multiple container pools 336 based on the types of multiple jobs 334 received for processing or the types of multiple jobs 334 expected to be received. Furthermore, predictive analytics can be used to predict when multiple jobs 334 will be received over time. This prediction can be used to adjust the multiple container pools 336 where the multiple jobs 334 are executed by the container orchestration platform 330 to improve efficiency.
[0062] As a result, the container orchestration platform 330 can improve performance when running multiple jobs within the network data processing system 300, thereby increasing at least one of the overall throughputs of job processing, reducing waiting times for running multiple jobs, or reducing the resources held by multiple job processors at any given time.
[0063] Referring now to Figure 4, a block diagram of a job processing environment according to an exemplary embodiment is shown. In this specific example, the job processing environment 400 comprises components that can be implemented in hardware, for example, the hardware shown in the cloud computing environment 100 in Figure 1 and the network data processing system 300 in Figure 3.
[0064] As shown in the diagram, the job processing environment 400 is an environment in which multiple jobs 402 are processed by multiple job processing pools 404, each having multiple job processors 406.
[0065] Multiple jobs 402 can take several different forms. For example, multiple jobs 402 can be selected from at least one of the following: a model training job for training machine learning models, a scheduling job for semiconductor manufacturing operations, an order processing job, or other appropriate types of tasks or operations.
[0066] Multiple job processors 406 can take several different forms. For example, multiple job processors 406 can be selected from at least one of the following: containers, pods, threads, processes, applications, operating system instances, virtual machines, hosts, clusters, processing units, or other appropriate types of processing components.
[0067] In this specific example, the job management system 412 can manage multiple job processing pools 404 and multiple jobs 402 that are currently running. In this specific example, the job management system 412 comprises a computer system 414 and a job manager 420.
[0068] The job manager 420 is located in the computer system 414 and can be implemented in software, hardware, firmware, or a combination thereof. When software is used, the operations performed by the job manager 420 can be implemented in hardware, such as a processor unit, using program instructions configured to run on it. When firmware is used, the operations performed by the job manager 420 can be implemented in program instructions and data, stored in persistent memory and executed on the processor unit. When hardware is used, the hardware may include circuitry that operates to perform operations in the job manager 420.
[0069] In specific examples, the hardware can take the form of at least one of a circuit system, an integrated circuit, an application-specific integrated circuit (ASIC), a programmable logic device, or other suitable types of hardware configured to perform several operations. A programmable logic device can be configured to perform several operations. The device can be reconfigured later or permanently configured to perform several operations. Programmable logic devices include, for example, programmable logic arrays, programmable array logic, field-programmable logic arrays, field-programmable gate arrays, and other suitable hardware devices. Additionally, the process can be implemented in organic components integrated with inorganic components and can be entirely composed of organic components excluding humans. For example, the process can be implemented as a circuit in an organic semiconductor.
[0070] The computer system 414 is a physical hardware system and includes one or more data processing systems. If two or more data processing systems are present in the computer system 414, those data processing systems communicate with each other using a communication medium. The communication medium may be a network. The data processing systems can be selected from at least one of a computer, a server computer, a tablet computer, or other suitable data processing systems.
[0071] As illustrated, the computer system 414 comprises several processor units 422 capable of executing program instructions 424 that implement a process in a specific example. As used herein, one of the several processor units 422 is a hardware device and consists of hardware circuitry, such as hardware circuitry in an integrated circuit that responds to and processes instructions and program code that operate the computer. When several processor units 422 execute program instructions 424 for a process, these several processor units 422 are one or more processor units that may be located on the same computer or on different computers. In other words, the process can be distributed among processor units in the same or different computers within the computer system. Furthermore, these several processor units 422 may be of the same type or different types. For example, several processor units can be selected from at least one of the following: single-core processors, dual-core processors, multi-processor cores, general-purpose central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), or several other types of processor units.
[0072] In this specific example, the job manager 420 in the job management system 412 manages multiple jobs 402 running using multiple job processing pools 404 based on multiple job types 426. Each job processing pool is configured to process multiple jobs 402 of the same job type. For example, one job processing pool 428 is for executing multiple jobs 402 of one job type 430 among the multiple job types 426. Multiple job processors 406 in one job processing pool 428 among the multiple job processing pools 404 are configured or designed to execute multiple jobs 402 which are all of job type 430. In this specific example, another job processing pool among the multiple job processing pools 404 may have other multiple job processors 406 for executing multiple jobs 402 of a different job type than job type 430 for the multiple job processors 406. As a result, different multiple job processing pools can process multiple jobs of various job types among the multiple job types 426.
[0073] In this specific example, the job manager 420 can receive one job 432 having one job type 430. The job manager 420 can identify one job processing pool 428 among multiple job processing pools 404 to execute multiple jobs 402 of one job type 430. One job processing pool 428 has multiple job processors 406 for executing multiple jobs 402 of one job type 430. The job manager 420 uses one job processor 434 within one job processing pool 428 for one job type 430 to execute one job 432 having one job type 430.
[0074] When executing a single job 432, the job manager 420 can execute the single job 432 using the single job processor 434 in the single job processing pool 428 for a single job type 430, in response to the fact that a single job processor 434 in the single job processing pool 428 for a single job type 430 is available to execute the single job 432.
[0075] In other specific examples, if one job processor is unavailable, the job manager 420 can perform several different steps to execute one job 432. For example, the job manager 420 can determine whether one job processor 434 of one job type 430 can be added to one job processing pool 428 to execute one job 432 having one job type 430. In response to the determination that one job processor 434 of one job type 430 can be added to one job processing pool 428 to execute one job 432 having one job type 430, and the determination that adding one job processor 434 of one job type 430 to one job processing pool 428 does not exceed one set of multiple resource limits 436, the job manager 420 can add one job processor 434 of one job type 430 to one job processing pool 428. The job manager 420 can then use one job processor 434 to execute one job 432.
[0076] When adding a job processor 434, the job manager 420 can remove a set of idle job processors 444 to free up resources for adding the job processor 434. In a specific example, a job processor may be idle, waiting for one job, or idle, executing one job. In both states, the job processor holds multiple processing resources. In response to the release of resources, the job manager 420 can add a job processor 434 of a job type 430 and execute a job 432 having a job type 430.
[0077] In another specific example, when adding one job processor 434, the job manager 420 can reorder the multiple job processing pools 404 in an order based on the aggregated job priority 448 for the multiple job processing pools 404 and the resource requirements for the multiple job processors 406 within the multiple job processing pools 404. In a specific example, the multiple jobs 402 may have different job types as well as different priorities.
[0078] Aggregated job priorities can be determined for each job processing pool. The aggregated job priority for a single job processing pool is an aggregation or a portion of the priorities for multiple jobs being processed in that single job processing pool.
[0079] The job manager 420 can remove multiple idle job processors 444 from multiple job processing pools 404 that have a higher rank in order 446 so that a sufficient number of resources are available to add one job processor 434. In response to removing multiple idle job processors 444 so that a sufficient number of resources are freed, the job manager 420 can add one job processor 434 of one job type 430 to one job processing pool 428.
[0080] In another exemplary embodiment, when adding a job processor 434, the job manager 420 can identify a set of multiple job candidates 452 having a lower priority 454 than multiple jobs 402. The resources freed by moving the set of multiple job candidates 452 are sufficient to add a job processor 434 of a single job type 430. In this example, the job manager 420 places the set of multiple job candidates 452 into a waiting queue 456 and adds a job processor 434 of a single job type 430 to a job processing pool 428.
[0081] When managing multiple resources, the job manager 420 can also selectively remove multiple job processors 406. For example, the job manager 420 can remove multiple job processors 406 that are running for longer than a time threshold 458. In another specific example, in response to a current job 460 running for longer than a time threshold 464, the job manager 420 can move the current job 460 from a single job processor 434 in a single job processing pool 428 to a dedicated job processor 462 that is not in multiple job processing pools 404.
[0082] In this specific example, the value of the execution time threshold 464 can be based on an analysis of past jobs. The value of the execution time threshold 464 can be selected to ensure that jobs with execution times longer than desired or expected are not held for longer than desired or expected by job processors other than the multiple job processing pools 404, so that multiple job processors 406 in the multiple job processing pools 404 are not held for longer than desired or expected in order to process multiple jobs 402.
[0083] Alternatively, a single dedicated job processor 462 can be used to handle these multiple long-running jobs. Due to factors such as inputs, delays, dependencies, interactions, or other factors, jobs may run longer than expected or anticipated. As a result, the job management system 412 can be optimized to handle multiple shorter-running jobs that have execution times shorter than the execution time threshold 464. Multiple long-running jobs can be handled using external job processors in multiple job processing pools 404.
[0084] Additionally, the job manager 420 may implement a time-based job forecast 466 to predict multiple jobs 402 that will be submitted over time. In this specific example, the job manager 420 can use the time-based job forecast 466 to manage the number of multiple job processors 406 and the number of multiple job processing pools 404 over different periods.
[0085] In this specific example, time-based job prediction 466 uses historical job processing data for jobs of different job types. Time-based job prediction 466 can be performed by building a mathematical model that captures trends in the historical job data. This model can then be used to predict future jobs. In other specific examples, a machine learning model can be trained using historical job data to perform time-based job prediction 466.
[0086] In one specific example, the job manager 420 can determine a time-based job forecast 466 for multiple jobs 402. The job manager 420 can use the time-based job forecast 466 for multiple jobs 402 to harmonize the sizes of multiple job processing pools 404. This harmonizing can be performed by adding one job processor to one or more of the multiple job processing pools 404 or removing one or more of the multiple job processing pools 404.
[0087] When harmonizing multiple job processing pools 404, the job manager 420 can remove multiple current job processors 468 from multiple job processing pools 404 that request size reduction based on time-based job forecasts 466. The job manager 420 can add multiple new job processors 470 to multiple job processing pools 404 based on a set of time-based job forecasts 466 and multiple resource limits 436. In other words, the addition of multiple new job processors 470 can be based on time-based job forecasts 466 such that the addition of the multiple new job processors 470 does not exceed the set of multiple resource limits 436.
[0088] When adding new multiple job processors 470, the job manager 420 can determine additional resource requirements 472 for the multiple job processing pools 404 requesting an increase in size, based on time-based job forecasts 466. In response to the determination that adding new multiple job processors 470 to the multiple job processing pools 404 does not exceed one set of multiple resource limits 436, the job manager 420 can add the new multiple job processors 470. In response to the determination that adding new multiple job processors 470 to the multiple job processing pools 404 exceeds one set of multiple resource limits 436, the job manager 420 can reorder the multiple job processing pools 404 in the order 446 based on aggregated job priorities 448 for the multiple job processing pools 404 and resource requirements 450 for the multiple job processors 406 within the multiple job processing pools 404 based on time-based job forecasts 466 for the multiple jobs 402. The job manager 420 can add multiple new job processors 470 to multiple job processing pools 404 based on the sequence 446 without exceeding one set of multiple resource limits 436.
[0089] In one specific example, there exists one or more technical solutions to overcome the technical problems related to multiple jobs running with a desired level of efficiency. As a result, one or more technical solutions can provide technical effects such as improving performance when running multiple jobs in a computer system where at least one of the overall throughputs of job processing is increased, reducing waiting times for running multiple jobs, or reducing the resources held by the job processor at any given time. In a specific example, the overall throughput when running multiple jobs can be increased within various time intervals. Furthermore, waiting times can also be reduced across various types of jobs. In other words, the job processor can be made to run jobs immediately or as quickly as possible. In addition, the management of multiple container resources held across different containers is reduced at any given time. Furthermore, these efficiencies can take into account job priorities when running multiple jobs.
[0090] The computer system 414 can be configured to perform at least one of a process, operation, or action using software, hardware, firmware, or a combination thereof, as described in several different specific examples. As a result, the computer system 414 operates as a purpose-specific computer system, in which a job manager 420 within the computer system 414, where jobs can be executed by the computer system 414, enables improved efficiency. In particular, the job manager 420 transforms the computer system 414 into a purpose-specific computer system compared to a currently available general-purpose computer system that does not have a job manager 420.
[0091] In a specific example, the use of the Job Manager 420 in computer system 414 integrates processes into an application that improves the performance of computer system 414 in order to execute multiple jobs. In other words, the Job Manager 420 in computer system 414 is directed towards a practical application of the integrated processes in the Job Manager 420 in computer system 414, which identifies a job processing pool based on the job type of the jobs to be executed.
[0092] The illustration of the job processing environment 400 in Figure 4 does not imply any physical or structural limitations on how the exemplary embodiment may be implemented. Other components may be used in addition to or instead of those shown. Some components may be unnecessary. Also, several blocks are presented to illustrate some functional components. One or more of these blocks may be combined, separated, combined or combined, and divided into different blocks when implemented in the exemplary embodiment.
[0093] For example, multiple job processing pools 404 may have two job processing pools that process multiple jobs 402 of the same job type. In this example, the multiple job processing pools 404 still cover all processing of multiple job types 426 for multiple jobs 402. Overlapping multiple job processing pools can exist in several specific examples.
[0094] Next, moving to Figure 5, a diagram of multiple container pools for multiple jobs having different job types is shown, according to an exemplary embodiment. In this example, jobs 500, 502, and 504 are examples of the multiple jobs 402 in Figure 4. These jobs have various job types. In this example, there are three unique job types for these multiple jobs. As shown in the figure, job 500 has job type 1, job 502 has job type 2, and job 504 has job type 3.
[0095] These multiple jobs can be executed by multiple containers, for example, multiple containers 506 in one container pool 508, multiple containers 510 in one container pool 512, and multiple containers 514 in one container pool 516. These containers are example implementations for multiple job processors 406, and the multiple container pools are example implementations for multiple job processing pools 404 in Figure 4.
[0096] Each container pool contains multiple containers, each unique in its purpose of running a single job of a specific type. For example, multiple containers 506 are for running type 1 jobs, multiple containers 510 are for running type 2 jobs, and multiple containers 514 are for running type 3 jobs. In this specific example, one container of a particular type can run a single job of the same type at any given time.
[0097] Multiple containers 506 of job type I can run multiple jobs 500 of job type I. Multiple containers 510 of job type 2 can run multiple jobs 502 of job type 2, and multiple containers 514 of job type 3 can run multiple jobs 504 of job type 3.
[0098] A diagram of a single container type for executing multiple jobs of different types is provided as one implementation example. This diagram is not intended to limit the ways in which other specific examples may be implemented. For example, in other examples, multiple pools may consist of pods, which are groups of multiple containers. In yet another specific example, other types of job processors, such as threads, processes, virtual machines, or other appropriate types of job processors, may be used. This example shows three container pools, but the number of container pools used may depend on the number of job types to be executed. In this example, multiple types of multiple containers associated with job types have a one-to-one relationship. For example, if there are five types of jobs, there may be five container pools, each with multiple containers specific to the execution of one particular job type.
[0099] Next, moving to Figure 6, a diagram of resource limits according to an exemplary embodiment is shown. In this specific example, resource limit 600 is an example of the set of multiple resource limits 436 in Figure 4. As illustrated, resource limit 600 is selected from at least one of the global resource limit 602, the job processor resource limit 604, and the consumer resource limit 606.
[0100] The global resource limit 602 can be a resource consumption limit for all job processors in all job processing pools. The job processor resource limit 604 can be a resource limit for each job processor. Different job processors may use different amounts of resources depending on the job type. The consumer resource limit 606 is a consumer limit and can be set for or by a consumer submitting a job. When adding a job processor, one or a combination of these resource limits may be considered.
[0101] In other specific examples, other types of resource limits may be used in addition to or instead of resource limit 600. For example, there may be pool resource limits, where resource limits are set for each job processing pool.
[0102] Next, referring to Figure 7, a flowchart of a process for managing multiple jobs, according to an exemplary embodiment, is shown. The process in Figure 7 can be implemented in hardware, software, or both. When implemented in software, the process can take the form of program instructions executed by one or more processor units located in one or more hardware devices within one or more computer systems. For example, the process in Figure 7 can be implemented in a job manager 420 within computer system 416 in Figure 4.
[0103] The process begins by receiving one job having one job type (step 700). The process identifies one job processing pool within the plurality of job processing pools to execute multiple jobs of the job type (step 702). In step 702, the one job processing pool comprises multiple job processors for executing multiple jobs of the job type, where another one job processing pool within the plurality of job processing pools comprises other multiple job processors for executing multiple jobs of different job types.
[0104] The process executes the job having the job type using one job processor for the job type in the one job processing pool for the job type (step 704). The process then terminates.
[0105] Referring now to Figure 8, a flowchart of the process for executing a job, according to an exemplary embodiment, is shown. This flowchart illustrates the implementation of step 704 in Figure 7.
[0106] The process executes the job having the job type using the one job processor for the job type in the one job processing pool for the job type, in response to the fact that the one job processor in the one job processing pool for the job type is available to execute the job (step 800). The process then terminates.
[0107] Referring to Figure 9, a flowchart of the process for performing a job, according to an exemplary embodiment, is shown. The process shown in Figure 9 is an example of step 704 in Figure 7.
[0108] The process begins by determining whether the one job processor of the job type can be added to the one job processing pool to execute the one job having the job type (step 900). The process adds the one job processor of the job type to the one job processing pool in response to the determination that the one job processor of the job type can be added to the one job processing pool to execute the job having the job type, and the determination that adding the one job processor of the job type to the one job processing pool does not exceed one set of multiple resource limits (step 902).
[0109] The process executes the job using the one job processor of the job type (step 904). The process then terminates.
[0110] Referring to Figure 10, a flowchart of the process for adding a job processor is shown, according to an exemplary embodiment. The process shown in Figure 10 is an example of step 902 in Figure 9.
[0111] The process begins by removing one set of idle job processors to free up resources to add the job processors (step 1000). In this specific example, step 1000 is any step. The process then adds the one job processor of the job type to execute the job having the job type (step 1002). The process then terminates.
[0112] Figure 11 shows a flowchart of the process for adding a job processor according to an exemplary embodiment. The process shown in Figure 11 is an example of step 902 in Figure 9.
[0113] The process begins by reordering the job processing pools in an order based on aggregated job priorities for the job processing pools and resource requirements for the job processors within the job processing pools (step 1100). The process then removes idle job processors from the job processing pools that have higher ranks in an order that sufficient resources are available to add one more job processor (step 1102).
[0114] The process adds the one job processor of the job type to the one job processing pool (step 1104). The process then terminates.
[0115] Next, referring to Figure 12, a flowchart of the process for adding a job processor according to an exemplary embodiment is shown. The process shown in Figure 12 is an example of step 902 in Figure 9.
[0116] The process begins by identifying a set of multiple job candidates having a lower priority than the job (step 1200). In step 1200, the resources freed by moving the set of multiple job candidates are sufficient to add the one job processor of the job type. The process then places the set of multiple job candidates into a waiting queue (step 1202).
[0117] The process adds the one job processor of the job type to the one job processing pool (step 1204). The process then terminates.
[0118] Referring to Figure 13, a flowchart of a process for managing multiple resources to process multiple pools is shown, according to an exemplary embodiment. The process shown in Figure 13 is an example of additional steps that may be used in the process in Figure 7.
[0119] The process begins by removing multiple idle job processors that have been idle for longer than a time threshold (step 1300). In response to the current job running for a period exceeding the execution time threshold, the process moves the current job from the one job processor in the one job processing pool to a dedicated job processor outside of the multiple job processing pools (step 1302). The process then terminates.
[0120] In step 1302, the execution time threshold can be selected based on various factors. For example, this threshold can be selected to increase performance across multiple jobs currently running. In a specific example, the threshold can be based on an analysis of several past jobs that were running where other execution time thresholds were used.
[0121] In this example, one job that exceeds the execution time threshold becomes a checkpoint. This job is stopped and placed in a queue to be executed using a dedicated container that was assigned to and started for that job. This container is outside of multiple container pools. Information about the job that exceeded the execution time threshold is recorded. This recorded information can be used to identify multiple similar jobs. These types of jobs can then be transferred to a dedicated job processor without needing to run within multiple container pools. In other words, multiple jobs that do not meet the characteristics required to run multiple jobs within the multiple container pools can be handled using a different mechanism.
[0122] In other specific examples, jobs that exceed their execution time do not need to be stopped. If a job that exceeds the execution time threshold has a high priority, the job may be allowed to continue running, and the job processor assigned to it may be removed from the processor pool.
[0123] Referring to Figure 14, a flowchart of a process for managing multiple jobs using time-based job forecasting management, according to an exemplary embodiment, is shown. The process shown in Figure 14 is an example of additional steps that may be used in the process in Figure 7.
[0124] The process begins by determining a time-based job forecast for the multiple jobs (step 1400). The process then uses the time-based job forecast for the multiple jobs to harmonize the sizes of the multiple job processing pools (step 1402). The process then ends.
[0125] Next, referring to Figure 15, a flowchart of the process for harmonizing a single job processing pool is shown, according to an exemplary embodiment. The process shown in Figure 15 is an implementation of step 1402 in Figure 14.
[0126] The process begins by removing several current job processors from the multiple job processing pools that request a reduction in size based on the time-based job forecast (step 1500). The process then adds several new job processors to the multiple job processing pools based on the time-based job forecast and a set of multiple resource limits (step 1502). The process then ends.
[0127] Referring now to Figure 16, a flowchart of the process for adding multiple new job processors according to an exemplary embodiment is shown. The process shown in Figure 16 is an example implementation of step 1502 in Figure 15.
[0128] The process begins by determining one additional resource requirement for the plurality of job processing pools that request an increase in size, based on the time-based job forecast (step 1600). In response to the determination that adding the new plurality of job processors to the plurality of job processing pools does not exceed one set of the plurality of resource limits, the process adds the new plurality of job processors without exceeding one set of the plurality of resource limits (step 1602).
[0129] In response to a decision that adding the new multiple job processors to the multiple job processing pools would exceed one set of multiple resource limits, the process reorders the job processing pools in an order based on aggregated job priorities for the multiple job processing pools and multiple resource requirements for the multiple job processors in the multiple job processing pools based on time-based job forecasts for the jobs (step 1604). The process adds the new multiple job processors to the multiple job processing pools in the order without exceeding one set of multiple resource limits (step 1606). The process then terminates. In step 1606, the one set of multiple resource limits is selected from one of global resource limits, consumer resource limits, or limits per container type.
[0130] Next, referring to Figure 17, a flowchart of a process for managing multiple jobs, according to an exemplary embodiment, is shown. The process in Figure 17 can be implemented in hardware, software, or both. When implemented in software, the process can take the form of program instructions executed by one or more processor units located in one or more hardware devices within one or more computer systems. For example, the process can be implemented in a job manager 420 within a computer system 416 in Figure 4.
[0131] The process begins by determining a time-based job forecast for multiple incoming jobs (step 1700). The process then harmonizes multiple container pools based on the time-based job forecast (step 1702). In this specific example, steps 1700 and 1702 are arbitrary steps.
[0132] The process receives one job to process (step 1704). In step 1704, the one job has one job type. The process determines whether one container of the job type is available to run the one job (step 1706). In step 1706, determining whether one container is available in one pool of the job type can be done by tracking multiple containers in the pool that are being used by multiple jobs at any given time and comparing the number of containers currently in use with the number of containers in the pool. If one container is available, the process runs the one job in the one container (step 1708), and then the process terminates.
[0133] Referring again to step 1706, if one container is unavailable, it is determined whether one container of the job type can be added, taking into account multiple resource limits (step 1710). If one container of the job type cannot be added (for example, because multiple resource limits have been exceeded), the process attempts to perform job preemption if job preemption is supported, or alternatively, if job preemption is not supported, the job is placed in a waiting queue (step 1712). The process for job preemption is shown in Figure 21 below. The process then terminates.
[0134] Referring again to step 1710, if one container of the job type can be added, the process removes idle containers to clear resources for the new container if additional resources are required (step 1714).
[0135] The process adds the one container of the job type for the job to the one container pool for the job type (step 1716). The process then proceeds to step 1708 as described above.
[0136] In addition, from step 1702, the process waits for a job completion event or timer to expire (step 1718). In response to receiving a job completion event or timer expiration, the process removes idle containers and frees up resources (step 1720). The targets for removing idle containers can be specified, for example, in terms of the requested amount of resources to be freed, the requested number of idle containers to be removed, or thresholds related to idleness in containers to be removed (e.g., idle for longer than a time threshold). The process identifies jobs that have exceeded the execution time threshold and switches these jobs to a single dedicated container other than the multiple container pools (step 1722). The process then terminates.
[0137] Referring now to Figure 18, a flowchart of the process for performing time-based job forecasting according to an exemplary embodiment is shown. The process in Figure 18 is one implementation example for process 1700 in Figure 17.
[0138] The process begins by identifying historical information about several previous jobs (step 1800). In step 1800, the historical information for each job includes attributes, e.g., (1) the calendar time when the job came in; (2) the type of job and the associated type of container for running the job; (3) the priority of the job; and (4) processing metrics, e.g., waiting time, execution time, resources consumed. The process uses the historical information about the several previous jobs to identify a function for predicting several future jobs (step 1802).
[0139] The process receives a period for which job predictions are desired (step 1804). The process predicts multiple jobs for the period using the function (step 1806). The process then terminates.
[0140] In this specific example, the output from the function may be a set of predicted jobs for the period. These predicted jobs may be tuples of time intervals within the period. Each tuple may include the required container type and the required number of containers. These tuples can be summarized or combined to identify the container type and the aggregated number of containers required for that container type during the period.
[0141] Referring to Figure 19, a flowchart of the process for harmonizing multiple container pools is depicted according to an exemplary embodiment. The process in Figure 19 is one implementation example for step 1702 in Figure 17. In this specific example, the time-based job forecast can be used to harmonize the current size of multiple pools to the predicted size for a future period. This harmonization can take into account various resource constraints. The process can add multiple containers to multiple pools where the forecast is larger than their current size, and remove multiple containers from multiple pools where the forecast is smaller than their current size. In this specific example, if there are no jobs of the type for multiple container pools in the forecast for the period, the multiple container pools can also be emptied and removed. If the predicted pool size is the same as the current pool size, the container pool size is not changed. Furthermore, if multiple resource constraints are met, the process can select the container pools to which multiple containers are added using the advantages over the cost metrics calculated for the multiple container pools.
[0142] The process begins by receiving a prediction result for a certain period (step 1900). The process identifies the number and type of containers required for the said period (step 1902). For example, it can be predicted that 10 containers of job type A, 5 containers of job type B, and 17 containers of job type C will be required for a given period.
[0143] The process determines the difference for each container pool (step 1904). In step 1904, the difference can be determined by subtracting the current size of each container pool from the predicted size of each container pool. The process removes multiple containers from multiple pools that are requesting a size reduction (step 1906).
[0144] Next, the process identifies several additional resource requirements for multiple container pools whose size is to be increased (step 1908). In step 1908, the additional resources required to run the required multiple containers are determined.
[0145] For each container pool where the required additional containers for that container pool exceed a container pool-based limit, the process reduces the required additional containers to satisfy the limit (step 1910). In step 1910, individual limits can be set for different container pools to limit the number of containers for specific job types. This type of limit is a container pool-based limit.
[0146] The process determines whether adding the required containers to the container pools exceeds the global resource limit or the consumer resource limit (step 1912). If it is determined that adding the required containers to the container pools does not exceed the global resource limit or the consumer resource limit, the process adds the required containers to the container pools (step 1914). The process then terminates.
[0147] Referring again to step 1912, if adding the required number of containers to multiple container pools exceeds the global resource limit or consumer resource limit, the process determines, on a container pool basis, the ratio of aggregated job priority to resource requirements (step 1916). In step 1916, the job priorities assigned to the multiple jobs are accumulated for each container pool over a period of time to provide an aggregated (e.g., average, mean, or median) job priority for each of the multiple container pools. The aggregated job priority for each container pool is divided by the resource requirements of one container of the container type in one of the container pools. In specific examples, a higher ratio indicates a better ratio of benefit to cost compared to a lower ratio.
[0148] The process sorts the multiple container pools by a ratio determined for each of the multiple container pools to generate an order (step 1918). In step 1918, the ranking of the multiple container pools can be in descending order from the highest ratio to the lowest ratio. The container pools ranked higher are better candidates for adding multiple containers in terms of the advantages and costs of adding multiple containers.
[0149] The process adds multiple containers to the multiple container pools using the sequence described above, without exceeding the consumer resource limit and until the global resource limit is reached (step 1920). The process then terminates.
[0150] Next, moving to Figure 20, a flowchart of the process for removing multiple idle containers is drawn according to an exemplary embodiment. This example shows one form in which step 1714 in Figure 17 can be implemented to remove multiple idle containers. The process begins by identifying multiple container pools having multiple idle containers (step 2000). In step 2000, one idle container may be one that has not run a job for longer than a threshold time. If time-based job forecasting is available, multiple container pools may be identified as multiple container pools having multiple idle containers that exceed the predicted number of required multiple containers or the predicted minimum limit of required containers for the one container pool. The forecasting function provides the number of containers of each container type to be used for a given future time interval. The function can also provide a range for each container type, where the range is the minimum and maximum number required, or the function can provide any combination of the minimum, maximum, and required numbers for each container type.
[0151] The process identifies the maximum number of containers that can be removed from each container pool (step 2002). In step 2002, the maximum number of containers is the number of idle containers in the container pool when time-based job forecasting is not available. When time-based job forecasting is available, the number of idle containers that can be removed is the number of idle containers that exceeds the predicted number of containers required or the predicted minimum limit of containers required for the container pool.
[0152] The process sorts the container pools in an order based on aggregated job priorities for multiple jobs within the container pools and the resource requirement sizes of multiple containers within the container pools (step 2004). In step 2004, a ratio of job priority to resource requirement size is generated for each container pool. These ratios are sorted in ascending order. As a result, the lower the job priority and the higher the resource requirement, the lower the cost-effectiveness ratio, and consequently, one container pool is ranked higher in the order of the container pools, and that one container pool is a more desirable candidate for removing multiple containers.
[0153] The process removes idle containers from the multiple container pools using the order from the sort of the multiple container pools, without exceeding the maximum number of containers that can be removed from each container pool, and until sufficient resources are available to add multiple containers (step 2006). In step 2006, the maximum number of containers that can be removed from each container pool, calculated in step 2002, is used. The process then terminates.
[0154] Referring to Figure 21, a flowchart of the process for performing job preemption is shown, according to an exemplary embodiment. The process in Figure 21 is an example of any steps that can be performed instead of step 1712 in Figure 21 when preemption is used.
[0155] The process begins by identifying the lowest priority job in all container pools having a lower priority than the current job received for processing (step 2100). A decision is made as to whether there are multiple low-priority jobs and whether the resources freed up by those multiple low-priority jobs are sufficient to create a container for the current job (step 2102). If there are multiple low-priority jobs and the resources freed up by those multiple low-priority jobs are sufficient, the process checkpoints and preempts the multiple low-priority jobs and places them in a waiting queue (step 2104). In step 2104, the multiple jobs are preempted to free up resources from the current single job.
[0156] The process uses the released resources to add a container having one job type of the current job (step 2106). The process then uses the added container to execute the job (step 2108). The process then terminates.
[0157] Referring again to step 2102, the process terminates if there are no multiple low-priority jobs or if there are not enough resources freed up by multiple low-priority jobs.
[0158] The flowcharts and block diagrams in the different embodiments shown illustrate the architecture, function, and operation of several possible implementations of the apparatus and method according to the exemplary embodiments. In this regard, each block in the flowchart or block diagram may represent at least one of a module, segment, function, or part of an operation or process. For example, one or more of the blocks may be implemented as program instructions, hardware, or a combination of program instructions and hardware. If implemented as hardware, the hardware may take the form of an integrated circuit manufactured or configured to perform one or more operations in the flowchart or block diagram. If implemented as a combination of program instructions and hardware, the implementation may take the form of firmware. Each block in the flowchart or block diagram may be implemented using a dedicated hardware system that performs various operations, or a combination of dedicated hardware and program instructions executed by the dedicated hardware.
[0159] In some alternative embodiments of the exemplary embodiments, one or more functions described in a block may occur in an order other than that shown in the figure. For example, in some cases, two blocks shown consecutively may be executed substantially simultaneously, or they may be executed in reverse order depending on the functions involved. In addition, other blocks may be added in addition to those shown in the flowchart or block diagram.
[0160] Referring now to Figure 22, a block diagram of a data processing system according to an exemplary embodiment is shown. The data processing system 2200 can be used to implement the cloud computing node 110, personal digital assistant (PDA) or smartphone 120A, desktop computer 120B, laptop computer 120C, or automotive computer system 120N in Figure 1, or a combination thereof. The data processing system 2200 can be used to implement the computer in the hardware and software layer 202 in Figure 2, as well as the server computer 304, server computer 306, and client device 310 in Figure 3. The data processing system 2200 can also be used to implement the computer system 414 in Figure 4. In this specific example, the data processing system 2200 includes a communication framework 2202, which provides communication between the processor unit 2204, memory 2206, persistent storage 2208, communication unit 2210, input / output (I / O) unit 2212, and display 2214. In this example, the communication framework 2202 takes the form of a bus system.
[0161] The processor unit 2204 is responsible for executing instructions for software that can be loaded into memory 2206. The processor unit 2204 comprises one or more processors. For example, the processor unit 2204 can be selected from at least one of the following: a multicore processor, a central processing unit (CPU), a graphics processing unit (GPU), a physics processing unit (PPU), a digital signal processor (DSP), a network processor, or any other suitable type of processor. Furthermore, the processor unit 2204 can be implemented using one or more heterogeneous processor systems in which the main processor resides together with secondary processors on a single chip. As another specific example, the processor unit 2204 can be a symmetrical multiprocessor system with multiple processors of the same type on a single chip.
[0162] Memory 2206 and persistent storage 2208 are examples of storage device 2216. A storage device is any part of hardware that can store information, such as, but not limited to, data, functional program instructions, or other suitable information, temporarily, permanently, or both. In these specific examples, storage device 2216 may also be referred to as a computer-readable storage device. In these examples, memory 2206 may be, for example, random-access memory or any other suitable volatile or non-volatile storage device. Persistent storage 2208 may take various forms depending on the specific implementation.
[0163] For example, persistent storage 2208 may comprise one or more components or devices. For instance, persistent storage 2208 may be a hard drive, a solid-state drive (SSD), flash memory, a rewritable optical disk, a rewritable magnetic tape, or any combination of the above. The media used by persistent storage 2208 may also be removable. For example, a removable hard drive may be used for persistent storage 2208.
[0164] In these specific examples, the communication unit 2210 provides communication with other data processing systems or devices. In these specific examples, the communication unit 2210 is a network interface card.
[0165] The input / output unit 2212 enables data input and output with other devices that can be connected to the data processing system 2200. For example, the input / output unit 2212 may provide a connection for user input via at least one of a keyboard, mouse, or some other suitable input device. Furthermore, the input / output unit 2212 may send output to a printer. The display 2214 provides a mechanism for displaying information to the user.
[0166] Instructions for at least one of the operating system, applications, or programs can be placed in the storage device 2216, which communicates with the processor unit 2204 via the communication framework 2202. Processes of different embodiments can be executed by the processor unit 2204 using instructions implemented in the computer, which can be placed in memory, for example, memory 2206.
[0167] These instructions are referred to as program instructions, computer-readable program instructions, or computer-readable program instructions that can be read and executed by a processor in the processor unit 2204. Program instructions in different embodiments may be embodied on different physical or computer-readable storage media, for example, on memory 2206 or persistent storage 2208.
[0168] The program instruction 2218 is arranged in a functional form on a computer-readable medium 2220, which is selectively removable and can be loaded or transferred to a data processing system 2200 for execution by a processor unit 2204. The program instruction 2218 and the computer-readable medium 2220 form a computer program product 2222 in these specific examples. In these specific examples, the computer-readable medium 2220 is a computer-readable storage medium 2224.
[0169] The computer-readable storage medium 2224 is not a medium for propagating or transmitting the program instructions 2218, but a physical or tangible memory device used to store the program instructions 2218. As used herein, the computer-readable storage medium 2224 should not be interpreted as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through optical fiber cables), or electrical signals transmitted via wires.
[0170] Alternatively, the program instruction 2218 can be transmitted to the data processing system 2200 using a computer-readable signal medium. This computer-readable signal medium is a signal, and may, for example, be a propagating data signal containing the program instruction 2218. For example, the computer-readable signal medium may be at least one of an electromagnetic signal, an optical signal, or any other suitable type of signal. These signals may be transmitted via a connection, such as a wireless connection, an optical fiber cable, a coaxial cable, a wire, or any other suitable type of connection.
[0171] Furthermore, as used herein, “computer-readable media 2220” may be singular or plural. For example, program instructions 2218 may be placed in computer-readable media 2220 in the form of a single storage device or system. In another example, program instructions 2218 may be placed in computer-readable media 2220 distributed across multiple data processing systems. In other words, some instructions within program instructions 2218 may be placed in one data processing system, while other instructions within program instructions 2218 may be placed in another data processing system. For example, some parts of program instructions 2218 may be placed in computer-readable media 2220 in a server computer, while other parts of program instructions 2218 may be placed in computer-readable media 2220 located in a set of client computers.
[0172] Accordingly, exemplary embodiments provide a computer-implemented method, computer system, and computer program product or computer program for the computer-implemented method for managing multiple jobs using multiple job processing pools. The computer system receives one job having one job type. The computer system identifies one job processing pool in the multiple job processing pools to execute the job of the job type, wherein the one job processing pool comprises job processors for executing multiple jobs of the job type, and another one job processing pool in the multiple job processing pools comprises other job processors for executing multiple jobs of different job types. The computer system executes the job having the job type using one job processor of the job type in the one job processing pool for the job type. According to other exemplary embodiments, computer systems and computer program products or computer programs for managing jobs are provided.
[0173] As a result, in specific examples, one or more technical solutions can provide technical effects such as improving performance when running multiple jobs in a computer system where at least one of the overall job processing throughputs is increased, reducing waiting times for running multiple jobs, or reducing the resources held by the job processor at any given time. In specific examples, the overall throughput when running multiple jobs can be increased at various time intervals. Furthermore, waiting times between multiple jobs of various types can also be reduced. In other words, the job processor is available to run jobs immediately or as quickly as possible. Furthermore, to manage containers, the resources held between different containers are reduced at any given time. Moreover, these efficiencies can take into account job priorities when running multiple jobs.
[0174] The descriptions of different exemplary embodiments are presented for illustrative and explanatory purposes and are not intended to be exhaustive or to limit the embodiments disclosed. Various specific examples describe components that perform actions or operations. In exemplary embodiments, components can be configured to perform the described actions or operations. For example, the component may have a configuration or design for a structure that provides the component with the ability to perform the actions or operations described in the specific examples when performed by the component. Furthermore, to the extent that the words “includes,” “including,” “has,” “contains,” and their variations are used herein, such words, like the word “comprises,” are intended to be inclusive as free-transition words without excluding additional or other elements.
[0175] The various embodiments of the present invention are presented for illustrative purposes only and are not intended to be exhaustive or to limit oneself to the disclosed embodiments. Not all embodiments include all features described in the specific examples. Furthermore, different exemplary embodiments may offer different features compared to other exemplary embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein have been selected to best describe the principles, practical applications, or technical improvements to the technologies available on the market of the embodiments, or to enable those skilled in the art to understand the embodiments disclosed herein.
Claims
1. A computer-implemented method for managing multiple jobs using multiple job processing pools, Receiving a single job with a single job type from a computer system. To execute multiple jobs of the job type, the computer system identifies one of the multiple job processing pools, where each job processing pool in the multiple job processing pools processes multiple jobs of the same job type, and at least two job processing pools in the multiple job processing pools are overlapping job processing pools for executing multiple jobs of the same job type, and one job processing pool comprises multiple job processors for executing multiple jobs of the job type, and another job processing pool in the multiple job processing pools comprises other multiple job processors for executing multiple jobs of different job types, and The computer system executes the job having the job type using one job processor of the job type within the one job processing pool for the job type. The method, including the method described above.
2. The computer system executes the job using the one job processor for the job type within the one job processing pool for the job type. In response to the fact that the one job processor in the one job processing pool for the job type is available to execute the job, the computer system executes the job using the one job processor for the job type in the one job processing pool for the job type. A computer-implemented method according to claim 1, including the following:
3. The computer system executes the job using the one job processor for the job type within the one job processing pool for the job type. The computer system determines whether the one job processor of the job type can be added to the one job processing pool to execute the job having the job type; The computer system adds the one job processor of the job type to the job processing pool in response to the determination that the one job processor of the job type can be added to the job processing pool to execute the job having the job type, and the determination that adding the one job processor of the job type to the job processing pool does not exceed one set of multiple resource limits; and Executing the job by the computer system using the one job processor of the job type A computer-implemented method according to claim 1, including the following:
4. In response to a decision that the one job processor of the job type can be added to execute the job having the job type, the computer system adds the one job processor of the job type. Removing one set of idle job processors by the computer system to free up resources to add the one job processor; and, The computer system adds one job processor of the job type and executes the job having the job type. A computer-implemented method according to claim 3, including the following:
5. In response to a decision that the one job processor of the job type can be added to execute the job having the job type, the computer system adds the one job processor of the job type. The computer system rearranges the multiple job processing pools in an order based on the aggregated job priority for the multiple job processing pools and the resource requirements for the multiple job processors within the multiple job processing pools; The computer system removes multiple idle job processors from the multiple job processing pools having higher ranks in the order so that sufficient resources are available to add the one job processor; and The computer system adds the one job processor of the job type to the one job processing pool. A computer-implemented method according to claim 3, including the following:
6. In response to a decision that the one job processor of the job type can be added to execute the job having the job type, the computer system adds the one job processor of the job type. The computer system identifies a set of multiple job candidates having a lower priority than the aforementioned job, wherein the resources freed up by moving the set of job candidates are sufficient to add one more job processor of the job type; The computer system places one set of multiple job candidates into a waiting queue; and, The computer system adds the one job processor of the job type to the one job processing pool. A computer-implemented method according to claim 3, including the following:
7. The computer system removes multiple idle job processors that have exceeded a time threshold. A computer-implemented method according to claim 1, further comprising:
8. Determining time-based job predictions for multiple jobs using the computer system; and, Using the time-based job forecast for the multiple jobs, the computer system harmonizes the size of the multiple job processing pools. A computer-implemented method according to claim 1, further comprising:
9. Using the time-based job forecast for the multiple jobs, the computer system can harmonize the sizes of the multiple job processing pools. The computer system removes a plurality of current job processors from the plurality of job processing pools that request a reduction in size based on the time-based job forecast; and, The computer system adds a number of new job processors to the number of job processing pools based on the time-based job forecast and a set of multiple resource limits. A computer-implemented method according to claim 8, including the following:
10. Based on the aforementioned set of multiple resource limits, the computer system may add the new multiple job processors to the multiple job processing pools. Based on the aforementioned time-based job forecast, the computer system determines one additional resource requirement for the multiple job processing pools that request an increase in size; The computer system adds the new job processors in response to a determination that adding the new job processors to the job processing pools does not exceed one set of resource limits; In response to a determination that adding the aforementioned new job processors to the aforementioned job processing pools exceeds one set of resource limits, the computer system reorders the aforementioned job processing pools in an order based on aggregated job priorities for the aforementioned job processing pools and resource requirements for the aforementioned job processors within the aforementioned job processing pools based on time-based job forecasts for the jobs; and, The computer system adds the new job processors to the job processing pools based on the order, without exceeding one set of multiple resource limits, wherein the one set of multiple resource limits is selected from one of global resource limits, consumer resource limits, or limits per container type. A computer-implemented method according to claim 9, including the following:
11. In response to the fact that the current job is running for a period exceeding the execution time threshold, the computer system moves the current job from the one job processor in the one job processing pool to a dedicated job processor other than the multiple job processing pools. A computer-implemented method according to claim 1, further comprising:
12. The computer-implemented method according to claim 1, wherein the plurality of job processors are selected from at least one of containers, pods, threads, processes, applications, operating system instances, virtual machines, hosts, or clusters.
13. A computer-implemented method for managing multiple jobs using multiple job processing pools, In response to receiving a job having a single job type, the computer system determines whether a container for that job type exists in a plurality of container pools to execute the job having that job type, where each container pool in the plurality of container pools executes multiple jobs for the selected job type, and at least two container pools in the plurality of container pools are overlapping job processing pools for executing multiple jobs of the same job type; In response to the availability of the aforementioned container, the computer system executes the job using the aforementioned container; In response to the unavailability of the aforementioned container, the computer system determines whether the aforementioned container for the job type can be added without exceeding one set of resource limits; and, In response to a decision that the one container for the job type can be added without exceeding the one set of resource limits, the computer system adds a container for the job type and executes the job; and, In response to the determination that the one container for the job type cannot be added without exceeding one set of multiple resource limits, the computer system places the job in a waiting queue. The method, including the method described above.
14. A computer system, It is equipped with several processor units, and here, the several processor units are Receive one job with one job type; To execute multiple jobs of the job type, one job processing pool is identified within a plurality of job processing pools, where each job processing pool within the plurality of job processing pools processes multiple jobs of the same job type, and at least two job processing pools within the plurality of job processing pools are overlapping job processing pools for executing multiple jobs of the same job type, and one job processing pool comprises multiple job processors for executing multiple jobs of the job type, and another job processing pool within the plurality of job processing pools comprises other multiple job processors for executing multiple jobs of different job types; and, Execute the job having the job type using one job processor in the one job processing pool for the job type. The computer system that executes program instructions in such a manner.
15. When executing the job using the one job processor of the job type in the one job processing pool for the job type, the several or more processor units, In response to the fact that the one job processor in the one job processing pool for the job type is available to execute the job, the job is executed using the one job processor for the job type in the one job processing pool for the job type. The computer system according to claim 14, which executes program instructions in such a manner.
16. When executing the job using the one job processor for the job type in the one job processing pool for the job type, the several processor units, Determine whether the one job processor of the job type can be added to the one job processing pool and execute the job having the job type; In response to the determination that the one job processor of the job type can be added to the one job processing pool to execute the job having the job type, and the determination that adding the one job processor of the job type to the one job processing pool does not exceed one set of multiple resource limits, the one job processor of the job type is added to the one job processing pool; and The job is executed using the one job processor of the job type. The computer system according to claim 14, which executes program instructions in such a manner.
17. In response to a decision that one job processor of the job type can be added to execute the job having the job type, when adding one job processor of the job type, the several or more processor units, Remove one set of idle job processors to free up resources to add the single job processor; and, Add the aforementioned job processor of the aforementioned job type and execute the job having the aforementioned job type. The computer system according to claim 16, which executes program instructions in such a manner.
18. In response to a decision that one job processor of the job type can be added to execute the job having the job type, when adding one job processor of the job type, the several or more processor units, The multiple job processing pools are sorted in an order based on the aggregated job priority for each of the multiple job processing pools and the resource requirements for one of the job processors within each of the multiple job processing pools; Remove multiple idle job processors from the multiple job processing pools having higher ranks so that sufficient multiple resources are available to add the one job processor; and, Add the one job processor of the job type to the one job processing pool. The computer system according to claim 16, which executes program instructions in such a manner.
19. In response to a decision that one job processor of the job type can be added to execute the job having the job type, when adding one job processor of the job type, the several or more processor units, Identify a set of multiple job candidates having a lower priority than the aforementioned job, wherein the resources freed by moving the set of job candidates are sufficient to add the aforementioned one job processor of the job type. The aforementioned set of multiple job candidates is placed in a waiting queue, and, Add the one job processor of the job type to the one job processing pool. The computer system according to claim 16, which executes program instructions in such a manner.
20. The aforementioned several processor units Remove multiple idle job processors that have exceeded a time threshold. The computer system according to claim 14, which executes program instructions in such a manner.
21. The aforementioned several processor units Determine time-based job forecasts for multiple jobs; and, The size of the multiple job processing pools is harmonized using the time-based job forecast for the multiple jobs. The computer system according to claim 14, which executes program instructions in such a manner.
22. When using the time-based job forecast for the multiple jobs to harmonize the sizes of the multiple job processing pools, the multiple processor units, Remove multiple current job processors from the multiple job processing pools that request a reduction in size based on the time-based job forecast; and, Based on the time-based job forecast and a set of resource limits, a new set of job processors is added to the set of job processing pools. The computer system according to claim 21, which executes program instructions in such a manner.
23. When adding the new job processors to the job processing pool based on the set of multiple resource limits, the several processor units, Based on the aforementioned time-based job forecast, determine one additional resource requirement for the multiple job processing pools that request an increase in size; In response to the determination that adding the new job processors to the job processing pools does not exceed one set of resource limits, the new job processors are added; In response to a determination that adding the new multiple job processors to the multiple job processing pools exceeds one set of multiple resource limits, the job processing pools are reordered in an order based on aggregated job priorities for the multiple job processing pools and resource requirements for one of the job processors in the multiple job processing pools based on time-based job forecasts for the jobs; and, Add the new job processors to the job processing pools based on the order, without exceeding one set of resource limits, where the set of resource limits is selected from one of global resource limits, consumer resource limits, or limits per container type. The computer system according to claim 22, which executes program instructions in such a manner.
24. The aforementioned several processor units In response to the current job running for a period exceeding the execution time threshold, the current job is moved from the one job processor in the one job processing pool to a dedicated job processor outside of the multiple job processing pools. The computer system according to claim 23, which executes program instructions in such a manner.
25. A computer program for managing multiple jobs using multiple job processing pools, wherein the computer program causes a computer system to perform each of the steps described in any one of claims 1 to 12.
Citation Information
Patent Citations
Job allocation method, device, and program
JP2013182569A
Stateful resource pool management for job execution
US20180060132A1