Content-Aware Node Selection Method, Program, and System

The content-aware node selection method optimizes container image deployment by selecting nodes based on manifest and mapping index comparisons, reducing network and storage needs, and enhancing performance through data reuse.

JP7710794B2Active Publication Date: 2025-07-22INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021174036
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-11-09
Filing Date
2021-10-25
Publication Date
2025-07-22
Estimated Expiration
2041-10-25

AI Technical Summary

Technical Problem

Current container images are slow to start up and require significant I/O and network resources due to large sizes, leading to high utilization of memory, storage, and network bandwidth, making them bulky and expensive to transfer and store.

Method used

A content-aware node selection method that identifies a computing node within a cluster based on a comparison of a container image's manifest and a mapping index to optimize storage and distribution, minimizing the need for redundant data transfer.

Benefits of technology

This approach reduces network traffic, storage requirements, and startup times by maximizing data reuse within the cluster, thereby improving performance and efficiency in container image deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007710794000001
    Figure 0007710794000001
  • Figure 0007710794000002
    Figure 0007710794000002
  • Figure 0007710794000003
    Figure 0007710794000003
Patent Text Reader

Abstract

To provide content-aware node selection for container creation.SOLUTION: A method, program and system for content-aware node selection includes: receiving a manifest for a container image of a container to be created; identifying a mapping index for a cluster of computing nodes; and selecting a computing node within the cluster of computing nodes to create the container, based on a comparison of the manifest to the mapping index.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to hardware virtualization, and more particularly, to the creation and deployment of container images.

Background Art

[0002] Modern application deployments often rely on the use of containers. For example, container images are distributed via a central registry, and a host pulls the container image and uses it to create the container's root file system in order to start the container. The number of container images as well as the speed of container deployment are increasing rapidly.

[0003] However, there are several problems with the current implementation of container images. For example, containers are currently slow to start up and I / O intensive because they require the download and storage of large container images, which results in high utilization of local memory or storage or both. Transferring large container images over a communication network also results in high network utilization and high load on the registry service storage subsystem. As a result, current container images are bulky and expensive to transfer and store.

[0004] Therefore, there is a need for a faster and more efficient way to store and distribute container images.

Summary of the Invention

Problems to be Solved by the Invention

[0005] Provide a content-aware node selection method, program, and system for container creation.

Means for Solving the Problems

[0006] The method according to one embodiment includes receiving a manifest of a container image of a container to be created, identifying a mapping index of a cluster of computing nodes, and selecting a computing node within the cluster of computing nodes based on a comparison of the manifest and the mapping index to create the container.

[0007] According to another embodiment, a computer program product for performing content-aware node selection (node selection aware of content) for container creation is a computer-readable storage medium in which program instructions are incorporated, the computer-readable storage medium not being a transient signal per se, and the program instructions causing a processor to receive a manifest of a container image of a container to be created, identify a mapping index of a cluster of computing nodes, and select a computing node within the cluster of computing nodes based on a comparison of the manifest and the mapping index to create the container, the computer-readable storage medium including a computer-readable storage medium executable by the processor for causing the processor to perform a method including the above.

[0008] According to another embodiment, a system includes a processor and logic integrated with the processor, logic executable by the processor, or logic integrated with and executable by the processor, the logic being configured to receive a manifest of a container image of a container to be created, identify a mapping index of a cluster of computing nodes, and select a computing node within the cluster of computing nodes based on a comparison of the manifest and the mapping index to create the container.

[0009] Other aspects and embodiments of the present invention will become apparent from the following detailed description, which, when considered in conjunction with the drawings, illustrate by way of example the principles of the present invention.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Modes for Carrying Out the Invention

[0011] The following description is made for the purpose of illustrating the general principles of the present invention and is not intended to limit the concepts of the invention claimed herein. Further, the specific features described herein can be used in combination with the other features described in each of various possible combinations and permutations.

[0012] Unless otherwise specifically defined herein, all terms should be given the broadest possible interpretation, including meanings implied from this specification, understood by those skilled in the art, or defined in a dictionary, treatise, etc., or both.

[0013] It should also be noted that as used in this specification and the appended claims, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. The terms "comprises", "comprising", or both as used herein specify the presence of the stated feature, integer, step, operation, element, or component, or combination thereof, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, or groups thereof, or combinations thereof.

[0014] The following description discloses several embodiments for performing content-aware node selection for container creation.

[0015] In one general embodiment, the method includes receiving a manifest of a container image of a container to be created, identifying a mapping index of a cluster of computing nodes, and selecting a computing node within the cluster of computing nodes based on a comparison of the manifest and the mapping index to create the container.

[0016] In another general embodiment, a computer program product for performing content-aware node selection for container creation is a computer-readable storage medium having program instructions incorporated therein, the computer-readable storage medium not being a transient signal per se, and the program instructions being executable by a processor to cause the processor to receive, at the processor, a manifest of a container image of a container to be created; identify, at the processor, a mapping index of a cluster of computing nodes; and select, at the processor, a computing node within the cluster of computing nodes based on a comparison of the manifest and the mapping index to create a container. The computer program product includes a computer-readable storage medium that is executable by a processor for causing the processor to execute a method including these operations.

[0017] In another general embodiment, a system includes a processor and logic integrated with the processor, logic executable by the processor, or logic integrated with and executable by the processor, the logic being configured to receive a manifest of a container image of a container to be created; identify a mapping index of a cluster of computing nodes; and select a computing node within the cluster of computing nodes based on a comparison of the manifest and the mapping index to create a container.

[0018] Although this disclosure includes a detailed description of cloud computing, it should be understood that the embodiments of the teachings recited herein are not limited to a cloud computing environment. Rather, embodiments of the invention may be implemented in connection with other types of computing environments now known or later developed.

[0019] Cloud computing is a service delivery model that enables convenient on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with a service provider. This cloud model can include at least five characteristics, at least three service models, and at least four deployment models.

[0020] The characteristics are as follows.

[0021] On-demand self-service: Cloud consumers can provision computing capabilities, such as server time and network storage, automatically as needed, without the need for human interaction with the service provider.

[0022] Broad network access: The capabilities are available over the network and accessed through standard mechanisms that promote use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).

[0023] Resource pooling: The provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, and various physical and virtual resources are dynamically assigned and reassigned according to demand. Consumers generally have no control over or knowledge of the exact location of the provided resources, but may have a sense of location independence in that they can potentially specify a higher level of abstraction location (e.g., country, state, or data center).

[0024] Rapid adaptability: Functions are provisioned quickly and adaptively, and in some cases automatically, scale out quickly, and are released quickly to scale in. To consumers, the functions available for provisioning often appear to be unlimited, and any amount can be purchased at any time.

[0025] Pay-per-use services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at any level of abstraction suitable for the type of service (e.g., storage, processing, bandwidth, and active user accounts). The usage of resources can be monitored, controlled, and reported, providing transparency to both the provider and consumer of the utilized services.

[0026] The service model is as follows.

[0027] Software as a Service (SaaS): The functions provided to consumers are to use the provider's applications running on the cloud infrastructure. The applications are accessible from various client devices through a thin-client interface such as a web browser (e.g., web-based email). Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application functions, with the exception of limited user-specific application configurations.

[0028] Platform as a Service (PaaS) as a service: The function provided to the consumer is to deploy consumer-created or acquired applications created using programming languages and tools supported by the provider onto the cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure including the network, server, operating system, or storage, but controls the deployed applications and, according to some, the application hosting environment configuration.

[0029] Infrastructure as a Service (IaaS) as a service: The function provided to the consumer is to provision processing, storage, network, and other basic computing resources, and the consumer can deploy and run any software that can include an operating system and applications. The consumer does not manage or control the underlying cloud infrastructure, but controls the operating system, storage, deployed applications, and, according to some, has limited control over selected networking components (e.g., host firewall).

[0030] The deployment models are as follows.

[0031] Private cloud: The cloud infrastructure is operated only for an organization. The cloud infrastructure can be managed by the organization or a third party and can exist inside or outside the facility.

[0032] Community cloud: The cloud infrastructure is shared by several organizations and supports a specific community with common concerns (e.g., mission, security requirements, policies, and compliance considerations). The cloud infrastructure can be managed by the organization or a third party and can exist inside or outside the facility.

[0033] Public cloud: The cloud infrastructure is available to the general public or a large industry group and is owned by an organization that sells cloud services.

[0034] Hybrid cloud: The cloud infrastructure remains a distinct entity but is a composite of two or more clouds (private, community, or public) joined by standard or proprietary technologies (such as cloud bursting for load balancing between clouds) that enable data and application portability.

[0035] Cloud computing environments are service-oriented, emphasizing statelessness, loose coupling, modularity, and semantic interoperability. At the center of cloud computing is an infrastructure that includes a network of interconnected nodes.

[0036] Next, referring to FIG. 1, an exemplary cloud computing environment 50 is shown. As illustrated, cloud computing environment 50 includes one or more cloud computing nodes 10 that can communicate with a local computing device used by a cloud consumer, such as a personal digital assistant (PDA) or cellular phone 54A, desktop computer 54B, laptop computer 54C, or in-vehicle computer system 54N, or a combination thereof. Nodes 10 can communicate with each other. Nodes 10 may be physically or virtually grouped in one or more networks such as a private, community, public, or hybrid cloud as described above, or a combination thereof (not shown). Thereby, cloud computing environment 50 can provide infrastructure, platform, software, or a combination thereof as a service such that a cloud consumer need not maintain resources on a local computing device. It is intended that the types of computing devices 54A-54N shown in FIG. 1 are merely exemplary, and that cloud computing nodes 10 and cloud computing environment 50 can communicate with any type of computerized device via any type of network or network addressable connection or both (e.g., using a web browser).

[0037] Next, referring to FIG. 2, a set of functional abstractions provided by cloud computing environment 50 (FIG. 1) is shown. It should be understood in advance that the components, layers, and functions shown in FIG. 2 are merely exemplary and that embodiments of the present invention are not limited thereto. As illustrated, the following layers and corresponding functions are provided.

[0038] The hardware and software layer 60 includes hardware components and software components. Examples of hardware components include mainframe 61, RISC (Reduced Instruction Set Computer) architecture-based server 62, server 63, blade server 64, storage device 65, and network and networking components 66. Depending on the embodiment, software components may include network application server software 67 and database software 68.

[0039] The virtualization layer 70 provides an abstraction layer that can include the following examples of virtual entities: virtual server 71, virtual storage 72, virtual network 73 including a virtual private network, virtual applications and operating systems 74, and virtual client 75.

[0040] In one example, the management layer 80 can provide the following functions. Resource provisioning 81 dynamically procures computing resources and other resources used to execute tasks within a cloud computing environment. Metering and pricing 82 tracks costs when resources are utilized within a cloud computing environment and sends invoices or bills for the consumption of these resources. In one example, these resources can include application software licenses. Security performs identity verification of cloud consumers and tasks, as well as protection of data and other resources. User portal 83 provides access rights to the cloud computing environment for consumers and system administrators. Service level management 84 performs cloud computing resource allocation and management such that the required service levels are met. Service Quality Assurance Agreement (SLA) planning and fulfillment 85 performs advance placement and procurement of cloud computing resources where future requirements are expected to comply with the SLA.

[0041] The workload layer 90 provides examples of functions that can utilize a cloud computing environment. Examples of workloads and functions that can be provided from this layer include mapping and navigation 91, software development and lifecycle management 92, virtual classroom education delivery 93, data analysis processing 94, transaction processing 95, and container image creation and deployment 96.

[0042] Referring now to Figure 3, a schematic diagram of an example of a cloud computing node is shown. The cloud computing node 10 is only one example of a suitable cloud computing node and is not intended to suggest any limitation as to the use or functionality scope of the embodiments of the present invention described herein. Nevertheless, the cloud computing node 10 can be implemented, or can execute any of the functions described above, or both.

[0043] The cloud computing node 10 includes a computer system / server 12 that operates in a number of other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, or configurations, or combinations thereof, that may be suitable for use in the computer system / server 12 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments including any of the above systems or devices.

[0044] The computer system / server 12 can be described in the general context of computer system-executable instructions, such as program modules executed by a computer system. Generally, program modules can include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. The computer system / server 12 can be practiced in a distributed cloud computing environment where tasks are performed by remote processing devices linked through a communications network. In a distributed cloud computing environment, program modules can be located in both local and remote computer system storage media, including a memory storage device.

[0045] As shown in FIG. 3, the computer system / server 12 in the cloud computing node 10 is shown in the form of a general-purpose computing device. The components of the computer system / server 12 can include, without limitation, one or more processors or processing units 16, a system memory 28, and a bus 18 that couples various system components including the system memory 28 to the processor 16.

[0046] The bus 18 represents any one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor bus or local bus using any of a variety of bus architectures. By way of example and not limitation, such architectures can include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Extended ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0047] Computer system / server 12 generally includes various computer system-readable media. Such media can be any available media accessible by computer system / server 12, including both volatile and non-volatile media, and removable and non-removable media.

[0048] System memory 28 can include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 30 or cache memory 32, or both. Computer system / server 12 can further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 can be provided for reading and writing to and from a non-removable non-volatile magnetic medium (not shown, commonly referred to as a "hard drive"). Although not shown, a magnetic disk drive for reading and writing to and from a removable non-volatile magnetic disk (e.g., a "floppy (R) disk"), and an optical disk drive for reading and writing to and from a removable non-volatile optical disk such as a CD-ROM, DVD-ROM, or other optical media can be provided. In such examples, each can be connected to bus 18 by one or more data media interfaces. As further illustrated and described below, memory 28 can include at least one program product having a set of program modules (e.g., at least one) configured to execute the functions of embodiments of the present invention.

[0049] A program / utility 40 having a set (at least one) of program modules 42, as well as an operating system, one or more application programs, other program modules, and program data can be stored, by way of example and not limitation, in a memory 28. Each of the operating system, one or more application programs, other program modules, and program data, or combinations thereof, can include an implementation of a networking environment. The program modules 42 generally execute the functions or methods or both of embodiments of the present invention as described herein.

[0050] Computer system / server 12 can also communicate with one or more external devices 14 such as a keyboard, a pointing device, a display 24, etc., one or more devices that enable a user to interact with computer system / server 12, or any device that enables computer system / server 12 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.), or a combination thereof. Such communication can be performed via input / output (I / O) interface 22. Furthermore, computer system / server 12 can communicate with one or more networks such as a local area network (LAN), a general wide area network (WAN), or a public network (e.g., the Internet), or a combination thereof, via network adapter 20. As shown, network adapter 20 communicates with other components of computer system / server 12 via bus 18. Although not shown, it should be understood that other hardware components or software components or both may be used with computer system / server 12. Examples include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archive storage systems, etc.

[0051] Next, referring to FIG. 4, a flowchart of method 400 is shown according to one embodiment. Method 400 can be implemented in various embodiments, in accordance with the present invention, in any of the environments shown in FIGS. 1-3 and FIG. 5. Of course, as will be understood by those skilled in the art upon reading this specification, method 400 may include more or fewer operations than those explicitly shown in FIG. 4.

[0052] Each of the steps of method 400 can be performed by any suitable component of the operating environment. For example, in various embodiments, method 400 can be performed, in part or in whole, by one or more servers, computers, or some other device having one or more processors. A processor, implemented, for example, in hardware or software or both, preferably a processing circuit, chip, or module having at least one hardware component, or a combination thereof, may be utilized in any device to perform one or more steps of method 400. Exemplary processors include, but are not limited to, a central processing unit (CPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), combinations thereof, or other suitable computing devices known in the art.

[0053] As shown in FIG. 4, method 400 can begin at operation 402, where a manifest of a container image for a container to be created is received. In one embodiment, the manifest can be obtained in response to a request to create a container within a cluster of computing nodes.

[0054] For example, the request can be received from a user, an application, etc. In another example, the request can be received in a cluster of computing nodes. In yet another example, the cluster of computing nodes can include a distributed computing network, a cloud-based computing environment, etc. In yet another example, a container can include a self-contained software package that implements operating system (OS) level virtualization.

[0055] Additionally, in one embodiment, creating a container can include mounting the container's file system on one of a plurality of computing nodes in a cluster and using the mounted file system or the like to load one or more libraries within the container, or execute one or more applications within the container, or a combination thereof.

[0056] Furthermore, in one embodiment, a container image can include all the files required to create a container on one of a cluster of computing nodes. In another embodiment, a container image can include multiple files (e.g., an executable package including code, runtime, system tools, system libraries, and settings, etc.). In yet another embodiment, a manifest of a container image can include metadata that describes the multiple files within the container image.

[0057] For example, the manifest can include file names (e.g., content identifiers), a list of ownership or permission data or both related to those files, etc. In another example, the manifest can include a content - based address of the file (e.g., a file hash, a pointer to where multiple files are stored, etc.). In yet another example, the manifest can store multiple file stubs each containing a pointer to where the file is stored.

[0058] Still further, in one embodiment, the manifest can include metadata that describes the multiple files within the container image but does not include the multiple files themselves. For example, the files can be remotely stored in a content store (e.g., a centralized object storage). In another example, one or more of the files can be stored locally (e.g., on the node where the container is to be created, etc.).

[0059] In another embodiment, after the container is created, the manifest retrieves individual files from the container image as needed. For example, it can be used by the container to retrieve them sequentially in a lazy manner when requested by an application running, for example, inside the container.

[0060] Furthermore, in one embodiment, the manifest may be retrieved by a scheduler module that is separated from the nodes in the cluster. In another embodiment, the manifest may be retrieved from a repository (e.g., a registry, etc.). For example, a scheduler module that is separated from the nodes in the cluster can retrieve the manifest of the container image of the container from a manifest repository (e.g., a database, etc.) that is physically separated from the scheduler module. In yet another embodiment, the manifest may be received from the scheduler by an overlap calculation module.

[0061] In addition, method 400 can proceed to operation 404, where a mapping index of a cluster of computing nodes is identified. In one embodiment, the mapping index can be identified in a repository. In another embodiment, the mapping index can be stored locally in the scheduler module. In one embodiment, the mapping index can be stored remotely (e.g., as a separate database, etc.). In yet another embodiment, the mapping index can store the identifiers of each node in the cluster of computing nodes.

[0062] Furthermore, in one embodiment, the mapping index can store identifiers (e.g., content identifiers, etc.) of all container image files currently stored within each node of a cluster of computing nodes. For example, each node within a cluster of computing nodes can store one or more portions of one or more container images. In another example, these portions include container image files. In yet another example, these portions can be stored within the cache of each node. In still another example, each node within a cluster of computing nodes can store different portions of one or more container images (e.g., different container image files, etc.) when compared to other nodes of the cluster.

[0063] Furthermore still, in one embodiment, the content identifier of each container image file currently stored within a node of a cluster of computing nodes can be linked (e.g., mapped, etc.) to an identifier of the node (e.g., node identifier, etc.) within the mapping index. In this way, the mapping index can store an index of all container image files stored within each node of the cluster.

[0064] Furthermore, in one embodiment, the mapping index can store only a portion (e.g., a prefix) of each content identifier in order to reduce the amount of stored data within the mapping index. In another embodiment, the mapping index may be obtained from a repository (e.g., a registry, etc.). For example, the mapping index may be obtained from the same repository from which the manifest is obtained. In yet another example, the mapping index may be received from a scheduler by an overlap calculation module.

[0065] Additionally, method 400 can proceed to operation 406, where a computing node is selected within a cluster of computing nodes based on a comparison of a manifest and a mapping index in order to create a container. In one embodiment, an overlap calculation module can compare a list of content identifiers within the manifest to content identifiers linked to node identifiers within the mapping index.

[0066] Furthermore, in one embodiment, for each node identifier within the mapping index, an overlap calculation module can determine the number of content identifiers within the manifest that are linked to the node identifier. For example, the number of content identifiers within the manifest linked to a node identifier can be totaled to create a score for the node identifier. In this way, the overlap calculation module can determine the number of container image files of a container image currently stored within each of the nodes of the cluster.

[0067] Still further, in one embodiment, a node linked to the largest number of content identifiers within the manifest (e.g., when compared to other nodes within the cluster) can be selected to create a container. For example, a node identifier having the highest score can be identified and returned. In another embodiment, a subset of node identifiers within the mapping index can be identified and compared to the manifest.

[0068] For example, only node identifiers of nodes that satisfy one or more additional resource requirements can be compared to the manifest. In another embodiment, the additional resource requirements can include a minimum amount of available cache memory of the node, a minimum amount of non-volatile storage of the node, and the like. For example, only node identifiers associated with nodes having a current available cache memory amount above a predetermined threshold can be compared to the manifest.

[0069] Furthermore, in one embodiment, a subset of all content identifiers in the manifest can be identified and compared to a mapping index. For example, each content identifier in the manifest can be assigned a weight value based on the amount of access to the history of the file represented by the content identifier. In another example, only content identifiers with a weight above a predetermined threshold may be compared to the mapping index. In this way, the node selection time can be reduced by limiting the number of node identifiers and content identifiers to be compared.

[0070] As a result, the node that currently stores the maximum number of container image files for the container's container image can be selected to create the container. This can minimize the amount of container image files that need to be obtained by the node during creation / execution of the container on the node, and as a result, reduce the amount of bandwidth used to transfer such container image files. This reduces the amount of network traffic between the node and the container image store, thereby improving the performance of one or more hardware components that perform such network communication. This can also reduce the waiting time for the node to obtain the container image files (by maximizing the number of local container image files), which can improve the performance of the node's computing hardware during container implementation.

[0071] In addition, in one embodiment, the creation of the container can be scheduled on the selected node. In another embodiment, the container's file system can be mounted on the selected node using the manifest. For example, the manifest can contain sufficient data to create (e.g., mount) the container's file system. In another example, the manifest can contain one or more inode descriptors and file hashes.

[0072] Further, in one example, the i-node descriptor can include metadata used to mount the container's file system. In another example, the file system can be mounted to a node of a cluster of computing nodes (e.g., a node to which a task of creating a container is assigned, etc.).

[0073] Still further, in one embodiment, the file system mounted for a container can identify a request to access data within the container's container image. For example, the request to access data can include a file read request. In another example, the request to access data can include a request from an application within the container to read data within the container image. In yet another example, the data can include a container image file.

[0074] Further, in one embodiment, the location of the data can be determined using a manifest. For example, the location of the data can be included within metadata stored within the manifest. For example, the metadata can describe multiple files within the container image. In another example, the manifest can include a content-based address of the file (e.g., a file hash, a pointer to a location where multiple files are stored, etc.).

[0075] Additionally, in one embodiment, data can be obtained using the location of the data. For example, the data can be obtained locally from the cache of a node, or remotely from a content store / repository / registry. For example, the cache can include high-speed low-latency memory that is faster than standard data storage within the nodes of the cluster. In another embodiment, the cache can include volatile memory. In yet another embodiment, the content store may be physically separated from the nodes of the cluster and may be accessed via a communication network. In yet another embodiment, the content store can further store the initially received manifest.

[0076] Furthermore, in one embodiment, in response to a determination by the manifest of an image that the data is stored locally in the cache of a node, the data can be obtained from the cache. In another embodiment, in response to a determination by the manifest of an image that the data is not stored locally in the cache, the data can be obtained from the content store using the communication network.

[0077] Still further, in one embodiment, the content store can store data related to the container image. In another embodiment, the obtained data can be used in a mounted file system. For example, the obtained data can be presented to an application running within the node using the mounted file system.

[0078] Furthermore, in one embodiment, the mapping index can be updated in response to the acquisition of a container image file and its storage on a node. For example, the content identifier of the acquired container image file can be linked / mapped to the identifier of the node within the mapping index. In another embodiment, the mapping index can be updated in response to the deletion of a container image file from a node. For example, a node can be able to eliminate a cached container image file after a predetermined time threshold has been exceeded.

[0079] In this way, the mapping index can be updated to accurately indicate all the container image files currently stored on all the nodes within the cluster.

[0080] Content-Aware Container Scheduling In one embodiment, containers are scheduled such that the overlap of the existing content of the target node is maximized. The mapping of the content is done at the container host. The optimal target host is calculated based on the overlap and additional constraints. The mapping is updated online when content is acquired / evicted from the host.

[0081] Additionally, as the amount of content that needs to be acquired decreases, the required network traffic decreases (e.g., some policies often re-pull the image every time a container is run (rerun)). When existing content can be reused and does not need to be stored redundantly, the required storage and memory decrease. When cached content can be reused, the required input / output (I / O) bandwidth decreases. Overall, by maximizing data reuse, a higher container density can be achieved.

[0082] Figure 5 shows an exemplary system architecture 500 according to one embodiment. As shown, the scheduler 502 receives a request to execute a container. In response to receiving the request, the scheduler 502 retrieves the container manifest 504 from the registry 506.

[0083] Next, the scheduler 502 sends the retrieved manifest 504 to the overlap calculator 508. In one embodiment, the scheduler 502 and the overlap calculator 508 may be located within a single computing system. In another embodiment, the scheduler 502 and the overlap calculator 508 may be located within separate computing systems.

[0084] In response to receiving the retrieved manifest 504, the overlap calculator 508 identifies the mapping index 510, compares the retrieved manifest 504 with the mapping index 510, and calculates the maximum overlap between the content identifiers within the mapping index 510 and the content identifiers within the retrieved manifest 504. For example, a node 514 within the mapping index 510 that is linked to the largest number of content identifiers found within the retrieved manifest 504 (e.g., when compared to other nodes within the cluster 512) can be selected to create the container.

[0085] In response to the identification of the node 514 with the maximum overlap, the node 514 is scheduled by the scheduler 502 to execute the container. The retrieved manifest 504 can be sent from the scheduler 502 to the node 514 and used to mount the container's file system within the node 514.

[0086] After the file system is mounted within node 514, additional container image data not found in the manifest may be required by the mounted file system. In response to this requirement, node 514 can obtain the requested content 516 from registry 506 on an on-demand basis. Node 514 can then send an index of this obtained content 516 to mapping index 510, and mapping index 510 can be updated to reflect this newly obtained content 516 within node 514. Mapping index 510 is also updated depending on node 514 having similarly deleted one or more instances of content, such that mapping index 510 can include an up-to-date snapshot of all container data stored at each node of cluster 512.

[0087] Identification of Target Node In one embodiment, the mapping index can be stored. For example, the mapping index can include N sub-indexes, one for each node within the cluster. Each index can include a hash set containing the content IDs of the files of that node. A lookup for a certain period of time can be performed for each index. The size of the index can be reduced by storing only the prefix of each content ID.

[0088] Calculation of Target In one embodiment, a node with the highest overlap with the manifest of existing files and that also meets other scheduling requirements (e.g., number of nodes, amount of available RAM, etc.) can be identified. Exemplary steps are as follows.

[0089] 1. For each node of the mapping index, investigate how many image files already exist at that node (e.g., O(N*M)). Here, N is the number of nodes and M is the number of files within the manifest.

[0090] 2. Rearrange the nodes according to the overlap size.

[0091] 3. Further, schedule the containers at the nodes with the highest overlap size that satisfy all other scheduling requirements.

[0092] Reduction of scheduling decision time In one embodiment, additional indexes can be added from the image to the nodes. Each entry can map an image to a list of nodes. For each node, the started image and the start time can be tracked. The node that last started the image can be selected to start a new instance of that image.

[0093] In another embodiment, the nodes can first be filtered according to other resource requirements. For example, only the nodes that satisfy other resource requirements can be checked for overlap (e.g., O(N*M) is reduced to O(N_R*M), where N_R = the set of remaining nodes that satisfy the resource requirements).

[0094] In yet another embodiment, the overlap can be calculated between the most frequently used subsets of the image. For example, each file entry of the image can be labeled with a priority based on its access pattern (e.g., the more frequently accessed a file is, the higher its priority). When calculating the overlap, only files with a priority above a threshold may be considered (e.g., O(N_R*M) -> O(N_R*M_p), where M_p = the number of files with priority > p).

[0095] In yet another embodiment, an inverse index of files can be used to list the nodes containing the files. For example, for each file i of the images, a list of nodes (N_Fi) containing the file can be identified. The list can be traversed to find the node with the maximum occurrence as the scheduling target (e.g., here, O(N*M) -> O(M*N_Fi+N), where in the worst case O(M*N+N)).

[0096] Fault tolerance In one embodiment, in response to a crash of an index node, the index can be periodically flushed to disk and restored upon restart. Since a slight inconsistency between the index and the actual cluster state only affects the quality of scheduling decisions and not the operation itself, consistency may not be important. Journaling, transactions, etc. may not be required. If the index is completely lost, the index can be reconstructed from the worker nodes.

[0097] In another embodiment, a worker node may crash. The scheduler can automatically stop scheduling to that node, but can retain the index of the node. Upon node recovery, the file system can send a local index checksum to the scheduler. The scheduler can then calculate the checksum of its own index entry. If the checksums are different, the scheduler can request the entire node index from the node.

[0098] In one embodiment, a method is provided for providing content-aware scheduling of container images within a containerized cluster. Additionally, a method is provided for maintaining a mapping index for tracking the location of files within the cluster. Further, a method is provided for using the mapping index to calculate an optimal scheduling target node that maximizes the sharing of existing data for new containers that need to be run on the cluster. Still further, a method is provided for reducing the scheduling decision time by limiting the number of nodes and files to check in the mapping index.

[0099] In one embodiment, maintaining a mapping between a cluster node and the image content it caches, identified by a content ID, updating the mapping whenever the cluster node acquires or deletes image data, downloading an image manifest that lists the content IDs of the content within the image as soon as a request to schedule a container is received, querying the node against the content mapping to compare the content IDs of the image manifest to the content IDs of the data stored on the cluster node, and scheduling the container on the node with the largest overlap between the cached content and the container image data. A method for performing content-aware container scheduling is provided.

[0100] The present invention can be a system, method, or computer program product, or a combination thereof, that can be integrated at a technical detail level. The computer program product can include one or more computer-readable storage media having computer-readable program instructions for causing a processor to execute embodiments of the present invention.

[0101] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy (R) disk, punch cards, or mechanically encoded devices such as raised structures within grooves in which instructions are recorded, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed to be a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.

[0102] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to respective computing / processing devices, or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface of each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on a computer-readable storage medium within each computing / processing device.

[0103] The computer-readable program instructions for carrying out the operations of the present invention may be source code or object code written in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or object-oriented programming languages such as Smalltalk(R), C++, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit for performing the embodiments of the present invention.

[0104] Embodiments of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0105] These computer-readable program instructions can be provided to a processor of a computer or other programmable data processing apparatus to create means for implementing the functions / operations specified in one or more blocks of a flowchart, a block diagram, or both, via the processor of the computer or other programmable data processing apparatus. These computer-readable program instructions may also be stored in a computer-readable storage medium that constitutes a product including instructions for implementing embodiments of the functions / operations specified in one or more blocks of a flowchart, a block diagram, or both, to instruct a computer, a programmable data processing apparatus, or other devices, or combinations thereof, to function in a particular manner.

[0106] The computer-readable program instructions may also be loaded onto a computer, other programmable apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices to create a computer-implemented process for implementing the functions / operations specified in one or more blocks of a flowchart, a block diagram, or both.

[0107] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible embodiments of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, segment, or portion of instructions that include one or more executable instructions for implementing the specified logical function. In some alternative embodiments, the functions noted in the blocks may be performed out of the order noted in the figures. For example, two blocks shown in succession may in fact be accomplished as one step, executed simultaneously, substantially simultaneously, in a partially or wholly temporally overlapping manner, or the blocks may sometimes be executed in the reverse order depending on the functionality involved. It should also be noted that each block of the block diagram or flowchart, or both, and combinations of blocks of the block diagram or flowchart, or both, can be implemented by a dedicated hardware-based system that performs the specified functions or operations or by a combination of dedicated hardware instructions and computer instructions.

[0108] In addition, a system according to various embodiments can include a processor and logic that is integrated with the processor, executable by the processor, or both, the logic being configured to execute one or more of the process steps described herein. What "integrated" means is that the processor has logic embedded therein as hardware logic such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or the like. What "executable by the processor" means is hardware logic, software logic such as firmware, a part of an operating system, a part of an application program, or the like, or a combination of hardware logic and software logic, that is accessible by the processor and configured to cause the processor to perform some functions when executed by the processor. Software logic can be stored in any type of local memory or remote memory or both, as is known in the art. Any processor known in the art such as a software processor module, or a hardware processor such as an ASIC, an FPGA, a central processing unit (CPU), an integrated circuit (IC), a graphics processing unit (GPU), or both, can be used.

[0109] It will be apparent that the various features of the aforementioned system or method or both can be combined in any manner to create a plurality of combinations based on the description presented above.

[0110] It will be further understood that embodiments of the present invention are provided in the form of services deployed for customers and can provide services on demand.

[0111] The descriptions of the various embodiments of the present invention are presented for illustrative purposes and are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein are chosen to best explain the principles of the embodiments, the practical application, or the technological improvements seen in the marketplace, or to enable those skilled in the art to understand the embodiments disclosed herein.

Description of Reference Numerals

[0112] 10 Cloud computing node 12 Computer system / server 14 External device 16 Processor or processing unit 18 Bus 20 Network adapter 22 Input / output (I / O) interface 24 Display 28 System memory 30 Random access memory (RAM) 32 Cache memory 34 Storage system 40 Program / utility 42 Program module 50 Cloud computing environment 54A Personal digital assistant (PDA) or mobile phone 54B Desktop computer 54C Laptop computer 54N Automotive computer system 60 Hardware and software layer 61 Mainframe 62 RISC (Reduced Instruction Set Computer) architecture-based server 63 Server 64 Blade server 65 Storage device 66 Networks and Networking Components 67 Network Application Server Software 68 Database Software 70 Virtualization Layer 71 Virtual Servers 72 Virtual Storage 73 Virtual Networks 74 Virtual Applications and Operating Systems 75 Virtual Clients 80 Management Layer 81 Resource Provisioning 82 Measurement and Pricing 83 User Portal 84 Service Level Management 85 Service Quality Assurance Contract (SLA) Planning and Fulfillment 90 Workload Layer 91 Mapping and Navigation 92 Software Development and Lifecycle Management 93 Virtual Classroom Education Delivery 94 Data Analysis Processing 95 Transaction Processing 96 Container Image Creation and Deployment 402 Operations 500 System Architecture 502 Scheduler 504 Manifest 506 Registry 508 Overlap Calculator 510 Mapping Index 512 Cluster 514 Node 516 Content

Claims

1. A method for information processing by a computer, comprising: Receiving a manifest of a container image of a container to be created; Identifying a mapping index of a cluster of computing nodes; Selecting a computing node within the cluster of computing nodes based on a comparison of the manifest and the mapping index for creating the container, wherein the manifest of the container image includes metadata describing a plurality of files within the container image, the metadata includes a list of content identifiers and content base addresses, and the mapping index stores content identifiers of all container image files currently stored within each node of the cluster of computing nodes, and selecting the computing node; A method comprising the above steps.

2. The method according to claim 1, wherein the content identifier of each container image file currently stored within a node of the cluster of computing nodes is linked to an identifier of the node.

3. The method according to claim 1, wherein the mapping index stores only a prefix of each content identifier of all container image files currently stored within each node of the cluster of computing nodes.

4. The method according to claim 1, wherein for each node identifier within the mapping index, the number of content identifiers within the manifest linked to the node identifier is determined.

5. The method according to claim 1, wherein the node linked to the maximum number of content identifiers within the manifest is selected for creating the container.

6. The method according to claim 1, further comprising identifying a subset of node identifiers within the mapping index for comparison with the manifest.

7. The method according to claim 1, further comprising identifying a subset of all content identifiers within the manifest for comparison with the mapping index.

8. The method according to claim 1, further comprising obtaining a container image file and updating the mapping index in accordance with storing the container image file in a computing node of the cluster.

9. The method according to claim 1, further comprising updating the mapping index in accordance with deleting a container image file from a computing node of the cluster.

10. The method according to claim 1, further comprising scheduling creation of the container at the selected computing node.

11. The method according to claim 10, further comprising mounting a file system of the container at the selected computing node using the manifest.

12. identifying a request to access data in a container image of the container by the file system mounted for the container; obtaining the data using a location of the data determined using the manifest; The method according to claim 11, further comprising.

13. A computer program that causes a computer to execute the method according to any one of claims 1 to 12.

14. A storage medium storing the computer program according to claim 13 in a computer-readable storage medium.

15. A computer system that executes the method according to any one of claims 1 to 12 by computer hardware.

Citation Information

Patent Citations

  • Representation of digital content metadata

    JP2009541839A

  • Global asset management

    JP2009544070A

  • System and method for reconstructable all-in-one content stream

    JP2016045944A

  • Volume arrangement management apparatus, volume arrangement management method and volume arrangement management program

    JP2020038421A

  • Software container registry container image deployment

    US20170177860A1