Systems and methods to dynamically isolate CPUS assigned to critical applications
The dynamic isolation of CPUs for critical applications addresses the limitations of static CPU allocation by allowing on-demand creation and release of isolated CPU cores, enhancing resource utilization and performance.
Patent Information
- Application Number
- PCT/IB2023/062332
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-06
- Publication Date
- 2025-06-12
AI Technical Summary
Existing systems face challenges in dynamically isolating CPUs assigned to critical applications, leading to underutilization of resources and performance degradation due to static pre-allocation of isolated CPUs and limitations in Cgroups management.
A method and system for dynamically isolating CPUs by receiving a request for CPU isolation, selecting candidate CPUs based on system administrator policies and application requirements, and controlling kernel/OS Cgroup parameters to restrict access to the isolated CPUs, allowing for on-demand creation and release of isolated CPU cores.
This approach eliminates the need for static CPU reservation, reduces underutilization, and provides isolation comparable to boot-time isolated CPUs, while allowing for seamless integration with existing systems and heterogeneous cloud infrastructure.
Smart Images

Figure IB2023062332_12062025_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS TO DYNAMICALLY ISOLATECPUS ASSIGNED TO CRITICAL APPLICATIONSTECHNICAL EIELD
[0001] Disclosed are embodiments related to systems and methods to dynamically isolate central processing units (CPUs) assigned to critical applications.BACKGROUND
[0002] Companies, such as telecommunication companies and others, are looking at possibilities of maximizing the performance of their critical applications while trying to ensure deterministic behavior for their critical applications.
[0003] Under normal circumstances, the kernel task scheduler treats all CPUs as available for scheduling process threads and preempts executing process threads thereby giving CPU time to other applications. The positive side-effect of this behavior is multitasking enablement and more efficient CPU resource utilization. The negative side-effect is non- deterministic performance, which makes it unsuitable for latency-sensitive workloads.
[0004] PIG. 1 illustrates an existing system where critical applications 106, best-effort applications 104, and operating system processes (such as kernel threads, interrupt requests (IRQs), timer ticks, timers, and Read Copy Update (RCU) callbacks) all share and compete for CPU 108 resources within server 102.
[0005] A prevalent solution for optimizing these workloads’ performance is to “isolate” a CPU core, or a set of CPU cores, from the kernel scheduler, such that it will not schedule a process thread on the isolated CPU cores. Then, latency-sensitive workload process threads can be pinned to execute on isolated CPU cores, providing them exclusive access to CPU cores. This results in more deterministic behavior due to reduced or eliminated thread preemption (context switching) and maximizes a CPU core’s cache utilization for the critical application.
[0006] Boot-time Isolated CPUs, commonly known as “Isolcpus”, is the solution predominately adopted to isolate CPU cores to run critical applications. Lor instance, in the telecommunication domain, some of the critical Radio Access Network (RAN) workloads may be provisioned to run on Isolcpus. CPU intensive workloads and low-latency network intensive(e.g., Data Plane Development Kit (DPDK)-based) workloads are good candidates for running on Isolcpus.
[0007] Isolcpus provides not only task-scheduling contention-free CPUs, but also migrates system IRQs (interrupts tasks), system timers and other housekeeping tasks such as Read-Copy-Update (RCU), kernel threads away from the isolated CPU cores.
[0008] FIG. 2 illustrates a performance comparison of a test application running on regular CPUs and on boot isolated CPUs. The test application is a CPU intensive benchmark application. The results of our experiments show that boot-isolated CPU cores can potentially enable deterministic performance for critical applications, as OS jitter can somehow be kept minimal on those CPU cores. This test application periodically reports OS jitter (Y-axis), which is a measure of time corresponding to the CPU cycles missed by a critical application running on a specific CPU core, as the missing CPU cycles would have been instead used by other co-hosted applications or kernel processes.
[0009] Another important feature used in embodiments disclosed herein is control groups (Cgroups). In Linux, Cgroups are a kernel feature that allows an administrator to allocate resources such as CPU, memory, and I / O bandwidth to groups of processes. Cgroups provide a way to control how much of the system’s resources a process or a group of processes can use. Cgroups may also be available in Windows or other operating systems.
[0010] Cgroups are organized in a hierarchy, with each Cgroup having a parent Cgroup. The child Cgroup inherits the resources of its parent Cgroup. This allows for fine-grained control over resources, as you can set limits on individual Cgroups or on entire Cgroup hierarchies. Cgroups help with better resource management, process isolation, and containerization.SUMMARY
[0011] Prior work tries to utilize boot-time allocated isolated CPUs (isolcpus) and statically provision critical workloads on the isolated CPUs. Even though Isolcpus provide the benefit of CPU isolation, there are some challenges associated with utilizing Isolcpus.
[0012] Isolcpus are statically pre-allocated at boot-time. Cloud operators / system administrators must specify the list of CPU cores to be isolated at system boot-time. The list ofCPUs to be isolated from kernel scheduling are specified in the GRUB using the following parameters:• “isolcpus=<cpu_list>”List of CPUs excluded from the general kernel symmetrical multiprocessing (SMP) balancing and scheduler algorithms.• “nohz_full=<cpu_list>”Kernel will stop sending timer ticks to those CPUs. CPUs will be made dynticks (i.e., supporting dynamic ticks, e.g., as specified by the Linux kernel).• “rcu_nocbs=<cpu_list>”Kernel will offload Read-Copy-Update (RCU) callbacks from the specified CPUs list.
[0013] The operating system scheduler will not schedule any tasks on isolated cores. Isolated cores remain idle, until the user explicitly binds (e.g., critical) tasks onto the isolated cores. Also, the list of isolated cores cannot be dynamically increased or decreased at runtime. These challenges can lead to underutilization of CPUs and can lead to resource imbalance within the server.
[0014] Using Isolcpus in Cloud environment is challenging. Isolcpus are not prioritized and well-utilized in cloud environments. For instance, vanilla Kubernetes does not prioritize and allocate isolated CPUs for critical workloads. Frameworks like Intel CMK and Nokia CPU Pooler in conjunction with Kubernetes CPU Manager can provide a workaround solution for utilizing isolated CPUs. Isolated CPUs are pre -reserved for critical workloads. At the time of scheduling, users must influence the Kubernetes scheduler to schedule the critical workloads on the server nodes with Isolcpus using labels and affinities. Finally, Intel CMK can allocate and schedule critical workload on Isolcpus.
[0015] This workaround is a static approach and can fail when nodes run out of Isolcpus. This static approach does not fit well with cloud-native approach. Reserving isolated CPUs and not utilizing until critical workloads are scheduled, can lead to under-utilization at server-level and at cluster-level.
[0016] There are also challenges with respect to non-uniform memory access (NUMA) and NUMA Input Output (IO). Critical workloads requiring CPU isolation can require the usage of hardware accelerators like SmartNICs, Graphical Processor Units (GPUs), and so on, present on the system. On a multi-socket or multi-NUMA system, it is possible that hardware accelerators and isolcpus can be allocated and available on different sockets / NUMA. This can lead to performance degradation of critical workloads.
[0017] Most importantly, Linux has deprecated the support for isolcpus. Hence it is important to find an alternative, possibly a dynamic approach to create isolcpus on-demand.
[0018] Challenges also exist with Cgroups management tools in achieving dynamic Isolation. Most available tools do static partitioning of Cgroups. For example, Cset partitions the system in two broader Cgroups - user and system Cgroup. It tries to shield processes only from operating system tasks and does not provide isolation between the tasks running in the user Cgroups. Also, Cset must be applied right after rebooting the system. Cset will not work well with system where multiple Cgroups are already active. Importantly, Cset does not migrate RCU callbacks and IRQs and doesn’t make the CPUs adaptive-ticks (NOHZ_FULL) on-demand.
[0019] Advantages of proposed embodiments include the following. Embodiments eliminate the need for reserving isolated CPUs on a set of nodes for the critical applications. Unlike the boot-time isolated CPUs where the list of isolated CPUs is fixed at boot-time, Embodiments can dynamically create new isolated CPUs on-demand. Dynamically created Isolcpus can be released back to the schedulable shared pool CPUs. This can reduce the underutilization of CPUs when critical workloads are not running. Results have shown that embodiments provide at least as good isolation as boot-time Isolcpus and can provide similar application performance on par with boot-time isolated CPUs (based on experimentation results). No changes are required to the applications, system components, and operating system. Embodiments can be seamlessly integrated into the existing systems. Embodiments can be deployed on heterogeneous cloud infrastructure, e.g., supporting different CPU architectures and accelerators. Embodiments provide an extra layer of abstraction for specifying how critical applications could request deterministic performance.
[0020] According to a first aspect, a method for providing isolation from an operating system (OS) scheduler is provided. The method includes receiving a request to isolate one or more CPUs for an application. The method includes selecting a set of CPUs to isolate based on the request. The method includes after boot-time, restricting tasks not associated with the application from accessing the set of CPUs to isolate.
[0021] In some embodiments, the request to isolate one or more CPUs for an application is received on-demand at run-time. In some embodiments, selecting a set of CPUs to isolate based on the request comprises selecting one or more exclusive CPU cores based on system administrator policies, information regarding reserved CPUs, and information regarding Non- Uniform Memory Access (NUMA) and NUMA Input-Output (IO) relationships. In some embodiments, the application has critical performance related requirements which must be met. In some embodiments, the tasks comprise one or more processes and / or threads. In some embodiments, restricting tasks not associated with the application from accessing the set of CPUs to isolate comprises restricting tasks in a control group not associated with the application from having access to the set of CPUs to isolate for one or more control groups. In some embodiments, restricting tasks in a control group not associated with the application from having access to the set of CPUs to isolate for one or more control groups comprises performing a depth first search (e.g., post-order) traversal of the one or more control groups and modifying CPU affinity for control groups not associated with the application to restrict those control groups from having access to the set of CPUs to isolate.
[0022] In some embodiments, restricting tasks not associated with the application from accessing the set of CPUs to isolate comprises offloading one or more interrupt requests from the set of CPUs to isolate to prevent the interrupt requests from being handled on the set of CPUs to isolate. In some embodiments, offloading one or more interrupt requests from the set of CPUs to isolate to prevent the interrupt requests from being handled on the set of CPUs to isolate comprises offloading all interrupt requests from being handled on the set of CPUs to isolate. In some embodiments, restricting tasks not associated with the application from accessing the set of CPUs to isolate comprises converting the set of CPUs to an adaptive tick mode. In some embodiments, the method further includes restricting tasks not associated with the application from accessing the set of CPUs to isolate comprises offloading read copy update (RCU)callbacks from the set of CPUs to isolate to alternative CPUs. In some embodiments, the method further includes taking the set of CPUs to isolate offline to force tasks running on the set of CPUs to migrate and then bringing the set of CPUs online. In some embodiments, the method further includes reverting the set of CPUs to isolate to a non-isolated state when the application is terminated or killed.
[0023] According to a second aspect, a server is provided. The server includes processing circuitry; and a memory. The memory contains instructions executable by the processing circuitry, whereby when executed the processing circuitry is configured to receive a request to isolate one or more CPUs for an application. The processing circuitry is configured to select a set of CPUs to isolate based on the request. The processing circuitry is configured to, after boot-time, restrict tasks not associated with the application from accessing the set of CPUs to isolate.
[0024] According to a third aspect, a computer program is provided, comprising instructions which when executed by the processing circuitry of a node cause the node to perform the method of any of the embodiments of the first aspect.
[0025] According to a fourth aspect, a carrier is provided, containing the computer program of the third aspect. The carrier is one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium.BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate various embodiments.
[0027] FIG. 1 illustrates an existing system where critical applications 106, best-effort applications 104, and operating system processes share and compete for resources.
[0028] FIG. 2 illustrates a performance comparison of a test application running on regular CPUs and on boot isolated CPUs.
[0029] FIG. 3 illustrates a server according to some embodiments.
[0030] FIG. 4 illustrates a server according to some embodiments.
[0031] FIG. 5 illustrates a control group hierarchy according to some embodiments.
[0032] FIG. 6 illustrates a control group hierarchy according to some embodiments.
[0033] FIG. 7 illustrates the performance comparison of a test application running on regular non-optimized CPUs and on the dynamically isolated CPUs from research experiments.
[0034] FIG. 8 illustrates a flowchart according to some embodiments.
[0035] FIG. 9 illustrates a high-level architecture of CPU Isolator 402 implemented as aCRI proxy according to some embodiments.
[0036] FIG. 10 illustrates a high-level architecture of CPU Isolator implementation as an NRI plugin inside container runtimes like Containerd or CRI-0 according to some embodiments.
[0037] FIG. 11 illustrates a flowchart according to some embodiments.
[0038] FIG. 12 is a block diagram of an apparatus according to some embodiments.DETAILED DESCRIPTION
[0039] Embodiments provide a system and a method to dynamically isolate CPUs, e.g., assigned to critical applications, at runtime, on-demand.
[0040] Embodiments include a new software module located on each server within a system, with access to the system’s CPU resources and control interfaces. Users may provide resource requirements for critical applications. Also, users may request isolated CPUs for critical workload, e.g., using intents. For example, a user can specify “CPU Isolation” as an intent in the application specification file for requesting isolated CPUs. Based on users’ specified requirements, the proposed module can optionally identify candidate CPUs (e.g., expressed as CPU cores) for running the critical workloads. The module may dynamically convert the candidate regular CPUs into isolated CPUs, such as by controlling the kernel / OS Cgroup parameters, migrating soft IRQs (interrupts) and timers away from the candidate CPUs, enabling adaptive-ticks (also referred to as dynticks, or nohz_full) on the CPUs, and offloading RCU callbacks from the candidate CPUs. Finally, the dynamically isolated CPUs may be exclusively allocated to the associated critical application. When the critical application runs to completion or is killed, the proposed module may reconvert the isolated CPUs to regular (non-optimized) CPUs and return the CPUs back to the shared schedulable pool of CPUs, where other best-effort applications of the system can utilize the CPUs.
[0041] Embodiments provide a system and a method to dynamically isolate the CPU cores used by critical applications from other tasks (e.g., processes and / or threads) running on the same server of a cloud environment. Embodiments are able to dynamically create the required number of isolated CPUs on-demand in a cloud environment, considering users’ application requirements and the system topology, either at schedule-time or deployment-time of the critical applications, or at runtime for already existing critical applications. A critical application has performance related requirements that must be met. For example, critical applications include applications with real-time requirements, or otherwise having a need to minimize jitter, maximize performance, and / or have a more deterministic performance.
[0042] Embodiments introduce a system component, referred to here as a “CPU Isolator”, for creating isolated CPUs dynamically by controlling host operating resource interfaces, with minimum impact to the operating system, cloud environment, or to the existing application’s architectural design.
[0043] Embodiments include an intent-driven framework, where the critical applications specify the need for CPU Isolation to get deterministic performance. Embodiments include identifying and / or allocating candidate CPUs based on application requirements, system administrator policies, and / or system topology. Embodiments include controlling Cgroups to prevent other tasks (e.g., processes and / or threads) from running on the isolated CPU cores. Embodiments include migrating interrupts to other CPU cores and preventing interrupt handlers from getting scheduled on the isolated cores by controlling the interrupts’ affinity. Embodiments include making the isolated CPU cores support adaptive-ticks for preventing periodic timer interrupts (ticks) from interrupting critical applications. Embodiments include offloading RCU callbacks kernel threads to other CPU cores. Embodiments include updating OS and kernel parameters, priority of the application, appropriate task scheduler for the application, and so on, on a per-application basis based on the system administrator policy.
[0044] FIG. 3 illustrates a system according to an embodiment. As shown, there is component, or module, referred to as “CPU Isolator” which is introduced into a server 302 to achieve dynamic CPU isolation. Achieving dynamic isolation requires modifying various parameters at user-space and in kernel-space. CPU Isolator will be realized as two components, one running at user-space (user space component 304) and another at kernel-space (kernel spacecomponent 306) as a loadable kernel module. In a cloud environment, CPU Isolator may be on each server 302 in the cluster that supports dynamic management of isolated CPUs. CPU Isolator may be running as a system daemon. Server 302 includes an OS Scheduler 308, CPUs 310, and memory 312. OS Scheduler 308 may schedule one or more best effort applications 104 and one or more critical applications 106. CPU Isolator ensures that the scheduling of these applications ensures that critical applications 106 are assigned to isolated CPUs.
[0045] FIG. 4 illustrates a high-level architecture of the proposed system according to an embodiment. FIG. 4 shows the inputs and control knobs that CPU Isolator 402 (including the user space component 304 and kernel space component 306) utilizes to dynamically create isolated CPUs.
[0046] CPU Isolator 402 may take various input parameters. For example, CPU Isolator 402 may take the following parameters as input in order to function efficiently.
[0047] Input parameters may include input related to system topology and available resources, such as available CPU resources within the system. If there are multiple entities allocating CPU resources to applications, the information about available CPUs may be periodically updated to CPU Isolator 402 to take informed decisions when allocating CPUs. The system topology may include a number of CPU sockets, NUMA nodes, boot isolated CPUs, CPUs capable of running at higher frequencies, and so on, some or all of which may be provided and / or may be obtained automatically via the existing (e.g., / sys) interfaces provided by the operating system. Also, CPU Isolator 402 may understand the existing Cgroup hierarchy, e.g., typically mounted under / sys / fs / cgroups location.
[0048] Input parameters may include input related to system administrator policies and requirements. In practice, it is common for system administrators to reserve certain CPUs for running housekeeping jobs or for running other workloads. These CPUs should be excluded as candidates for dynamic Isolcpus, and such exclusion can be specified by the input parameters. Socket / NUMA directly connected to hardware accelerator(s) can also be another possible input from a system administrator. Also, a system administrator may specify to enforce kernel and OS parameters when critical applications are provisioned on the system through CPU Isolator 402. For example, disabling NUMA balancing of memory pages, virtual memory usage statscollection frequency, setting Isolcpus to run at higher CPU frequency, disabling c-states on isolated CPUs, and so on may be specified via input parameters. An administrator may also provide hints about what priority a particular process should run at, task scheduling algorithm(s) to use (e.g., Completley Fair Scheduler (CFS), Realtime - SCHED_FIFO, SCHED_RR, or SCHED_DEADLINE, and so on). For example, Container Runtime Interface (CRI) Open Container Initiative (OCI) (CRI-O) container runtime provides enabling or disabling CFS on a container basis, or on a per-isolated core basis. Inputs may be considered when CPU Isolator 402 creates dynamic isolated CPUs.
[0049] Input parameters may include input related to application resource specifications. For example, a number of CPUs required by the application may be specified. Typically, all the cores requested by the application will be isolated from running other processes in the system. A possible variant could be, for example, an application requesting a subset of CPUs (from the total CPUs requested by the application) to be isolated and providing the details of the isolated CPUs to the application. Also, a user may explicitly specify if the application requires isolation, as part of the application specifications, e.g., if it cannot be somehow derived from other means.
[0050] Note: In scenarios where CPU Isolator 402 cannot read an application’s request directly, an agent that is responsible for handling a user request (e.g., Kubelet in Kubernetes) can signal CPU Isolator 402 with the required application information, if the application requests for CPU isolation. CPU Isolator 402 may also implicitly identify if an application requires isolated CPUs. This can be deduced, for example, based on profiling the application.
[0051] A method to create dynamic Isolcpus is provided. The following steps describe a method to dynamically select and create Isolcpus.
[0052] The method may include selection of candidate CPUs. An administrator may provision the CPU Isolator 402 (user-space 304 and kernel-space 306 components) on the interested nodes in the cluster. CPU Isolator 402 may automatically identify and map the system topology (e.g., layout of the CPUs). Based on the system administrator policies, CPUs that are reserved for other workloads may be excluded from allocation. CPU Isolator 402 may also apply the system parameters (OS and Kernel parameters) provided by the system administrator in the policy. When the application requests isolated CPUs, the request is forwarded to CPU Isolator402. CPU Isolator 402 may identify a number of CPUs to be allocated for the application and a number of CPUs to be isolated (these numbers may be the same, or one may be smaller than the other). In some embodiments, NUMA and NUMA-10 relationships are considered when identifying candidate CPUs for critical application. For example, allocating CPUs from the same NUMA / Socket where the SmartNIC and / or GPU is located based on an application’s requirements, or co-location with another peer-application on the same NUMA node when a shared-memory is used, and so on. During selection of candidate CPUs, CPU Isolator 402 may try to prioritize the selection of boot-time Isolated CPUs (if they are available). If the available number of boot-time isolated CPUs are less than the request number of isolated CPUs by the application, CPU Isolator 402 may dynamically create Isolated CPUs for the remaining count.
[0053] The method may include managing Cgroups (Domains). One possible way to control the operating system scheduler from scheduling processes on the candidate CPU cores is by controlling cpusets in Control Groups - Cgroups in Linux or Job objects in Windows (referred to herein generically by Cgroups). Typically, Cgroups are mounted in / sys / fs / cgroup in Linux. Each Cgroup (folder) can inherit CPUs (or a subset of CPUs) from the parent Cgroup. CPUs available to the Cgroup are controlled via a cpuset. cpus file present in each Cgroup. To control the processes (PIDs) which can be part of a particular Cgroup, tasks (for CgroupsV 1 ) or cgroup. procs (for CgroupsV2) files can be utilized. Processes belonging to a Cgroup have access only to the CPUs allocated to the Cgroup. This provides an opportunity to control the available CPUs for a particular Cgroup, by which CPUs can be isolated and made available only to certain processes - providing isolation from other processes. FIG. 5 illustrates an example of a Cgroups tree hierarchy in Linux.
[0054] Applications provisioned using cloud-native methods like Docker or Kubernetes follow a pre-defined hierarchy of folders under Cgroups. Each application has a dedicated Cgroup folder associated with it. For example, docker.slice, kubernetes. slice, and so on. Kubernetes and container runtimes take care of creating corresponding Cgroup folders when a Pod is provisioned on the node. If applications are provisioned directly, then CPU Isolator 402 may create a new folder for the application below the Cgroup’ s root folder in the hierarchy. The candidate CPUs allocated for the critical application may be removed from other Cgroups, that is, from those Cgroups not associated with the critical application. This will remove thepermission for other processes to run on the candidate cores. CPU Isolator 402 may perform a Depth-First-Search (DFS) (e.g., post-order) traversal and remove the candidate CPUs from each Cgroup (cpuset. cpus) from bottom to top of the Cgroups tree hierarchy not associated with the critical application. CPU Isolator 402 may append the candidate CPUs only to the path in the hierarchy that leads to the Cgroup where the critical application tasks (e.g., processes, threads) will be running.
[0055] FIG. 6 illustrates a modified Cgroups after assigning candidate CPUs to the critical application (running in “Isolated Container- A”). In this example, the application requests for two CPUs, hence core number “5” and its sibling hyper-thread “37” are assigned to the container-A. It is to be noted that, except for the path that leads to the container-A, cores “5” and “37” were excluded from all other Cgroups dynamically. This provides isolation - other processes do not have access to CPU cores 5 and 37, as shown in FIG. 6.
[0056] The aforementioned steps help to isolate user-space and some system processes on-demand from the CPUs on which the critical application will be provisioned. To reach the level of isolation provided by boot-time isolcpus, the next suggested steps may be taken for shielding the CPUs from system processes.
[0057] The method may include moving the interrupts (interrupt requests (IRQs)) away from candidate CPUs. In Linux, IRQs are used by devices to request service from the CPU. When a hardware device needs to signal the CPU that it requires processing or service, it sends an interrupt request corresponding to the interrupt ID. The CPU stops its current task, saves its context, and jumps to the interrupt handler routine associated with that IRQ to service the interrupt. Typically, these interrupts can be triggered and handled on any available CPU in the system. Handling interrupts on the isolated cores will preempt critical application tasks with the high-priority interrupt handler, leading to CPU jitter. A simple and effective way to prevent an IRQ from running on a particular CPU is to control the IRQ’s CPU affinity. Iproclirql< irq_no > / smp_affinity_list provides a list of CPUs on which an IRQ can be handled. CPU Isolator 402 may dynamically remove candidate CPUs from some or all of the IRQs in the system. This stops interrupts getting serviced on isolated CPUs. Also, CPU Isolator 402 may disable the “irqbalance” daemon program from interfering with the CPU Isolator’s 402 decisions.
[0058] In some embodiments, CPU Isolator 402 may selectively migrate IRQs based on an application’s demand. Certain applications, for example, might rely on interrupts. For example, a network packet processing application might rely on interrupts from SmartNICs, or a data intensive application might rely on interrupts from storage, and so on. Migrating all interrupts can deteriorate an application’s performance. A possible solution is to allow developers to specify the interrupts that are required for an application to function efficiently. CPU Isolator 402 can then migrate other interrupts out of the isolated cores, while allowing the user specified interrupts to be serviced on the isolated cores where the critical application is running.
[0059] The method may include converting candidate CPUs to Adaptive-ticks (NOHZ_FULL). “NoHZ_Full" is a Linux kernel feature related to adaptive-ticks. It stands for "No-Hz Full Dynticks". Other operating systems may have similar features. The Linux kernel traditionally used periodic timer interrupts (ticks) to manage various tasks, including timekeeping and scheduling. However, these periodic interrupts can interrupt the critical process. NOHZ_FULL is a kernel parameter that can be used to disable the scheduler tick on a CPU. In short, the CPU will not be interrupted by the scheduler at regular intervals.
[0060] Unlike previous optimizations, Linux doesn’t provide system calls or APIs to enable / disable NOHZ_FULL (adaptive-ticks mode) from user-space. In reference to FIG. 3, embodiments may utilize the CPU Isolator’s 402 (loadable) kernel module 306 to control the NOHZ_FULL parameter. CPU Isolator’s 402 user-space component 304 may signal the list of candidate CPUs selected for dynamic isolation to its counterpart kernel space component 306, possibly via Netlink sockets. Kernel space component 306 may use kernel APIs to control the NOHZ_FULL CPUs. The Linux kernel maintains the list of CPUs using a variable “tick jwhz_f ull_mask” . CPU Isolator’s 402 kernel component 306 may control this mask to dynamically to enable / disable a CPU into adaptive-ticks.
[0061] Note: With limitations in Linux architecture, modifying NOHZ_FULL through kernel parameters may not work. This is due to a fact that variables (specifically CPUs list) corresponding to NOHZ_FULL are allocated during boot-time. Allocating these variables at runtime leads to unexpected system state. One workaround proposed is to add the following parameters in grub:GRUB_CMDLINE_LINUX = " ... isolcpus = nohz, < cpu_list > “All the CPUs available in the system (except CPU-0) may be added to the “cpu_list” parameter. The kernel will boot the CPUs specified in the list as dyntick (adaptive-ticks) and corresponding memory for the CPU list will be allocated. Running the system with all CPUs in adaptive-ticks mode can degrade the overall system performance. In some embodiments, only the required CPUs may be made adaptive-ticks on-demand. Hence, CPU Isolator 402 may revert the adaptive-ticks CPUs to regular CPUs when the system boots. Upon the request for isolated CPUs, CPU Isolator 402 may make the candidate CPUs adaptive-ticks (dynticks) on-demand.
[0062] The method may include offloading Read-Copy-Update (RCU) Callback workers threads. RCU_NOCBS stands for "RCU No callbacks" is a kernel parameter in Linux that controls whether RCU callbacks are offloaded to threads. RCU callbacks are a way for the RCU (Read-Copy-Update) system to perform cleanup operations after RCU-protected updates have been committed. By default, RCU callbacks are executed on the same CPU where the RCU- protected update was performed. However, if RCU_NOCBS is set, then the RCU callbacks will be offloaded to kernel threads that can run on any CPU. Offloading RCU callbacks to threads can improve performance by reducing the number of interrupts that are generated by RCU. This is especially beneficial for CPUs that are running other high-priority tasks. Most of the existing work uses boot-parameters to control RCU offloads.
[0063] Newer kernel versions have exposed APIs to dynamically offload and de-offload RCU callbacks through rcu_nocb_cpu_offloadQ and rcu_nocb _cpu_deof floadQ APIs. CPU Isolator’s 402 user-space component 304 may signal the list of candidate CPUs selected for dynamic isolation to its counterpart kernel space component 306, possibly via Netlink sockets. Kernel-space component 306 may utilize kernel APIs to migrate RCU callbacks away from isolated CPUs.
[0064] The method may include assigning the critical application to dynamic Isolcpus. In some embodiments, the isolated CPUs may be taken offline and then brought back online. This is recommended as a practice to force tasks running on isolated CPUs to migrate to other CPUs with immediate effect. CPU isolator 402 may migrate the application to the Cgroup where the isolated cores are assigned. Since the Cgroup has only the dynamically isolated CPUs, theapplication will be automatically scheduled on isolated CPU cores. In case of docker or Kubernetes, the container runtime takes care of creating the Cgroup folder for the pods and container. Container runtimes automatically provision the container processes onto their corresponding Cgroups.
[0065] When a new critical application arrives, CPU Isolator 402 may repeat steps for creating a dynamic isolated CPU for the new critical application.
[0066] The method may include decommissioning dynamic Isolcpus assigned to the critical application. When the critical application terminates or is killed, the CPU Isolator 402 may decommission Isolcpus by one or more of:• De-offloading RCU callbacks to the isolated CPUs.• Allowing periodic timer ticks on the isolated CPUs.• Rebalancing and allowing the IRQs on the isolated CPUs.• Reassigning isolated CPUs to other Cgroups, making Isolcpus schedulable for other processes in the system (becomes shared CPUs).• If there are no tasks running under the Cgroup where critical application was provisioned, CPU Isolator 402 can delete the Cgroup folder from the Cgroup hierarchy. If the Cgroup folders are created by Kubernetes or Docker, typically the orchestrator takes care of deleting the Cgroup folder when all the process in the Cgroup folder terminates. But if the Cgroup folder is created by CPU Isolator 402 then CPU Isolator 402 may delete the Cgroup folder. To avoid active monitoring of process liveness in a Cgroup, CPU Isolator 402 may use a notification-based approach. On CgroupVl a "notify_on_release" flag can be used to notify CPU Isolator when the last process in Cgroup exits. On CgroupV2 a “inotify" Linux functionality can be used on "cgroup. event" Cgroup file to reports the number of active process on the Cgroup.
[0067] FIG. 7 illustrates the performance comparison of a test application running on regular non-optimized CPUs and on the dynamically isolated CPUs from research experiments. The results of dynamic isolcpus are in par with boot-time isolcpus CPUs (achieving Ops CPU jiter).
[0068] FIG. 8 illustrates a flowchart showing one sequence involved in the creation of dynamic Isolcpus during an initialization phase and at a runtime phase of an embodiment. The user space component 304 of CPU Isolator 402 may take a number of inputs, including, for example, application resource specifications, system administrator policies and requirements, and application requirements (such as described above). Similarly, the kernel space component 306 of CPU Isolator 402 may also take a number of inputs, including, for example, boot-time kernel parameters (such as described above). User space component 304 may begin an initialization phase 802. As part of the initialization phase 802, user space component 304 may load the kernel space component 306. User space component 304 may also build system topology 804. Part of building system topology 804 may include identifying system resources and mappings of the system resources. These resources and mappings may include CPUs, memory, Cgroups, NUMA, NUMA-10, and relationships between these resources. Building system topology 804 may also include building available resource topology based on an administrator’s policies. During the initialization phase, kernel space component 306 may revert boot parameters 820. Part of reverting boot parameters 820 may include reversing the adaptive tick (NOHZ_FULL) CPUs to regular CPUs with ticks enabled. Reverting boot parameters 820 may also include deoffloading the RCU callbacks to the CPUs where the RCU callbacks were offloaded. At initialization, in some embodiments, the servers may be started normally, and the CPUs may be in a normal (non-isolated) state. As described herein, some CPUs may also be boot-time isolated.
[0069] During the runtime phase, user space component 304 may select candidate CPUs 806, for example, based on system topology and available resources (e.g., from building system topology 804) and / or an application’s CPU and other resource requirements. Selecting candidate CPUs 806 may proceed as described above. The success 808 of the selecting candidate CPUs 806 is checked, and if not successful, an error is returned to the caller. Otherwise, if successful, operation continues to managing Cgroups 810. Managing Cgroups 810 may include performing a depth-first search postorder traversal to remove candidate CPUs from Cgroups not associated with the critical application. Managing Cgroups 810 may also include performing a depth- first search preorder traversal on the path leading to the Cgroup associated with the critical application and appending the candidate CPUs to this path. The Cgroup associated with thecritical application may be provided (e.g., by an application, parameter, user, administrator, and so on), or may be determined by analyzing the Cgroup hierarchy and application. Operation may continue to migrating IRQs away from candidate CPUs 812. Migrating IRQs away from candidate CPUs 812 may include, for each IRQ in / proc / irq, checking if the IRQ is allowed on the candidate CPUs. If not, the next IRQ may be processed. If so, then the corresponding smp_affinity_list file may be modified to exclude the candidate CPUs. This process is continued until all the IRQs have been processed. After this, a message may be sent 814 to the kernel component 306 (e.g., a Netlink message).
[0070] During the runtime phase, kernel space component 306, after receiving the message sent at 814, may receive a list of candidate CPUs 820. Kernel space component 306 may then convert the candidate CPUs to adaptive tick (NOHZ_FULL) 822. Converting candidate CPUs to adaptive tick 822 may include modifying the tick_nohz_full_mask variable. Kernel space component 306 may also offload RCU callbacks from candidate CPUs 824. Offloading RCU callbacks from candidate CPUs 824 may include using “rcu_nocb_cpu_offload()” function to offload RCU callbacks to kernel threads on other (nonisolated) cores. Kernel space component 306 may signal the completion of dynamic isolcpus creation to user space component 304. Following this, user space component 304 may assign dynamic isolcpus to an application’s Cgroup 816. The CPU Isloator 402 may then enter a wait state until a new request for creating dynamic isolcpus arrives 818.
[0071] Variations to the embodiments discussed above are within the scope of this disclosure. For example, some embodiments described above involve creating dynamic Isolcpus for applications at schedule time. But the same concept can be applied on existing workloads in the system (i.e., on running applications) if proper inputs are provided to the CPU Isolator 402. If required, CPU Isolator 402 can reallocate CPUs for critical workloads or can convert the preallocated CPUs to isolated CPUs on-demand.
[0072] Due to the current limitations with the existing Linux architecture, proposed embodiments must split the functionalities between user-space 304 and kernel-space 306 components. But if Linux exposes control for NOHZ_FULL and RCU callbacks to user-space, all functionalities could be handled by the user-space component 304. Also, in certain scenarios CPU Isolator 402 can offload the CPU selection for critical application to another entity. Forexample, if Kubernetes uses CPU Manager for allocation of CPUs to workloads, then management of CPUs can be offloaded to Kubernetes CPU Manager. Kubernetes CPU manager allocates the CPUs, while CPU Isolator 402 isolates the CPUs using the dynamic isolation steps described above.
[0073] As another example, it is important to provide a high-level view of CPU Isolator 402 in the context of Cloud-native environments like Kubernetes. Embodiments provide at least two possible implementations for CPU Isolator 402. One possible implementation is as a nontransparent proxy using CRI interface. FIG. 9 illustrates a high-level architecture of CPU Isolator 402 implemented as a CRI proxy. This is the method we have followed in our research prototype. CRI proxy acts as a non-transparent proxy, intercepting the CRI messages sent between Kubelet and the container runtime. CRI proxy runs as a system daemon having access to the host operating system resource control interface. An Example of CRI proxy-based is Intel CRI-RM. This approach doesn’t require changes to code or functionality of Kubernetes.
[0074] CRI messages contain the information about the workload, especially the number of CPUs required and so on. As shown in the figure, one non-limiting method is to use “Annotation” fields in the Kubernetes Pod specification as a hint / intent to specify if the application (containers) requires CPU isolation. CPU Isolator can parse this information and can provide isolation only to the container requesting isolation. Another possible method is by the cloud operator specifying explicitly which application requires CPU isolation. Based on the user intents, cloud operator policies and application requirements (e.g., CPUs required, request for any of hardware accelerator, and so on), CPU Isolator identifies the suitable CPUs and makes them isolated dynamically while scheduling the container. When the container starts, it runs on isolated cores. Note: This solution is also compatible with Kubernetes CPU Manager, where Kubernetes CPU Manager assigns the CPU cores, and CPU Isolator isolates the assigned cores.
[0075] Another possible implementation is as a container runtime plugin using Node Resource Interface (NRI). FIG. 10 illustrates a high-level architecture of CPU Isolator implementation as an NRI plugin inside container runtimes like Containerd or CRI-O. This functionality is the same as CRI proxy intercepting CRI messages and configuring CPU resources, but using Node Resource Interface (NRI). An example of NRI plugin is Containerd’ s NRI plugin.
[0076] FIG. 11 is a flowchart illustrating a process 1100 for providing isolation from an operating system (OS) scheduler, according to an embodiment. Process 1000 may begin in step si 102, and may be performed, e.g., by a server 302, CPU isolator 402, user space component 304, and / or kernel space component 306.
[0077] Step si 102 comprises receiving a request to isolate one or more CPUs for an application.
[0078] Step si 104 comprises selecting a set of CPUs to isolate based on the request.
[0079] Step si 106 comprises, after boot-time, restricting tasks not associated with the application from accessing the set of CPUs to isolate.
[0080] In some embodiments, the request to isolate one or more CPUs for an application is received on-demand at run-time. In some embodiments, selecting a set of CPUs to isolate based on the request comprises selecting one or more exclusive CPU cores based on system administrator policies, information regarding reserved CPUs, and information regarding Non- Uniform Memory Access (NUMA) and NUMA Input-Output (IO) relationships. In some embodiments, the application has critical performance related requirements which must be met. In some embodiments, the tasks comprise one or more processes and / or threads. In some embodiments, restricting tasks not associated with the application from accessing the set of CPUs to isolate comprises restricting tasks in a control group not associated with the application from having access to the set of CPUs to isolate for one or more control groups. In some embodiments, restricting tasks in a control group not associated with the application from having access to the set of CPUs to isolate for one or more control groups comprises performing a depth first search (e.g., post-order) traversal of the one or more control groups and modifying CPU affinity for control groups not associated with the application to restrict those control groups from having access to the set of CPUs to isolate.
[0081] In some embodiments, restricting tasks not associated with the application from accessing the set of CPUs to isolate comprises offloading one or more interrupt requests from the set of CPUs to isolate to prevent the interrupt requests from being handled on the set of CPUs to isolate. In some embodiments, offloading one or more interrupt requests from the set of CPUs to isolate to prevent the interrupt requests from being handled on the set of CPUs to isolatecomprises offloading all interrupt requests from being handled on the set of CPUs to isolate. In some embodiments, restricting tasks not associated with the application from accessing the set of CPUs to isolate comprises converting the set of CPUs to an adaptive tick mode. In some embodiments, the method further includes restricting tasks not associated with the application from accessing the set of CPUs to isolate comprises offloading read copy update (RCU) callbacks from the set of CPUs to isolate to alternative CPUs. In some embodiments, the method further includes taking the set of CPUs to isolate offline to force tasks running on the set of CPUs to migrate and then bringing the set of CPUs online. In some embodiments, the method further includes reverting the set of CPUs to isolate to a non-isolated state when the application is terminated or killed.
[0082] FIG. 12 is a block diagram of apparatus 1200 (e.g., server 302, CPU Isolator 402, user space component 304, kernel space component 306), according to some embodiments, for performing the methods disclosed herein. As shown in FIG. 12, apparatus 1200 may comprise: processing circuitry (PC) 1202, which may include one or more processors (P) 1255 (e.g., a general purpose microprocessor and / or one or more other processors, such as an application specific integrated circuit (ASIC), field-programmable gate arrays (FPGAs), and the like), which processors may be co-located in a single housing or in a single data center or may be geographically distributed (i.e., apparatus 1200 may be a distributed computing apparatus); at least one network interface 1248 comprising a transmitter (Tx) 1245 and a receiver (Rx) 1247 for enabling apparatus 1200 to transmit data to and receive data from other nodes connected to a network 1210 (e.g., an Internet Protocol (IP) network) to which network interface 1248 is connected (directly or indirectly) (e.g., network interface 1248 may be wirelessly connected to the network 1210, in which case network interface 1248 is connected to an antenna arrangement); and a storage unit (a.k.a., “data storage system”) 1208, which may include one or more non-volatile storage devices and / or one or more volatile storage devices. Interface 1260 may connect PC 1202 and storage unit 1208, interface 1262 may connect PC 1202 and network interface 1248, and interface 1264 may connect network interface 1248 and network 1210. In embodiments where PC 1202 includes a programmable processor, a computer program product (CPP) 1241 may be provided. CPP 1241 includes a computer readable medium (CRM) 1242 storing a computer program (CP) 1243 comprising computer readable instructions (CRI) 1244.CRM 1242 may be a non-transitory computer readable medium, such as, magnetic media (e.g., a hard disk), optical media, memory devices (e.g., random access memory, flash memory), and the like. In some embodiments, the CRI 1244 of computer program 1243 is configured such that when executed by PC 1202, the CRI causes apparatus 1200 to perform steps described herein (e.g., steps described herein with reference to the flow charts). In other embodiments, apparatus 1200 may be configured to perform steps described herein without the need for code. That is, for example, PC 1202 may consist merely of one or more ASICs. Hence, the features of the embodiments described herein may be implemented in hardware and / or software.
[0083] While various embodiments are described herein, it should be understood that they have been presented by way of example only, and not limitation. Thus, the breadth and scope of this disclosure should not be limited by any of the above described exemplary embodiments. Moreover, any combination of the above-described embodiments in all possible variations thereof is encompassed by the disclosure unless otherwise indicated herein or otherwise clearly contradicted by context.
[0084] Additionally, while the processes described above and illustrated in the drawings are shown as a sequence of steps, this was done solely for the sake of illustration. Accordingly, it is contemplated that some steps may be added, some steps may be omitted, the order of the steps may be re-arranged, and some steps may be performed in parallel.
Claims
CLAIMS1. A method for providing isolation from an operating system (OS) scheduler, the method comprising: receiving a request to isolate one or more central processing units (CPUs) for an application; selecting a set of CPUs to isolate based on the request; and after boot-time, restricting tasks not associated with the application from accessing the set of CPUs to isolate.
2. The method of claim 1, wherein the request to isolate one or more CPUs for an application is received on-demand at run-time.
3. The method of any one of claims 1-2, wherein selecting a set of CPUs to isolate based on the request comprises selecting one or more exclusive CPU cores based on system administrator policies, information regarding reserved CPUs, and information regarding Non- Uniform Memory Access (NUMA) and NUMA Input-Output (IO) relationships.
4. The method of any one of claims 1-3, wherein the application has critical performance related requirements which must be met.
5. The method of any one of claims 1-4, wherein the tasks comprise one or more processes and / or threads.
6. The method of any one of claims 1-5, wherein restricting tasks not associated with the application from accessing the set of CPUs to isolate comprises restricting tasks in a control group not associated with the application from having access to the set of CPUs to isolate for one or more control groups.
7. The method of claim 6, wherein restricting tasks in a control group not associated with the application from having access to the set of CPUs to isolate for one or more control groups comprises performing a depth first search traversal of the one or more control groups and modifying CPU affinity for control groups not associated with the application to restrict those control groups from having access to the set of CPUs to isolate.
8. The method of any one of claims 1-7, wherein restricting tasks not associated with the application from accessing the set of CPUs to isolate comprises offloading one or more interrupt requests from the set of CPUs to isolate to prevent the interrupt requests from being handled on the set of CPUs to isolate.
9. The method of claim 8, wherein offloading one or more interrupt requests from the set of CPUs to isolate to prevent the interrupt requests from being handled on the set of CPUs to isolate comprises offloading all interrupt requests from being handled on the set of CPUs to isolate.
10. The method of any one of claims 1-9, wherein restricting tasks not associated with the application from accessing the set of CPUs to isolate comprises converting the set of CPUs to an adaptive tick mode.
11. The method of any one of claims 1-10, wherein restricting tasks not associated with the application from accessing the set of CPUs to isolate comprises offloading read copy update (RCU) callbacks from the set of CPUs to isolate to alternative CPUs.
12. The method of any one of claims 1-11, further comprising taking the set of CPUs to isolate offline to force tasks running on the set of CPUs to migrate and then bringing the set of CPUs online.
13. The method of any one of claims 1-12, further comprising reverting the set of CPUs to isolate to a non-isolated state when the application is terminated or killed.
14. A server (302) comprising: processing circuitry (1202); and a memory, the memory containing instructions (1244) executable by the processing circuitry (1202), whereby when executed the processing circuitry (1202) is configured to: receive a request to isolate one or more central processing units (CPUs) for an application; select a set of CPUs to isolate based on the request; and after boot-time, restrict tasks not associated with the application from accessing the set of CPUs to isolate.
15. The server of claim 14, wherein the request to isolate one or more CPUs for an application is received on-demand at run-time.
16. The server of any one of claims 14-15, wherein selecting a set of CPUs to isolate based on the request comprises selecting one or more exclusive CPU cores based on system administrator policies, information regarding reserved CPUs, and information regarding Non- Uniform Memory Access (NUMA) and NUMA Input-Output (IO) relationships.
17. The server of any one of claims 14-16, wherein the application has critical performance related requirements which must be met.
18. The server of any one of claims 14-17, wherein the tasks comprise one or more processes and / or threads.
19. The server of any one of claims 14-18, wherein restricting tasks not associated with the application from accessing the set of CPUs to isolate comprises restricting tasks in a control group not associated with the application from having access to the set of CPUs to isolate for one or more control groups.
20. The server of claim 19, wherein restricting tasks in a control group not associated with the application from having access to the set of CPUs to isolate for one or more control groups comprises performing a depth first search traversal of the one or more control groups and modifying CPU affinity for control groups not associated with the application to restrict those control groups from having access to the set of CPUs to isolate.
21. The server of any one of claims 14-20, wherein restricting tasks not associated with the application from accessing the set of CPUs to isolate comprises offloading one or more interrupt requests from the set of CPUs to isolate to prevent the interrupt requests from being handled on the set of CPUs to isolate.
22. The server of claim 21, wherein offloading one or more interrupt requests from the set of CPUs to isolate to prevent the interrupt requests from being handled on the set of CPUs to isolate comprises offloading all interrupt requests from being handled on the set of CPUs to isolate.
23. The server of any one of claims 14-22, wherein restricting tasks not associated with the application from accessing the set of CPUs to isolate comprises converting the set of CPUs to an adaptive tick mode.
24. The server of any one of claims 14-23, wherein restricting tasks not associated with the application from accessing the set of CPUs to isolate comprises offloading read copy update (RCU) callbacks from the set of CPUs to isolate to alternative CPUs.
25. The server of any one of claims 14-24, wherein the processing circuitry (1202) is further configured to take the set of CPUs to isolate offline to force tasks running on the set of CPUs to migrate and then bring the set of CPUs online.
26. The server of any one of claims 14-25, wherein the processing circuitry (1202) is further configured to revert the set of CPUs to isolate to a non-isolated state when the application is terminated or killed.
27. A computer program (1243) comprising instructions which when executed by processing circuitry (1202) of a node (1200), causes the node (1200) to perform the method of any one of claims 1-13.
28. A carrier containing the computer program (1243) of claim 27, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium (1242).
Citation Information
Patent Citations
A method, apparatus and system for real-time virtual network function orchestration
WO2019084793A1
Cited By
Core isolation for errors
US20250390371A1