Lock-free work deployment thread scheduler
The lock-free thread scheduling mechanism addresses the inefficiencies in conventional multi-compute engine systems by dynamically distributing threads across multiple CPUs through a shared ring buffer, resulting in improved I/O performance and consistent system performance.
Patent Information
- Application Number
- DE102021108963
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-08-31
- Filing Date
- 2021-04-11
- Publication Date
- 2025-06-05
- Estimated Expiration
- 2041-04-11
AI Technical Summary
Conventional multi-compute engine systems lack an efficient mechanism to distribute work among multiple CPUs, leading to CPU saturation and bottlenecks, which result in lower I/O performance and performance variations between boot cycles or array nodes.
A lock-free thread scheduling mechanism that dynamically allocates threads to local CPU run queues and utilizes a shared ring buffer for thread distribution, allowing idle CPUs to steal threads from other CPUs and ensuring balanced workload distribution.
The proposed solution effectively distributes work among multiple CPUs, preventing saturation and bottlenecks, thereby enhancing I/O performance and achieving more consistent performance across system boot cycles and array nodes.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Description of the Prior ArtThe emerging technology has led to exponential growth in computing power of computing systems. The use of multi-processor devices (e.g., multi-computer processing unit or CPU) and multi-core processors (including a number of cores or processors) in computing systems has also contributed to increasing computing power of computing systems. Each of the cores or processors may include an independent cache memory.Brief Description of the DrawingsThe present disclosure will be described in detail according to one or more different embodiments with reference to the following figures. The figures are for illustrative purposes only and are merely representative or exemplary embodiments. FIG. 1A shows an example of a hardware computer system in which various embodiments may be implemented. FIG. 1B illustrates an example one-to-many core processing system architecture of a processor in the hardware computing system of FIG. 1A. FIG. 2A illustrates an example workflow for thread scheduling, in accordance with various embodiments. FIG. 2B illustrates an example query workflow in accordance with various embodiments. FIG. 3 illustrates an example computer component capable of executing instructions to perform thread scheduling, in accordance with various embodiments. FIG. 4 illustrates an example computer component capable of executing instructions to perform queries in accordance with various embodiments. FIG. 5 illustrates an example computer component with which various features and / or functions described herein may be implemented.The figures are not intended to be exhaustive and do not limit the present disclosure to the precise form illustrated.Detailed DescriptionProcessors or CPUs refer to electronic circuits within a computer that execute the instructions of a computer program by performing the basic arithmetic, logical, control, and input / output (I / O) operations specified by the instructions. The processing performance of computers can be increased by using multi-core processors or CPUs, which essentially results from two or more individual processors (called cores in this sense) being plugged into an integrated circuit. Ideally, a dual-core processor would be almost twice as powerful as a single-core processor, although in practice the actual power gain may be lower. Increasing the number of cores in a processor (i.e., dual-core, quad-core, etc.) increases the workload that can be processed in parallel. This means that the processor can now process numerous asynchronous events, interrupts, etc. However, multiprocessor computers or systems may support more than one processor or CPU, e.g., two to eight or even many more CPUs, as may be the case with Petascale supercomputers and Exascale supercomputing systems.Generally, the memory of a computer system includes main memory, e.g., non-volatile memory (NVM), and cache memory (or simply cache). The main memory may be a physical device used to store application programs or data in the computing system. The cache stores data that is frequently accessed, so that no time is required to access the data from the main memory. Typically, the data is transferred between the main memory and the cache in fixed size blocks referred to as cache lines. When a processor of the computer system must read from or write to a location in main memory, the processor reads from or writes to the cache if the data is already present in the cache, which is faster than reading from or writing to main memory. Data written to the cache is generally written back to main memory.In a multi-compute engine system, each compute engine, e.g., a core or processor, may have multiple threads and include one or more caches, and generally, a cache is organized as a hierarchy with one or more levels of cache. In conventional multi-compute engine systems, threads typically run until the thread enters, sleeps, or exits, and each logical thread can be assigned to a CPU at the time the logical thread was created. Successive CPUs are selected, e.g., in a round robin fashion, for assignment to a logical thread, and only the assigned CPU can execute the thread. However, such a scheduling mechanism lacks a way to spread work among multiple CPUs, and some CPUs may saturate while others remain virtually idle. The saturated CPUs may become a bottleneck, resulting in lower I / O performance that could be achieved if the same amount of work were distributed to more CPUs. In addition, performance variations may occur between boot cycles or between array nodes because the round robin CPU assignment may vary every boot of a node.Accordingly, various embodiments are directed to thread scheduling that is lock-free and allows threads to be dynamically executed by different CPUs at runtime. In some embodiments, a scheduler may follow a particular scheduling algorithm that performs: (1) threads are allocated to local CPU run queues; (2) threads are placed in a ring buffer shared by all CPUs when the threads are ready to run; (3) when the local run queue of a CPU is emptied, that CPU checks the shared ring buffer to determine whether any threads are waiting to run on that CPU, and when so, the CPU pulls a stack of threads associated with that ready to run thread and places the threads in their local run queue; (4) If not, an idle CPU randomly selects another thread stealing CPU (preferably, a closer CPU), and the idle CPU attempts to discard a thread stack connected to the CPU from the shared circular buffer - this can be repeated if a selected CPU does not have associated threads in the shared circular buffer; (5) the idle CPU executes the thread stack stolen from the other CPU in its local execution queue in priority order; (6) the process repeats for each idle CPU. As described further below, a scheduler may be implemented in executable software, e.g., at node level.FIG. 1A shows an example of a hardware computer system 1 with multi-core processors or CPUs, the processing threads of which may be compensated according to various embodiments. A system, such as hardware computer system 1, may include various elements or components, circuits, software stored and executable thereon, etc. However, for simplicity and ease of explanation, only some aspects are shown and described.The hardware computer system 1 may include an operating system or OS 2 in which one or more processes including zero or more threads run (or are idle) on multiple CPUs ( 10 and 30). As described herein, a process may have the ability to generate a thread that in turn generates another "lightweight" process that shares the data of the "parent" process, but that may be independently executed on another processor at the same time as the parent process. For example, a process N4 (along with other processes) in operating system 2 may be idle, while other processes, e.g., processes 12, 32, are running on CPUs 10, 30, respectively. FIG. 1A further illustrates a high level flow for memory maps of each executing process that includes virtual memory map 40 for CPU 10 and virtual memory map 42 for CPU 30, as well as physical memory map 50.FIG. 1B illustrates an example architecture of the CPU 10 of FIG. 1A. In one example, CPU 10 may operate with a directory-based protocol to achieve cache coherency. CPU 10 may be, for example, a multiprocessor system, a system having one or more cores, or a system having multiple cores. Accordingly, the CPU 10 may include multiple processors or cores, and in this example, multiple cores 10A, 10B,... 10N.Each of the cores 10A, 10B,... 10N, one or more cache levels 14A, 14B,...14N may be associated. A network 16 (which may be a system bus) allows the cores 10A, 10B,...10N to communicate with each other as well as with a main memory 18 of the CPU 10. The data of the main memory 10 may be temporarily stored by each core of the CPU 10, for example, each of the cores 10A, 10B... 10N.Returning to FIG. 1A, the CPUs 10, 30 each implement a virtual memory management system. For example, CPU 10 generates memory references by first forming a virtual address representing the address within an entire range of addresses by the architectural specifications of the computer or the portion thereof permitted by operating system 2. The virtual address may then be translated into a physical address in the physical memory map 50 that is limited by the size of the main memory. In some embodiments, the translation is with pages such that a virtual page address for a page in virtual memory map 40 is translated to a physical address for a page in physical memory map 50. A page table is maintained in memory to provide the translation between virtual address and physical address, and typically a translation buffer (not shown) is included in the CPU to hold the last used translations so that a table reference in memory 18 need not be made to obtain the translation before a data reference can be made.As indicated above, various embodiments are work-stealing to allow threads, e.g., threads, to be dynamically executed by different CPUs at runtime. In this way, work is distributed among the available CPUs without overloading a particular CPU and / or creating / creating a bottleneck in processing. In some embodiments, the work-stealing algorithm, after which a local CPU scheduler operates (e.g., scheduler 314 (FIG. 3 )), embodies a non-blocking approach. Locking may refer to a mechanism by which a serializing resource protects data from access by many threads. That is, typically a core may acquire spin lock to allow access to data structures to be synchronized and to avoid processing conflicts between incoming / outgoing packets. However, locking may result in wasted CPU cycles / spinning while waiting for the acquisition of a lock. In particular in modern systems with more than e.g. 100 cores per system, conventional thread scheduling would greatly impede (hardware) scaling and / or lead to high latencies.FIG. 2A illustrates an example workflow for implementing a non-blocking work-stealing approach to scheduling threads, according to an embodiment. A local CPU 200 is shown in FIG. 2A, where the local CPU 200 may have a local run queue 205 with threads 210, 212, 214 to be executed by the local CPU 200. As mentioned above, a CPU may have one or more cores. Threads for processing requests may be created in each CPU core. One thread can be created per CPU core, but a plurality of threads are also possible. Threads typically continue to run as long as there is work, i.e., a request must be processed. A thread may refer to a basic unit of CPU utilization and may include a program counter, a stack, and a set of registers. Threads run in the same memory context and may share the same data during execution. A thread is what a CPU is actually allowed to execute, and access to shared resources, e.g., the CPU, is scheduled accordingly. In contrast, a process includes or initiates a program, and a process may include / include one or more threads (the thread is an execution unit in the process). And while threads use address spaces of the process, when a CPU switches from one process to another, the current information is stored in a process descriptor and the information of the new process is loaded.As shown in FIG. 2A, a shared ring buffer may be implemented and made accessible to each CPU. That is, in some embodiments, each CPU, such as CPU 200, may be implemented therein or have access to a cache or queue of threads to be executed by the CPU. Unlike conventional systems, where threads may be allocated to a particular CPU, threads are queued in a shared circular buffer, e.g., shared circular buffer 215. It should be appreciated that any suitable data structure or mechanism may be used to queue pending threads to be executed by the CPU 200. The shared ring buffer is merely one way of implementing such a thread queue. Other data structures or storage mechanisms may be used to queue threads in accordance with other embodiments.It should be appreciated that the shared circular buffer 215 may be implemented in physical / main memory (e.g., memory 18 of FIG. 1B ) that is accessible by each CPU in a system. In some embodiments, the shared ring buffer 215 may be implemented locally in software, e.g., as buffers that may be accessed by local CPUs. It should also be noted that data coherence may be maintained in hardware, i.e., the work stealing algorithm need not take into account data coherence. For example, a node controller controlling a node to which a CPU may belong may maintain a coherency directory cache to ensure data coherency between main memory and local memory, and the atomic instructions used to implement the work-stealing algorithm already act on the shared memory. Accordingly, no locks are used or needed.In operation, local CPU 200 may place all woken threads into their shared circular buffer, in this case shared circular buffer 215. Threads may be generated by a fork or similar function(s) in a computer program. In other cases, threads may be woken by, for example, a cron job in which the thread is woken after a certain timer / after a certain period of time has elapsed. After waking up, threads may be placed into the shared circular buffer 215 from the local CPU 200. As shown in FIG. 2A, threads 220- 228 are currently located in shared circular buffer 215. At this time, threads 220- 228 are not yet dedicated to any particular CPU.It should be noted that each thread may have an allowed set of CPUs on which it may run, and this allowed set may be programmable and determined, e.g., by a developer who creates the thread and defines its attributes. For example, a thread may be set to be created on a particular node. In this case, a particular non-uniform memory access (NUMA) affinity may be set for a particular CPU / memory. For example, a thread may be set for creation on a particular CPU (CPU affinity). For example, a thread may not have affinity, in which case each CPU may execute the thread. It should be understood that NUMA refers to a memory design, the architecture of which may include a plurality of nodes interconnected via a symmetric multi-processing (SMP) system. Each node itself may be a small SMP that includes multiple processor sockets with processors / CPUs and associated memory interconnected, where the memory within the node is shared by all CPUs. The memory within a node may be considered local memory for the CPUs of the node, while the memory of other nodes may be considered remote memory. Node controllers within each node allow the CPUs to access remote memory within the system. A node controller may be considered an extended memory controller that manages access to some or all of the local memory and access of the CPUs of the node to the remote memory. A scheduler according to various embodiments may be implemented in software executing on the node, e.g. in a node controller.Take the case of a thread with a particular CPU affinity as an example: When the thread is woken up (as described above), the thread is mapped to the CPU to which the woken-up thread is bound (due to the indicated affinity). As mentioned above, the CPU, in this case the local CPU 200, is, at random, the CPU on which the thread is to be executed. Therefore, the CPU 200 places the thread in the shared circular buffer 215.To execute one or more threads currently queued in the shared circular buffer 215, the local CPU 200 may dequeue threads in batches or individually (although dequeuing in batches may be more efficient). That is, each dequeue operation on the shared ring buffer 215 may have some latency (memory overhead), and processing thread stacks may amortize this latency / delay by fetching multiple elements in a single dequeue operation. It should be appreciated that a stack may be any number between one and a shared circular buffer size, although a standard maximum thread stack size may be selected, e.g., based on tests and the tuning of the algorithm upon which the embodiments operate. It should be appreciated that the implementation of a shared circular buffer, e.g., shared circular buffer 215, provides an interface through which up to N elements (threads) may be removed from the queue in a single call and placed in a return buffer for processing by the caller. After a stack of threads is extracted from the shared circular buffer 215, the extracted threads may be placed in the local run queue 205. Depending on the indicated affinity, e.g., to a particular CPU or node or processor socket (in the case of NUMA affinity), only a particular CPU / set of CPUs may be able to queue a stack of threads. In the case of a thread without affinity, each CPU may dequeue the thread.When threads are in the local execution queue 205, the local CPU 200 may execute these threads, in this example, threads 210- 214. In some embodiments, local CPU 200 may execute threads 210- 214 in a priority order, but this is not necessarily required, and the threads may be executed in a different order or in any order. That is, local CPU 200 may execute a thread, e.g., thread 210. When a thread returns to a current context, local CPU 200 may execute another thread, which in this case may be thread 212, 214, etc.When CPUs are idle, i.e., have no threads to execute, such as idle CPU 230, it may "steal" work or a stack of threads from shared circular buffer 215. Idle CPUs, such as idle CPU 230, may steal a stack of threads from the shared circular buffer if the stack of threads is not assigned an affinity that would prevent idle CPU 230 from executing that stack of threads. For example, the idle CPU 230 may fall within a set of CPU / NUMA affinities. If the idle CPU 230 cannot execute a stack of threads, these threads remain in the shared circular buffer 215. Note that all threads may be sent to the shared circular buffer 215 regardless of the affinity, because the shared circular buffer 215 into which the threads are sent corresponds to the affinity of a thread. That is, for each shared ring buffer, only CPUs capable of executing threads (also referred to as tasks) having an affinity corresponding to the shared ring buffer will attempt to steal threads therefrom. After stealing a stack of threads, similar to the operation of the local CPU 200, the idle CPU 230 takes the stack of threads from the shared circular buffer 215. The idle CPU 230 may place the extracted stack of threads into its own local queue (not shown) for execution, and the (no longer inactive) CPU 230 may execute the threads.As mentioned above, each CPU may have at least one shared circular buffer that an idle CPU may steal if there is no affinity that an idle CPU prevents from executing a stack of threads. An idle CPU may be local to the CPU from whose shared ring buffer the idle CPU may steal, or the idle CPU may be remote from the CPU from whose shared ring buffer the idle CPU may steal, depending on the architecture of the system to which the CPUs belong. For example, as indicated above, a system may have a plurality of nodes interconnected, each of the nodes including one or more CPUs.When an idle CPU attempts to steal work (threads) from another CPU, the idle CPU may randomly select another CPU in the system from which it may attempt to steal work. In some embodiments, a hyperthread and NUMA-enabled selection algorithm may be used. In some embodiments, such an algorithm "prefers" to steal CPUs that may be closer with respect to CPU cache location. In some embodiments, an idle CPU may preferably attempt to steal from a shared ring buffer connected to a non-idle CPU located in the same processor socket as the idle CPU.If a shared ring buffer from which an inactive CPU wants to steal work is empty, the inactive CPU may go to another randomly selected CPU from which it wants to steal. It should be noted that, since each CPU is / has associated with at least one shared circular buffer, threads may be placed on and stolen / taken from the thread, the more hardware (CPUs) is added to a system, the more scalability is possible / the better the processing speed is.Considering that the primary mechanism to take better performance out of a system is to add more CPUs, the corresponding addition of more shared ring buffers will result in potentially more CPUs sharing the threads dynamically in a more balanced manner than is possible with conventional systems. Moreover, simply adding more CPUs without a non-blocking work-stealing mechanism, such as that described herein, tends to result in more conflicts between CPUs attempting to execute the same threads and in more latency because more CPU cycles may be missed due to blocking. According to various embodiments, there are no conflicts between the CPUs with respect to the shared buffer ring, except when stealing work and / or when a remote wake is involved. In particular, remote wake may refer to a situation where a CPU cannot place a completed thread on its own shared circular buffer due to affinity constraints (or for some other reason), and the thread is therefore placed on the shared circular buffer of another CPU. If the other CPU accesses the same shared ring buffer at the same time, cache line contention occurs.If the search of the idle CPU for the operation of other CPUs that it can steal does not result in other threads, the idle CPU can go idle until new operation arrives. This prevents an idle CPU from consuming too many system CPU cycles. When an idle CPU goes idle may be based on a certain number of CPU cycles that the CPU requests for work, e.g., a threshold number of CPU cycles. In some embodiments, a predictive algorithm may be used that relies on the time an inactive CPU goes idle on the number of CPU cycles it previously needed to find (steal) work. There are still other conceivable mechanisms that may be used to determine how long a CPU searches for work and / or how long it takes for a CPU to go to sleep without finding work for stealing.If a non-idle CPU, e.g., CPU 200, ceases to work, i.e., its local run queue 205 becomes empty, the CPU first checks its shared ring buffer, in this case, shared ring buffer 215, to determine whether any threads queued therein are waiting to run on CPU 200. If so, the CPU 200 pulls a stack of threads waiting to execute and places them in the local queue 205. CPU 200 may then execute each thread in priority order.In FIG. 2B, a polling architecture is described. A poller may refer to a callback function that is periodically invoked by the scheduler described herein to check a condition, e.g., work, that needs to be performed. Polers can often replace interrupt service routines that typically exist only in kernel mode code.Previous kernel drivers used device interrupts, but interrupts directly to user mode were not feasible, therefore a polling model was used in which polling routines did as little as possible to determine whether work to be done is present, and if so, a "lower half" thread could be triggered to perform the actual work. Polling are distinguished from threads in that they typically consist of a single routine that is called repeatedly, typically runs to completion, and then returns to the scheduler, although polling may be implemented in threads. Thus, the polling model mentioned above executes polling threads on dedicated CPUs in each processor socket. The lower half threads could be executed on other CPUs in the same physical processor socket as the Poller threads (among other threads, but with higher priority). However, similar to the CPU bottleneck problems described above, using dedicated CPUs for polling threads may result in under or over supply of CPU resources for polling.Accordingly, the polling architecture illustrated in FIG. 2B and described below uses a distributed scheme in which polling work can be nested with the "regular" thread execution across all CPUs in a system. In this way, a system can be scaled more easily by dynamically distributing polling across all CPUs in a processor socket, as opposed to using a dedicated set of CPUs for handling the polling.In some embodiments, a shared array of polling 240 may be implemented in a system, where each of the pollings in the array may be associated with a particular socket, e.g., a NUMA socket. As shown in FIG. 2B, a shared array of polling 240 includes polling 240 a, 240 b, and 240 c, for example, each of which may be registered with a particular processor socket that all CPUs in the processor socket may read. Each of the polling devices 240 a, 240 b, and 24 cmay have an atomic flag indicating whether or not it is running on a particular CPU, i.e., a running flag that may be implemented as an atomic variable. It should be noted that all changes to a shared polling array, such as shared polling array 240, may be made when polling are registered or unregistered (which typically occurs rarely), and using a read copy update (RCU) type scheme.Each time a thread returns a CPU, in this example local CPU 200, to the scheduler / scheduling algorithm, the scheduler attempts to execute a polling before scheduling the next thread to execute. That is, when the local CPU 200 experiences / executes a thread or thread scheduling loop, a polling may be randomly selected from the common polling array 240. One of the polling 240a, 240b, or 240c may be selected, e.g., polling 240a, and an attempt is made to set the run flag of that polling 240a atomically using a compare-and-swap (CAS) operation. If the CAS operation succeeds, polling 240a is "claimed" as being executed on the local CPU 200, and no other CPU can execute polling 240a in parallel. Local CPU 200 may then execute polling 240a, clear its run flag, and enable it to be run elsewhere (execution on another CPU). If the CAS operation fails, the local CPU 200 may either retry executing the polling thread with another randomly selected polling, or the local CPU 200 may suspend execution of the polling thread (for the moment) and schedule the next thread to be executed (as described above). Note that the number of attempts to perform a bollard may be configured at runtime. In some embodiments, the attempt to run a bollard per scheduling cycle may be the default mode. In this way, all CPUs in a processor socket can efficiently share the dispatch of multiple pollings with minimal overhead. The cost of attempting to find a bollard to execute may simply be to generate a random index in the shared bollard array, e.g., the shared bollard array 240, and then execute a single CAS instruction to attempt to claim a bollard at that index. This incurred cost is low enough that the Poller selection can be made between each thread execution / yield cycle without introducing significant additional processing latency in the scheduler.It should be noted that using the above polling architecture may result in a lower likelihood of selecting a polling if the polling has not found work for a period of time. This allows the polling to have a preference for more active devices. For example, a memory array on a particular front-end or back-end device port(s) may have higher activity while other ports are idle. Rather than waste time from polling unused ports too often, the resources may be better used for polling active ports, thereby achieving lower latency for requests at those ports. In some embodiments, this preference for active devices may be achieved by implementing a counter for each poller, where a counter is accumulated based on the number of cycles a particular poller was invoked, but the poller has not found any work it can perform. For example, if the accumulated count for a particular bollard reaches or exceeds a threshold, a "delay order" counter may be incremented for that particular bollard. On the other hand, if a randomly selected bollard finds work, the delay order counter may be reset to zero. In addition, each CPU making a query may maintain a sequence number that may be incremented each time the CPU attempts to find a poller for execution (as described above). Thus, if the CPU selects a polling, e.g., if the local CPU 200 selects the polling 240a, the local CPU 200 checks whether its current sequence number is divisible by 2N, where N may refer to the delay order of the selected polling (polling 240a in this case). In other words, the local CPU 200 may check whether the following expression is true: (my_sequence_number MOD (1<< poller_delay_order )) ==0If the term is true, the selected bollard, e.g., bollard 240a, may be executed. If the term is incorrect, local CPU 20 may skip polling 240a and select another polling, in this case either polling 240b or polling 240c. Such a mechanism or logic results in a bollard typically being executed about every 2N times when considered. If N = =0 (the bollard is active), the bollard is executed each time it is considered. However, the longer a bollard is idle, the further its delay order counter increases, and thus this bollard is less frequently executed over time. It should be noted that a maximum upper limit or threshold value may be set up to which the retard order may grow / accumulate, so that each bollard is still executed at an appropriate frequency.FIG. 3 is an example of a computing device 300 in accordance with embodiments of the present disclosure. Where operations and functionality of the computing device 300 are the same or similar to those discussed with respect to FIG. 2A, the description should be interpreted as appropriate. For example, computing device 300 may be an embodiment of a node, a node controller, a CPU such as CPU 10 or 30. The computing device 300 includes one or more hardware processors 302, which may be one or more central processing units (CPUs), semiconductor-based microprocessors, and / or other hardware devices suitable for fetching and executing instructions stored in a machine-readable storage medium 304. The one or more hardware processors 302 may fetch, decode, and execute instructions, such as instructions 306- 314, to control processes or operations for causing error detection and control in the context of coherency directory caches, according to one embodiment. Alternatively or additionally to fetching and executing instructions, the one or more hardware processors 302 may include one or more electronic circuits including electronic components for executing the functionality of one or more instructions, such as a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or other electronic circuits.The one or more hardware processors 302 are configured to execute instructions stored on a machine readable medium 304. The machine-readable medium 304 may be one or more types of non-transitory computer storage media. Non-limiting examples include flash memory, solid state storage devices (SSDs), storage area network (SAN), removable memory (e.g., memory stick, CD, SD cards, etc.), or internal computer RAM or ROM, among other types of computer storage media. The instructions stored on the machine readable medium 304 may include various sub-instructions for executing the function embodied by the identified functions.The one or more hardware processors 302 may execute the instruction 306 to execute threads in a local run queue. As mentioned above, a CPU, e.g., local CPU 200 (FIG. 2A ), may place threads to be executed in a local run queue, e.g., local run queue 205, specific to the CPU (meaning that other CPUs cannot steal or execute work for threads in the local run queue of another CPU. In some embodiments, the CPU may execute the threads in the local run queue in priority order.The one or more hardware processors 302 may execute the instruction 308 to check a shared circular buffer after clearing the local queue. For example, if the local CPU 200 has no more threads to execute in the local queue 205, it has become an idle CPU and the local CPU 200 may check the shared ring buffer 215 to which it has access. The local CPU 200 checks the shared circular buffer 215 to determine whether there are threads to be executed by the local CPU 200, e.g., a particular thread or threads may have a particular affinity for the local CPU 200. If there are no particular threads for the local CPU 200 to execute specifically, the local CPU 200 may steal all threads from the shared circular buffer 215 that it can execute. That is, the local CPU 200 may steal all threads that have an affinity for a group of CPUs to which the local CPU 200 belongs, have a particular NUMA affinity, or have no affinity at all (i.e., any available CPU may execute the thread / s). It should be appreciated that other idle CPUs, such as idle CPU 230, may also check the shared circular buffer to determine whether they may steal threads.Accordingly, the one or more hardware processors 302 may execute the instruction 310 to remove a stack of threads from the shared circular buffer. For example, local CPU 200 may remove a stack of threads it is allowed to execute from the shared circular buffer. The one or more hardware processors 302 may further execute the instruction 312 to place the stack of threads into the local run queue. That is, the local CPU 200 may place this stack of threads that has been removed from the queue into its own local queue 205 to execute, e.g., in the order of priority. The one or more hardware processors 302 may execute the instruction 312 to execute / launch the stack of threads. Now, because the extracted stack of threads that the local CPU 200 has extracted from the shared circular buffer 215 may be executed, the local CPU 200 may execute one thread until the thread gives way, execute a subsequent thread, and so on.FIG. 4 is an example of a computing device 400 in accordance with embodiments of the present disclosure. Where operations and functionality of computing device 400 are the same or similar to those discussed with respect to FIG. 2B, the description should be interpreted as appropriate. For example, computing device 400 may be an embodiment of a node, a node controller, a CPU such as CPU 10 or 30. The computing device 400 includes one or more hardware processors 402, which may be one or more central processing units (CPUs), semiconductor-based microprocessors, and / or other hardware devices suitable for fetching and executing instructions stored in a machine-readable storage medium 404. The one or more hardware processors 402 may fetch, decode, and execute instructions, such as instructions 406- 414, to control processes or operations for causing error detection and control in the context of coherency directory caches, according to one embodiment. Alternatively or additionally to fetching and executing instructions, the one or more hardware processors 402 may include one or more electronic circuits including electronic components for executing the functionality of one or more instructions, such as a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or other electronic circuits.The one or more hardware processors 402 are configured to execute instructions stored on a machine readable medium 404. The machine-readable medium 404 may be one or more types of non-transitory computer storage media. Non-limiting examples include flash memory, solid state storage devices (SSDs), storage area network (SAN), removable memory (e.g., memory stick, CD, SD cards, etc.), or internal computer RAM or ROM, among other types of computer storage media. The instructions stored on the machine readable medium 404 may include various sub-instructions for executing the function embodied by the identified functions.The one or more hardware processors 402 may execute the instruction 406 to group a plurality of polling in a common polling array. As described above, and similar to the placement of threads in a shared buffer ring, to allow unused CPUs to steal work, polling agents, e.g., polling agents 240 a- 240 c(FIG. 2B ) may be merged into a shared polling array 240 to be selected by a CPU, e.g., local CPU 200, for execution to dynamically distribute polling threads to the CPUs. Each poller in the shared poller array may be assigned an atomic running flag.The use of an atomic run tag allows for the random selection of a bollard, and a CAS operation can be used to attempt to set the run tag of a selected bollard atomically. The plurality of hardware processors 402 may execute the instruction 408 to randomly select a first bollard from the plurality of bollards and attempt to execute the first bollard. If the CAS operation is successful and the CPU can request the first polling for execution. As described above, a CPU may attempt to execute a polling between scheduling execution of threads.Thus, the one or more hardware processors 402 may execute the instruction 410 to execute the first polling after the first polling is successfully requested for execution. After execution of the first polling, the CPU may enable the first polling so that it may be randomly selected for execution by, e.g., another CPU.On the other hand, if the CAS operation fails and the first poller cannot be requested for execution, the CPU has two possibilities, in accordance with some embodiments. That is, the one or more hardware processors 402 may execute the instruction 412 to either randomly select a second polling from the plurality of polling and attempt to execute the second polling, or the one or more hardware processors 402 may return to scheduling a thread to be executed. Again, the thread to be executed may be a next or subsequent thread, recalling that the polling architecture is nested between thread scheduling as disclosed above.FIG. 5 shows a block diagram of an example computer system 500 in which various embodiments described herein may be implemented. Computer system 500 includes a bus 502 or other communication mechanism for communicating information, one or more hardware processors 504 coupled to bus 502 for processing information. The hardware processor(s) 504 may be, for example, one or more general purpose microprocessors.The computer system 500 also includes a main memory 506, such as random access memory (RAM), a cache, and / or other dynamic storage devices, coupled to the bus 502 for storing information and instructions to be executed by the processor 504. Main memory 506 may also be used to store temporary variables or other intermediate information during execution of instructions to be executed by processor 504. Such instructions, when stored in storage media accessible by the processor 504, make the computer system 500 a special purpose machine adapted to perform the operations specified in the instructions.Computer system 500 also includes read only memory (ROM) 508 or other static storage device coupled to bus 502 for storing static information and instructions for processor 504. A storage device 510, e.g., a magnetic disk, optical disk, or USB stick (flash drive), etc., is provided and coupled to bus 502 to store information and instructions.In general, the word "component", "system", "database", and the like, as used herein, may refer to logic embodied in hardware or firmware, or to a collection of software instructions that may have entry and exit points and are written in a programming language such as Java, C, or C++. A software component may be compiled and linked into an executable program, installed in a dynamic link library, or written in an interpreted programming language such as BASIC, Perl, or Python. Software components may be callable from other components or from themselves and / or may be invoked in response to detected events or interrupts. Software components configured for execution on computing devices may be provided on a computer readable medium, such as a compact disc, a digital video disc, a flash drive, a magnetic disc, or other tangible medium, or as a digital download (and may be originally stored in a compressed or installable format that requires installation, decompression, or decryption prior to execution). Such software code may be partially or completely stored on a storage device of the executing computing device for execution by the computing device. Software instructions may be embedded in firmware such as an EPROM. It is understood that hardware components may consist of connected logic units such as gates and flip-flops and / or may be composed of programmable units such as programmable gate arrays or processors.The computer system 500 may implement the techniques described herein using custom hard-wired logic, one or more ASICs or FPGAs, firmware, and / or program logic that, in combination with the computer system, makes or programs the computer system 500 a special-purpose machine. According to one embodiment, the techniques described herein are performed by the computer system 500 in response to the processor(s) 504, executing / executing one or more sequences of one or more instructions contained in the main memory 506. Such instructions may be read into main memory 506 from another storage medium, such as storage device 510. Execution of the sequences of instructions contained in main memory 506 causes processor(s) 504 to perform the process steps described herein. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions.The term "non-transitory media" and similar terms as used herein refer to any media that stores data and / or instructions that cause a machine to operate in a particular manner. Such non-transitory media may include non-transitory media and / or volatile media. Non-volatile media includes, for example, optical or magnetic disks, such as storage device 510. The volatile media includes dynamic memory, such as main memory 506. Common forms of non-volatile media include, for example, a floppy disk, a flexible disk, a hard disk, a solid state drive, magnetic tape or other magnetic data storage medium, a CD-ROM, other optical data storage medium, any physical medium with patterns of holes, a RAM, a PROM and EPROM, a FLASH-EPROM, NVRAM, any other memory chip or cartridge, and networked versions thereof.Non-transitory media are different from, but may be used in conjunction with, transmission media. Transmission media are involved in the transmission of information between non-transitive media. Transmission media includes, for example, coaxial cables, copper wires, and optical fibers, including the wires making up bus 502. Transmission media may also take the form of acoustic or light waves, such as those generated in radio wave and infrared data communication.As used herein, the term "or" may be understood in both an inclusive and an exclusive sense. Moreover, the description of resources, acts, or structures in the singular is not to be understood as excluding the plural. Conditional terms such as "may", "could", "could" or "permitted" unless expressly stated otherwise or otherwise understood in the context, are generally to be understood such that certain embodiments include certain features, elements and / or steps, while other embodiments do not include these. Terms and expressions used in this document and variations thereof should be understood as open and not restrictive unless expressly stated otherwise. The term "including" should be taken as an example in the sense of "including, without limitation", or the like. The term "example" is used to give exemplary examples of the subject matter discussed, not as an exhaustive or limiting list thereof. The terms "a" or "an" are to be understood in the sense of "at least one", "one or more" or the like. The presence of extending words and expressions such as "one or more", "at least", "but not limited to", or other similar expressions in some cases is not to be understood as the narrower case is intended or required if such extending expressions may be missing.
Claims
A central processing unit, CPU (10, 20) comprising: processing circuitry; A controller that extracts instructions from a memory unit (18), the instructions causing the controller to: execute respective threads (210, 212, 214) using processing circuitry maintained in a local run queue (205) associated with the processing circuitry and with a polling, wherein the controller interleaves execution of the respective threads (210, 212, 214) and the polling, randomly selects the polling to be executed after execution of a thread in the local run queue (205) from a shared polling array (240), and wherein execution of the polling by the processing circuitry causes the polling to determine whether the CPU is to execute a task for a device external to the CPU; after clearing the local run queue (205), checking a buffer (215) for threads to execute shared by a group of CPUs comprising the CPU and one or more additional CPUs; taking a stack of threads (220, 222, 224, 226, 228) associated with the group of CPUs from the run queue of the buffer (215); placing the stack of threads (220, 222, 224, 226, 228) into the local run queue (205); and executing, by the processing circuitry, a respective thread in the stack of threads (220, 222, 224, 226, 228) from the local run queue (205).The CPU of claim 1, wherein the CPU and the one or more additional CPUs are located in the same physical core or in the same non-uniform memory access (NUMA) socket.The CPU of claim 2, wherein the instructions further cause the controller to select the stack of threads (220, 222, 224, 226, 228) based on a correspondence with the CPU, physical core, or NUMA socket.The CPU of claim 3, wherein the membership indicates a higher priority for the same NUMA socket compared to the same physical core but not in the same NUMA socket.The CPU of claim 1, wherein the instructions further cause the controller to check the buffer (215) for all remaining threads waiting on the CPU to run and execute the remaining threads before removing the stack of threads (220, 222, 224, 226, 228) from the queue.The CPU of claim 1, wherein the instructions further cause the controller to randomly select a first CPU from the one or more additional CPUs to which the stack of threads (220, 222, 224, 226, 228) is coupled.The CPU of claim 6, wherein the instructions further cause the controller to repeatedly perform random selection of a subsequent CPU of the one or more additional CPUs in response to the first CPU not being associated with at least one executable thread.The CPU of claim 7, wherein the instructions further cause the controller to enter a sleep state upon reaching a threshold number for attempts for random selection.The CPU of claim 1, wherein the instructions further cause the controller to execute a respective thread of the stack of threads (220, 222, 224, 226, 228) based on a priority order.The CPU of claim 1, wherein the instructions further cause the controller to execute the polling for execution upon successful assertion of the polling from the shared polling array (240).The CPU of claim 10, wherein execution of the poller further causes the controller to clear a running flag associated with the poller.The CPU of claim 11, wherein the instructions further cause the controller to clear the run flag when the polling is requested to be executed due to a successful compare and swap operation.The CPU of claim 1, wherein the instructions further cause the controller to randomly select a second polling from the shared polling array (240) to be executed or determine a subsequent thread for execution.The CPU of claim 1, wherein the instructions further cause the controller to increment a delay order counter associated with a respective polling of the shared polling array (240) in response to the selected polling not finding a thread to execute.The CPU of claim 14, wherein the instructions further cause the controller to maintain a sequence number that is incrementable by the controller each time the controller randomly selects a polling from the shared polling array (240), such that the controller has a preference for polling with a smaller value for a corresponding delay order.The CPU of claim 15, wherein the instructions further cause the controller to ignore the preference until the delay order counter reaches a maximum delay order value.The CPU of claim 1, wherein randomly selecting a bollard from the bollard array (240) comprises generating a random index in the common bollard array (240) and a CAS instruction to attempt to claim a bollard at that index.
Citation Information
Patent Citations
Multi-dimensional thread grouping for multiple processors
US20090307704A1
Method and system for scheduling a thread in a multiprocessor system
US20110004882A1