Adaptive Thread Management for Heterogeneous Computing Architectures

A dynamic thread scheduling method for heterogeneous computing architectures optimizes thread assignment by using performance counters to reassign tasks from high-performance to low-power cores, addressing inefficiencies in existing scheduling methods and improving power efficiency.

JP2025522498APending Publication Date: 2025-07-15ADVANCED MICRO DEVICES INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024574607
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-22
Filing Date
2023-05-03
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Existing scheduling methods for heterogeneous computing architectures, which include architecturally compatible cores with varying throughput and power consumption, lead to performance limitations or excessive power consumption due to mismatches between scheduling methods and system resources.

Method used

A dynamic thread scheduling mechanism that utilizes hardware performance counters to classify thread behavior and reassign threads from high-performance cores to low-power cores based on measured metrics, optimizing resource utilization and power efficiency.

Benefits of technology

Enhances performance by preferentially assigning non-scalable, I/O-bound, or memory-bound threads to low-power cores, reducing overall power consumption while maintaining or improving system throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025522498000001_ABST
    Figure 2025522498000001_ABST
Patent Text Reader

Abstract

An apparatus and method are provided for dynamically and efficiently scheduling tasks to a plurality of cores that support heterogeneous computing architectures. The computing system includes a plurality of cores having at least two cores that are capable of executing instructions of the same instruction set architecture (ISA) and are thus architecturally compatible. In one embodiment, each of the at least two cores is a general-purpose central processing unit (CPU) core capable of executing instructions of the same ISA. However, throughput and power consumption vary significantly between at least two cores based on their hardware designs. The operating system scheduler assigns a thread to a first core, and the first core measures the thread's dynamic behavior over a time interval. Based on the thread's dynamic behavior, the scheduler reassigns the thread to a second core that is different from the first core.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] In some designs, the microarchitecture of each of the first core and the second core can execute a particular type of task. For example, each of the first core and the second core can execute instructions of the same instruction set architecture (ISA), and thus is architecturally compatible. The first core and the second core have permitted access to the same region of memory and can thus swap workloads between them. However, the hardware design of the first core emphasizes high performance (high throughput), which also increases power consumption. In contrast, the hardware design of the second core emphasizes low power consumption.

[0002] The first core can use a general-purpose microarchitecture and can include hardware that emphasizes high performance rather than low power consumption. In contrast, the second core includes hardware that emphasizes low power consumption rather than high performance. For example, the first core utilizes hardware of standard cells from a cell library focused on high performance (high throughput), such as using transistors having a shorter channel length, higher doping concentrations in the source and drain regions, a thinner gate oxide thickness, etc. In contrast, the second core utilizes hardware of standard cells from a cell library focused on low power consumption, such as using transistors having a longer channel length, lower doping concentrations in the source and drain regions, a larger gate oxide thickness, etc. In some designs, the first core can use more hardware than the second core to support the issue, dispatch, execution, and retirement of multiple instructions, the simultaneous determination of data transfers for multiple instructions per clock cycle, extra routing and logic, complex branch prediction schemes, support for deep pipelines, support for simultaneous multithreading, access to a relatively large cache, and other design features. In contrast, the second core uses the same general-purpose microarchitecture but includes significantly less hardware than the enumerated hardware of the first core. When an architecture includes at least two architecturally compatible cores but there are significant differences in throughput and power consumption between them, that architecture is referred to as a "heterogeneous computing architecture" or a "heterogeneous processing architecture".

[0003] The above labeling of the architecture should not be confused with a "heterogeneous architecture" that includes at least two cores that are architecturally incompatible, which is such that at least two cores use different microarchitectures such as a general-purpose microarchitecture and a relatively extensive single instruction multiple data (SIMD) microarchitecture. To schedule the workload executed on a computer system having a heterogeneous computing architecture, an operating system (OS) scheduler uses a round-robin method or a method based on core availability. However, these scheduling methods limit performance or consume a significant amount of power if there is a mismatch between the scheduling method and the system resources.

[0004] In view of the above, an efficient method and mechanism for efficient dynamic scheduling of tasks to multiple cores are desired.

Brief Description of the Drawings

[0005]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

DETAILED DESCRIPTION OF THE INVENTION

[0006] Although the present invention has various modifications and alternative forms, specific embodiments are shown by way of example in the drawings and are described in detail herein. However, it should be understood that the drawings and their detailed description are not intended to limit the present invention to the particular forms disclosed, but on the contrary, the present invention is intended to cover all modifications, equivalents, and alternatives falling within the scope of the present invention as defined by the appended claims.

[0007] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, one of ordinary skill in the art should recognize that the present invention may be practiced without these specific details. In some instances, well-known circuits, structures, and techniques are not shown in detail to avoid obscuring the present invention. Further, for simplicity and clarity of explanation, it should be understood that the elements shown in the figures are not necessarily drawn to scale. For example, the dimensions of some elements are exaggerated relative to other elements.

[0008] An apparatus and method for efficiently scheduling tasks in a dynamic manner to a plurality of cores that support heterogeneous computing architectures are contemplated. In various embodiments, a computing system includes a plurality of processor cores (or cores) that support heterogeneous computing architectures. At least two of the plurality of cores are capable of executing instructions of the same instruction set architecture (ISA) and are thus architecturally compatible. At least two cores have permitted access to the same region of memory and are thus able to swap workloads therebetween. However, the hardware design of one of the at least two cores emphasizes high performance (high throughput), which also increases power consumption. In contrast, the hardware design of the other of the at least two cores emphasizes low power consumption. In one embodiment, each of the at least two cores is a general-purpose central processing unit (CPU) core capable of executing instructions of the same instruction set architecture (ISA).

[0009] If an architecture includes at least two cores that are architecturally compatible, but there are significant differences in throughput and power consumption between them, then that architecture is referred to as a "heterogeneous computing architecture" or a "heterogeneous processing architecture". Such a type of architecture should not be confused with a "heterogeneous architecture" that includes at least two cores that are not architecturally compatible such that the at least two cores use different microarchitectures. An example of a heterogeneous architecture is a computing system having a first core such as a CPU core that includes a general-purpose microarchitecture and a second core such as a graphics processing unit (GPU) core that includes a relatively extensive single instruction multiple data (SIMD) microarchitecture. It should be noted that a computing system can include both a heterogeneous computing architecture (or heterogeneous processing architecture) and a heterogeneous architecture simultaneously. For example, in one embodiment, a computing system includes at least two CPU cores that are architecturally compatible, but the throughput and power consumption are significantly different between them. Also, this computing system includes at least one GPU core or other core that uses a relatively extensive SIMD microarchitecture.

[0010] The computing system uses at least a multi-core heterogeneous computing architecture. The hardware such as the circuitry of a specific core among the multiple cores of the computing system executes the instructions of the operating system scheduler. This specific core can be a general-purpose CPU core. Over time, it should be noted that one or more cores execute the instructions of the operating system scheduler. For example, the hardware of one or more cores such as CPU cores individually executes the instructions of the kernel at different times and accesses one or more specific regions of the shared memory. Any of the shared data structures maintained by one or more cores executing the operating system kernel is a list of threads ready to execute. Therefore, at a later time, it is possible and contemplated that the hardware of a core different from the specific core executes the instructions of the operating system scheduler. As used herein, the terms "scheduler" and "scheduling core" refer to one or both of the software and hardware that execute to schedule threads for execution. In some embodiments, the core that acts as a scheduler or otherwise executes scheduling software is referred to as a scheduling core. For example, any of the multiple cores may be executing the instructions of the operating system scheduler. The scheduler assigns a thread to a first core among the multiple cores of the computing system. The first core can be any of the multiple cores of the computing system. The scheduler assigns the thread to the first core based on the availability of the multiple cores, the round-robin scheduling method, or another initial scheduling method. The first core determines the classification of the thread based on the hardware performance metric measured over a certain time interval. In various embodiments, the first core includes a hardware performance counter that monitors the hardware events occurring within the first core while the first core is executing the thread assigned to it.In some embodiments, the first core includes one or more multi-bit registers used as hardware performance counters that can count multiple hardware-related activities.

[0011] The first core compares the value stored in one or more of the hardware performance counters with a threshold value to determine the dynamic behavior of the assigned thread. The number of events reaching at least a predetermined threshold value is used to determine the classification of the assigned thread. Examples of classifications are non-scalable threads, input / output (I / O) or memory-bound threads, vector threads, and the like. The first core communicates the classification, which is an indicator of the thread's dynamic behavior, to the scheduler. As previously explained, the scheduler is any one of the multiple cores that execute the instructions of the current operating system scheduler. For communication, in some embodiments, the first core writes a specific value (or values) to a predetermined configuration register, such as a machine-specific register (MSR) assigned to the first core. The scheduler checks, inspects, or polls this predetermined configuration register to determine the value stored in the predetermined configuration register. Based on the stored value previously written by the first core, the scheduler determines the classification, which is an indicator of the dynamic behavior of the thread determined by the first core. The scheduler reassigns the thread to a second core different from the first core based at least in part on the indicator of the thread's dynamic behavior.

[0012] In one embodiment, the classification of the thread indicates that the thread is part of a process that executes a background task. However, the scheduler first assigns the thread to a first core, which is a high-performance CPU core, based on availability, round-robin, etc. The first core uses a hardware performance counter to classify the thread as a non-scalable thread or another type of thread that is preferably preferentially assigned to a low-power core. The first core communicates the classification to a scheduling core (the core that runs the OS scheduler), and the scheduling core removes this thread from the high-performance core and reassigns this thread to a low-power core. It should be noted that as used herein, a "high-performance" core refers to a core designed to operate at a higher performance and / or power level than a "low-power" core. Further details of efficiently scheduling tasks in a dynamic manner to a plurality of cores that support heterogeneous computing architectures are provided in the following description.

[0013] Referring to FIG. 1, a generalized block diagram of a computing system 100 that efficiently schedules threads on cores that support heterogeneous computing architectures is shown. The computing system 100 includes a semiconductor chip 110 and a memory 130. The semiconductor chip 110 (or chip 110) includes multiple types of integrated circuits. For example, the chip 110 includes cores within a processing unit 150, and at least multiple processor cores (or cores) such as cores 112 and 116 of processing units (units) 115 and 119. There are various options for arranging the circuits of the chip 110 in system packaging to integrate multiple types of integrated circuits. Some examples are system-on-a-chip (SOC), multi-chip module (MCM), and system-in-package (SiP). A clock source, for example, a phase lock loop (PLL), an interrupt controller, a power controller, an interface for input / output (I / O) devices, etc. are not shown in FIG. 1 for ease of explanation.

[0014] The chip 110 includes a memory controller 120 for communicating with the memory 130, an interface logic 140, processing units (units) 115 and 119, a communication fabric 160, a shared cache memory subsystem 170, and a shared processing unit 150. The processing unit 115 includes a core 112 and a corresponding cache memory subsystem 114. The processing unit 119 includes a core 116 and a corresponding cache memory subsystem 118. The memory 130 is shown to include operating system code 132. Note that various portions of the operating system code 132 can reside in the memory 130, one or more caches (114, 118), and can be stored in a non-volatile storage device such as a hard disk (not shown). In one embodiment, the described functions of the chip 110 are incorporated on a single integrated circuit.

[0015] Two units 115 and 119 are shown, each having a single core (cores 112 and 116 respectively), but chip 110 is contemplated to include a different number of four units, each of which can include a different number of cores, based on design requirements. Similarly, while a single processing unit 150 is shown, in other embodiments the chip includes a different number of other processing units that can communicate with other components of chip 110 via communication fabric 160. In various embodiments, cores 112 and 116 can execute one or more threads and share at least a shared cache memory subsystem 170, a processing unit 150, and a combined input / output (I / O) device connected to an interface (IF) 140.

[0016] In various embodiments, the two cores 112 and 116 can execute instructions of the same instruction set architecture (ISA) and are thus architecturally compatible. The two cores 112 and 116 have permitted access to the same regions of memory and can thus swap workloads between them. In one embodiment, each of the two cores 112 and 116 is a general-purpose central processing unit (CPU) core that can execute instructions of the same ISA. However, in one embodiment, the hardware design of core 112 emphasizes high performance (high throughput), which also increases power consumption. In contrast, the hardware design of core 116 emphasizes low power consumption. In some embodiments, the high-performance core 112 utilizes hardware of standard cells from a cell library focused on high performance (high throughput), such as using transistors having a shorter channel length, higher doping concentrations in the source and drain regions, a thinner gate oxide thickness, etc. In contrast, the low-power core 116 utilizes hardware of standard cells from a cell library focused on low power consumption, such as using transistors having a longer channel length, lower doping concentrations in the source and drain regions, a larger gate oxide thickness, etc. It is also possible and contemplated that the high-performance core 112 includes hardware such as circuitry for issuing, dispatching, executing, and retiring multiple instructions. The first core may use more hardware than the second core to support extra routing and logic for simultaneously determining data transfers for multiple instructions per clock cycle, complex branch prediction schemes, support for deep pipelines, support for simultaneous multithreading, access to a relatively large cache, and other design features.

[0017] In contrast to the high-performance core 112, the low-power core 116 can and is intended to include significantly fewer of the above-listed hardware of the high-performance core 112. Accordingly, the low-power core 116 consumes considerably less power but also has a throughput much lower than that of the high-performance core 112. When an architecture includes at least two cores that are architecturally compatible but have significantly different throughputs and power consumptions between them, that architecture is referred to as a "heterogeneous computing architecture" or a "heterogeneous processing architecture". Such a type of architecture should not be confused with a "heterogeneous architecture" that includes at least two cores that are not architecturally compatible such that at least two cores use different microarchitectures. An example of a heterogeneous architecture is a computing system having a first core such as the high-performance CPU core 112 and a second core such as the core of the processing unit 150. In one embodiment, one or more cores of the processing unit 150 use a relatively wide single instruction multiple data (SIMD) microarchitecture. Other types of data processing and corresponding microarchitectures are possible and are intended.

[0018] Note that the computing system 100 can include a heterogeneous computing architecture (or a heterogeneous processing architecture) and a heterogeneous architecture at the same time. For example, in various embodiments, the computing system 100 includes a high-performance CPU core 112 and a low-power CPU core 116 that are architecturally compatible, but the throughput and power consumption are significantly different between them. Also, the computing system 100 includes one or more cores of the processing unit 150 that use a microarchitecture different from the general-purpose CPU microarchitecture. An example of a different microarchitecture is a wide SIMD microarchitecture.

[0019] As described above, the hardware of a specific core among a plurality of cores, such as a circuit, executes the instructions of the current operating system scheduler, and this specific scheduling core assigns threads to the high-performance core 112. The scheduling core assigns threads to the high-performance core 112 based on the availability of a plurality of cores within the computing system 100, the round-robin scheduling method, or another initial scheduling method. It is possible and contemplated that other cores among the plurality of cores include hardware that executes the instructions of the operating system scheduler at other times. The high-performance core 112 determines the classification of threads based on hardware performance metrics measured over a certain time interval. In various embodiments, the high-performance core 112 includes a hardware performance counter that monitors hardware events occurring within the high-performance core 112 while the first core is executing the assigned threads. In some embodiments, the high-performance core 112 includes one or more multi-bit registers used as hardware performance counters that can count a plurality of hardware-related activities.

[0020] Alternatively, the counter counts the number of clock cycles spent by the high-performance core 112 in performing a given event. Examples of events include pipeline flushes, data cache snoops and snoop hits, cache and translation lookaside buffer (TLB) misses, read and write operations, data cache line write-backs, branch operations, branch executions, the number of instructions in the integer or floating-point pipeline, and bus utilization. Several other events are possible and contemplated. In addition to storing the absolute numbers corresponding to hardware-related activities, the counter of the high-performance core 112 determines and stores relative numbers such as the percentage of cache read operations that hit in the cache. In addition to the hardware performance counter, in one embodiment, the high-performance core 112 includes a timestamp counter that is used to determine the time rate or frequency of hardware-related activities. For example, the first core uses the timestamp counter to store and update the number of cache read operations per second, the number of pipeline flushes per second, the number of floating-point operations per second, and so on. In various embodiments, the low-power core 116 includes one or more of these types of registers that are used to implement a hardware performance counter.

[0021] The high-performance core 112 compares the values stored in one or more of the hardware performance counters to a threshold value to determine the dynamic behavior of the assigned thread. The number of events that reach at least a given threshold value is used to determine the classification of the assigned thread. Examples of classifications are non-scalable threads, input / output (I / O) or memory-bound threads, vector threads, and the like. The high-performance core 112 communicates the classification, which is an indicator of the thread's dynamic behavior, to the scheduling core. In one embodiment, the high-performance core 112 writes the indicator to a given configuration register that is later read by the scheduling core.

[0022] The scheduling core can remove threads from the high-performance cores 112 and reassign them to the low-power cores 116, at least partially based on metrics of the dynamic behavior of the threads. In various embodiments, note that one or more of the high-performance cores 112 and the low-power cores 116 include hardware capable of executing multiple threads at a particular point in time due to supporting simultaneous multi-threading (SMT). If the high-performance core 112 supports SMT for four threads, the scheduling core uses four individual logical processor identifiers (IDs) for thread assignment and thread reassignment. However, the scheduling core does not reassign a thread from a first logical processor ID of the high-performance core 112 to a second logical processor ID of the high-performance core 112 based on metrics of the dynamic behavior of the thread.

[0023] In one embodiment, the above classification of threads indicates that the thread is part of a process that executes video playback, a process that executes input / output (I / O) bound or memory bound tasks, or a process that executes a pause or spin loop (busy spin or wait spin). However, the scheduling core initially assigns threads to the high-performance cores 112 based on availability, round-robin, etc. Threads of these types are preferably preferentially assigned to the low-power cores 116. The high-performance core 112 communicates the classification (metrics of dynamic thread behavior) to the scheduling core, and the scheduling core removes this thread from the high-performance core 112 and reassigns this thread to the low-power core 116.

[0024] In other embodiments, the reassignment is from the low-power core 116 to the high-performance core 112. For example, if the scheduling core first assigns a thread of a real-time process such as a real-time audio / visual (A / V) process of a video game to the low-power core 116 based on availability, round-robin, etc., after a certain time interval, the low-power core 116 sends an indicator of dynamic thread behavior to the scheduling core. Based on this indicator, the scheduling core removes the real-time thread from the low-power core 116 and reassigns the real-time thread to the high-performance core 112. Before continuing with more details of efficiently scheduling tasks in a dynamic manner to multiple cores that support heterogeneous computing architectures, further description of the components of the computing system 100 is provided.

[0025] Interface 140 generally provides an interface for various types of input / output (I / O) devices remote from chip 110 to shared cache memory subsystem 170 and processing unit 115. Generally, interface logic 140 includes buffers for receiving packets from corresponding links and buffering packets to be transmitted on corresponding links. Any suitable flow control mechanism can be used to transmit packets between chip 110. Memory 130 can be used as system memory for chip 110 and can include any suitable memory device such as one or more RAMBUS dynamic random access memories (DRAMs), synchronous DRAMs (SDRAMs), DRAMs, static RAMs, etc.

[0026] The address space of chip 110 is divided among multiple memories corresponding to multiple cores. In one embodiment, the coherence point of an address is memory controller 120 that communicates with the memory storing the byte corresponding to the address. Memory controller 120 includes control circuitry for interfacing with the memory and a request queue for queuing memory requests. Generally speaking, communication fabric 160 responds to control packets received on the links of IF140, generates control packets in response to cores 112 and 116 and / or cache memory subsystems 114 and 118, generates probe commands and response packets in response to transactions selected by memory controller 120 for servicing, and routes packets to other nodes through interface logic 140. The communication fabric supports various packet transmission protocols and includes one or more of a system bus, packet processing circuitry and packet selection arbitration logic, and queues for storing requests, responses and messages.

[0027] Cache memory subsystems 114 and 118 include relatively high-speed cache memories for storing blocks of data. Cache memory subsystem 114 can be integrated within respective high-performance cores 112. Alternatively, cache memory subsystem 114 can be connected to processor cores 112 in a backside cache configuration or an in-line configuration as needed. Cache memory subsystem 114 can be implemented as a cache hierarchy. In one embodiment, cache memory subsystem 114 represents an L2 cache structure and shared cache subsystem 170 represents an L3 cache structure. Cache memory subsystem 118 can be implemented in any of the manners described above for cache memory subsystem 114.

[0028] Referring to FIG. 2, another generalized block diagram of a computing system 200 is shown that efficiently schedules threads on a core that supports heterogeneous computing architectures. The computing system 200 includes a microcontroller 210, an assigned processor core 220 (or core 220), a classification table 230, and a scheduler core 240. Although various references are made herein to a "table" or "tables", it should be noted that in various embodiments, the data / content identified as being stored by such a table can be stored in any of various forms other than the table itself. For example, various data structures (trees, indexed, relational searches, etc.) allocated in memory can be used to store the data. All such embodiments are possible and contemplated. A clock source, e.g., a phase lock loop (PLL), an interrupt controller, a power controller, an interface for input / output (I / O) devices, etc., are not shown in FIG. 2 for ease of explanation. Additionally, in various embodiments, the computing system 200 includes other cores in addition to the assigned core 220 and the scheduler core 240, and the computing system supports a multi-core heterogeneous computing architecture. In some embodiments, the assigned core 220, the scheduler core 240, and other cores are general-purpose CPU cores. In another embodiment, the functionality of the microcontroller 210 is implemented by a core or another integrated circuit.

[0029] The assigned core 220 includes a thread class table 222 (or table 222) that maps a class identifier (ID) or an enumerated value (Enum) to a specific thread cluster type. As shown, examples of thread cluster types or classifications are non-scalable threads, input / output (I / O) or memory-bound threads, vector threads, etc. Non-scalable threads are from processes that execute a pause or a spin loop (busy spin or wait spin). In one embodiment, non-scalable threads are preferably assigned preferentially to low-power cores. In some embodiments, threads from processes that execute input / output (I / O) bound or memory-bound tasks are also preferably assigned preferentially to low-power cores. Threads from processes that execute vector and / or floating-point operations are preferably assigned preferentially to high-performance cores. Threads that are not identified as being in any other class are identified as a default thread cluster type (or default classification). In one embodiment, default threads are preferably assigned preferentially to high-performance cores. The "HP" indicator designates that thread cluster types 0, 1 are preferably assigned preferentially to high-performance cores, and the "LP" indicator designates that thread cluster types 2, 3 are preferably assigned preferentially to low-power cores. Other assignments and thread cluster types are possible and contemplated.

[0030] In one embodiment, the microcontroller 210 executes firmware instructions implementing an algorithm for generating and updating values stored in the classification table 230 (or table 230). The table 230 stores information mapping the cores of the computing system 200 participating in heterogeneous computing architectures to identified thread cluster types. The mapping includes a ranking that is a weight value determining how strongly a particular thread cluster type is prioritized for assignment to a high-performance core or a low-power core. The cores are identified by core identifiers (IDs) shown as "c0", "c1", etc. The thread cluster types use the same identifiers as those used in table 222.

[0031] The ranking or weight of the mapping of the assignment between core "c0" and thread cluster type 0 is shown as "eff_0". This ranking "eff_0" indicates that thread cluster type 0 is preferably not preferentially assigned to a low-power core but rather to a high-performance core. Here, the low-power cores are shown as "e" and "eff" of the efficient cores, and the high-performance cores are shown as "p" and "perf" of the performance cores. The ranking or weight regarding the mapping of the assignment between core "c0" and thread cluster type 0 is shown as "perf_3". Here, the larger the numerical value, the higher the priority of the assignment. Therefore, this ranking "perf_3" indicates that thread cluster type 0 is strongly prioritized to be assigned to a high-performance core such as core "c0". Here, a ranking or weight of "0" indicates a strong avoidance of making the corresponding assignment between the core and the thread cluster type. Other value types, ranges, and indicators for ranking in table 230 are possible and contemplated.

[0032] In some embodiments, hardware such as the circuitry of microcontroller 210 sends table update 212 to a specific region of memory that stores table 230. Copies of tables 222 and 230 are stored in various types of data storage areas. For example, copies of the data stored in tables 222 and 230 are stored in one or more of registers, flip-flop circuits, any of various random access memories, content addressable memory (CAM), etc. Also, copies of the contents of tables 222 and 230 are stored in a hard disk, solid state drive, read only memory (ROM), etc. When microcontroller 210 generates or later updates the content stored in table 230, microcontroller 210 sends notification 214 to scheduler core 240. In some embodiments, notification 214 is an interrupt sent from microcontroller 210 to scheduler core 240. When scheduler core 240 receives notification 214, scheduler core 240 retrieves the latest copy of the updated content and updates its own local copy of table 230.

[0033] Microcontroller 210 determines whether a system event has occurred in computing system 200. Examples of system events are detection of a throttling state notification from a power manager due to a high heat measurement value exceeding a threshold, notification of a specific type of application being executed such as a real-time audio / visual application, etc. When microcontroller 210 determines that a system event has occurred, microcontroller 210 updates the content within table 230 and sends notification 214 to scheduler core 240.

[0034] The scheduler core 240 executes the instructions of the operating system scheduler and assigns threads to the assigned cores 220 based on availability, round-robin, etc. At this point, the scheduler core 240 does not recognize the dynamic thread behavior of the assigned threads. In various embodiments, the assigned core 220 includes a hardware performance counter that monitors hardware events occurring within the assigned core 220 while the assigned core 220 is executing the assigned thread. After a certain time interval has elapsed, the assigned core 220 compares the value stored in one or more of the hardware performance counters with a threshold value to determine the dynamic behavior of the assigned thread. The assigned core 220 accesses the table 222 to perform a classification of the dynamic behavior of the assigned thread. The assigned core 220 reports the dynamic thread behavior to the scheduler core 240. In some embodiments, the assigned core 220 writes a specific value (or values) indicating the thread classification 224 to a predetermined configuration register such as a machine-specific register (MSR) assigned to the assigned core 220. The scheduler core 240 polls this configuration register periodically and reads the stored value. The thread classification 224 provides scheduling feedback to the scheduler core 240. This scheduling feedback depends on the measured dynamic behavior of the threads executing on the assigned core 220.

[0035] The scheduler core 240 accesses a copy of table 230 that has information regarding the identifier of the assigned core 220 and the received thread classification 224. In some embodiments, the scheduler 240 first assigns a thread to the assigned core 220 and then verifies whether enough time has elapsed for an accurate thread classification to occur. If so, the scheduler core 240 accesses a copy of table 230. If not, the scheduler core 240 sends an indication to the assigned core 220 to perform the classification again. Using the rankings (weights) within table 230, the scheduler core 240 determines whether the thread should remain executed on the core 220 to which it was assigned. If not, the scheduler core 240 identifies a different core, other than the assigned core 220, to execute the thread. The scheduler core 240 sends a thread assignment 242 to the core, and the core removes the thread from the assigned core 220 and assigns the thread to execute on the identified different core.

[0036] Referring to FIG. 3, a generalized block diagram of a thread assignment 300 for a computing system that uses cores to support heterogeneous computing architectures is shown. Here, the partitioning of hardware and software resources during the execution of one or more software applications 320, as well as their interrelationships and assignments, are shown. In one embodiment, the circuitry of the scheduling core that executes the operating system 318 allocates regions of memory for processes 308a - 308q. When an application 320 or computer program is executed, each application includes a plurality of processes such as processes 308a - 308j and 308k - 308q. In such an embodiment, each of the processes 308a - 308q owns its own resources, such as a memory image or an instance of instructions and data, prior to application execution. Also, each of the processes 308a - 308q includes process-specific information such as code, data, and an address space that may optionally address a heap and a stack, variables in data and control registers such as stack pointers, general-purpose and floating-point registers, program counters, etc., operating system descriptors such as stdin, stdout, and security attributes such as a processor owner and a process permission set.

[0037] Within each of processes 308a to 308q, there is one or more software threads. For example, process 308a includes software (SW) threads 310a to 310d. Threads execute independently of other threads within their corresponding processes, and threads can execute simultaneously with other threads within their corresponding processes. Generally speaking, each of threads 310a to 310q belongs to only one of processes 308a to 308q. Therefore, for multiple threads of the same process, such as SW threads 310a to 310d of process 308a, the same data content of a memory line, for example, the line at address 0xff38, can be the same for all threads. This handles the contention of the first thread, for example, SW thread 310a, which writes to the memory line read by the second thread, for example, SW thread 310d, assuming that inter-thread communication is securely performed.

[0038] However, for multiple threads of different processes, such as SW thread 310a within process 308a and SW thread 310e within process 308j, the data content of the memory line associated with address 0xff38 can be different for the threads. However, multiple threads of different processes can access the same data content at a specific address if they share the same portion of the address space. Generally, for a given application, the scheduling core that executes operating system 318 sets up the application's address space, loads the application's code into memory, sets up the program stack, branches to a predetermined location within the application, and starts the execution of the application. Typically, the part of operating system 318 that manages such activities is operating system kernel 312.

[0039] As described above, an application can be split into two or more processes, and the hardware computing system 302 can execute two or more applications. Therefore, there can be several processes running in parallel. The scheduling core that executes the kernel 312 always determines which of the concurrent processes should be allocated to the processor core (or cores). The scheduler 316 within the operating system 318 that can be within the kernel 312 includes deterministic logic for allocating the threads of a process to the cores. Also, the scheduling core that executes the scheduler 316 determines the allocation of a particular thread among the software threads 310a - 310q to a particular core among the hardware processor cores (or cores) 314a - 314g and 314h - 314r within the hardware computing system 302, as further described below.

[0040] In various embodiments, the cores 314a - 314g and 314h - 314r are general - purpose CPU cores that include hardware capable of processing the execution of one or more of the threads 310a - 310q within any of the processes 308a - 308q. The hardware computing system 302 supports heterogeneous computing architectures. Dashed lines are used to separate the high - performance cores 314a - 314g from the low - power cores 314h - 314r. In FIG. 3, there are also dashed lines indicating the assignments, and these do not necessarily indicate direct physical connections. Thus, for example, SW thread 310d can be assigned to the high - performance core 314a. However, later (e.g., after thread re - allocation), the scheduling core removes the SW thread 310d from the high - performance core 314a and re - assigns the SW thread 310d to the low - power core 314h based on the metrics of the dynamic behavior of the SW thread 310d reported by the high - performance core 314a.

[0041] In one embodiment, an identifier (ID) is assigned to each of cores 314a - 314g and 314h - 314r. This hardware thread ID (not shown) is used to assign one of cores 314a - 314g to execute any one of SW threads 310a - 310q. A scheduling core that executes scheduler 316 within kernel 312 processes this assignment. For example, the scheduling core uses the hardware thread ID to assign SW thread 310m to low - power core 314r based on the availability of high - performance cores 314a - 314g and low - power cores 314h - 314r, a round - robin scheduling method, or another initial scheduling method. The scheduling core that executes scheduler 316 performs thread re - assignment based on the dynamic thread behavior reported by high - performance cores 314a - 314g and low - power cores 314h - 314r. For example, after a time interval during which SW thread 310m is executed, low - power core 314r determines the dynamic behavior of SW thread 310m. In one embodiment, low - power core 314r determines the dynamic behavior of SW thread 310m by comparing a value stored in one or more of the hardware performance counters with a threshold value. Low - power core 314r writes a specific value (or values) to a predetermined configuration register such as a machine - specific register (MSR) assigned to low - power core 314r. In one embodiment, the predetermined configuration register is a specific area of memory.

[0042] The scheduler checks, inspects, or polls this predetermined configuration register to determine the value stored in the predetermined configuration register. Based on the stored value previously written by the low-power core 314r, the scheduler determines a classification that is an indicator of the dynamic behavior of the thread determined by the low-power core 314r. Based on this classification indicating the dynamic behavior of the SW thread 310m reported by the low-power core 314r, the scheduling core removes the SW thread 310m from the low-power core 314r and reassigns the SW thread 310m to the high-performance core 314g. To assist with thread migration, user data allocated by one thread is used only by that thread, and data sharing between threads occurs via read-only global variables and fast local messages passed through the scheduling core that executes the thread scheduler 316. Also, the scheduling core processes any changes to the stack pointer, if any.

[0043] In various embodiments, the high-performance cores 314a - 314g and the low-power cores 314h - 314r include hardware performance counters that monitor hardware events occurring during a predetermined time interval while executing the threads to which the high-performance cores 314a - 314g and the low-power cores 314h - 314r are assigned. After a certain time interval has elapsed, the high-performance cores 314a - 314g and the low-power cores 314h - 314r compare the value stored in one or more of the hardware performance counters to a threshold value to determine the dynamic behavior of the assigned threads. The high-performance cores 314a - 314g and the low-power cores 314h - 314r report the dynamic thread behavior to the scheduling core, and the scheduling core determines whether to reassign the thread.

[0044] In one embodiment, the scheduling core prefers to allocate threads of a process that executes video playback, a process that executes an input / output (I / O) bound or memory bound task, or a process that executes a pause or spin loop (busy spin or wait spin) to the low-power cores 314h-314r. In contrast, the scheduling core prefers to allocate threads of a real-time process, such as a real-time A / V video game, to the high-performance cores 314a-314g. However, in some embodiments, the priorities used by the scheduling core are programmable, and the updated values of these priorities are stored in a table or other data storage area accessible by the scheduling core.

[0045] Referring to FIG. 4, a generalized block diagram of a method 400 for allocating threads in a computing system that uses cores that support a heterogeneous computing architecture is shown. For purposes of explanation, the steps in this embodiment (as well as FIGS. 5-6) are shown in order. However, in other embodiments, some steps occur in a different order than shown, some steps are executed simultaneously, some steps are combined with other steps, and some steps do not exist.

[0046] The computing system includes a plurality of cores that support heterogeneous computing architectures. The hardware of a particular general-purpose CPU processor core among the plurality of cores executes the instructions of the operating system scheduler, and this CPU processor core (or scheduler) allocates a thread to a particular processor core among the plurality of processor cores of the computing system (block 402). The scheduler allocates the thread to this particular processor core (or particular core) based on the availability of the plurality of cores, the round-robin scheduling method, or another initial scheduling method. This particular core determines the classification of the thread assigned to it based on the hardware performance metric measured over a certain time interval (block 404). In various embodiments, the particular core includes a hardware performance counter that monitors hardware events occurring within the particular core while the particular core is executing the thread assigned to it. In some embodiments, the particular core includes one or more multi-bit registers used as hardware performance counters that can count a plurality of hardware-related activities.

[0047] Alternatively, the counter counts the number of clock cycles spent by a particular core in performing a given event. Examples of events include pipeline flushes, data cache snoops and snoop hits, cache and TLB misses, read and write operations, data cache line writebacks, branch operations, branch executions, the number of instructions in an integer or floating point pipeline, and bus utilization. Several other events are possible and contemplated. In addition to storing the absolute numbers corresponding to hardware-related activities, the counter for a particular core determines and stores relative numbers such as the percentage of cache read operations that hit in the cache. In addition to the hardware performance counter, in one embodiment, a particular core includes a timestamp counter that is used to determine the time rate or frequency of hardware-related activities. For example, a particular core uses the timestamp counter to store and update the number of cache read operations per second, the number of pipeline flushes per second, the number of floating point operations per second, and so on.

[0048] A particular core compares the value stored in one or more of these hardware performance counters to a threshold value to determine the dynamic behavior of the assigned thread. The number of events reaching at least a predetermined threshold is used to determine the classification of the assigned thread. Examples of classifications are non-scalable threads, input / output (I / O) or memory-bound threads, vector threads, etc. The particular core communicates the classification to the scheduler (block 406). In some embodiments, the particular core writes a particular value (or values) to a predetermined configuration register such as a machine-specific register (MSR) assigned to the particular core. The scheduler verifies whether the particular core has hardware that matches the dynamic behavior of the thread by examining a classification table (block 408).

[0049] The classification table stores a ranking mapping indicating a defined correspondence between a plurality of cores and a plurality of predefined classifications of thread dynamic behavior. In some embodiments, the classification table is generated and updated by another processing unit assigned to cause a particular firmware to be executed on the circuit using another core, a microcontroller, or instructions of an algorithm for determining how and when to update the contents of the classification table among the plurality of cores. The scheduler loads the contents of this classification table into a local cache and accesses the classification table when evaluating the original thread assignment.

[0050] If the scheduler determines that there is a match between the assigned thread and a particular core based on the classification table (condition block 410: "Yes"), the scheduler maintains the thread as being assigned to the particular core (block 412). A match indicates that the original thread assignment is found to be a valid assignment within the classification table. If the scheduler determines that there is no match based on the classification table (condition block 410: "No"), the scheduler assigns the thread to a different core that matches the dynamic behavior of the thread based on the classification table (block 414). The different core is another core among the plurality of cores of a computing system using a multi-core heterogeneous computing architecture, rather than the particular core. The scheduler reassigns the thread from the particular core to a different core. Note that the scheduler may include other decision-making logic of a particular scheduling policy in addition to the steps performed in blocks 408 - 414 of method 400 and is contemplated.

[0051] Referring to FIG. 5, a generalized block diagram of a method 500 for updating information used when allocating threads in a computing system that uses cores supporting heterogeneous computing architectures is shown. The computing system includes a plurality of cores that support heterogeneous computing architectures. The hardware such as the circuitry of a particular processor core (or cores) among the plurality of cores or microcontrollers used in the computing system generates a ranking mapping between the plurality of cores and the types of thread dynamic behaviors (block 502). These ranking mappings are stored in a classification table. The circuitry notifies the operating system scheduler of the mapping (block 504). In some embodiments, a particular core sends an interrupt to another core that executes the operating system scheduler.

[0052] The circuitry determines whether a system event has occurred. Examples of system events include detection of a throttling state notification from a power manager due to a high thermal measurement value exceeding a threshold, notification of a particular type of application being executed, etc. If the circuitry determines that a system event has occurred (conditional block 506: "Yes"), the circuitry updates the ranking mapping in the classification table based on the system event (block 508). For example, the circuitry updates the ranking mapping to indicate that more threads and possibly all threads are allocated to low-power cores rather than high-performance cores in a computing system that uses a heterogeneous computing architecture. The circuitry notifies the scheduler of the updated mapping (block 510). If the circuitry determines that a system event has not occurred (conditional block 506: "No"), the circuitry maintains the ranking mapping in the classification table (block 512).

[0053] Referring to FIG. 6, a generalized block diagram of a method 600 for allocating threads in a computing system that uses cores supporting heterogeneous computing architectures is shown. The computing system includes a plurality of cores that support heterogeneous computing architectures. Hardware such as the circuitry of a particular processor core (or cores) among the plurality of cores determines the quality of service (QoS) of a thread (block 602). In some embodiments, the particular core that determines the QoS value of a thread is the core that executes instructions of the operating system kernel. In other embodiments, a thread (or process) includes an indicator of a QoS value. If a particular core determines that there is no indicator of a QoS value (conditional block 604: "yes"), the particular core allocates the thread to a high-performance core (block 606). In some embodiments, each thread has a QoS value and conditional block 604 is unnecessary. In other embodiments, a check is performed to determine whether the thread has a QoS value assigned thereto as implemented in conditional block 604.

[0054] When a specific core determines that there is an indicator of a QoS value (condition block 604: "No") and that the indicated QoS value is greater than a threshold (condition block 608: "Yes"), the specific core assigns the thread to a high-performance core (block 606). However, when the specific core determines that the indicated QoS value is below the threshold (condition block 608: "No"), the specific core assigns the thread to a low-power core (block 610). It should be noted that in other embodiments, these assignments can be interchanged according to the result of the comparison with the threshold. It should also be noted that the specific core can combine the result of the assignment step of method 600 with the assignment steps of methods 400-500 (in FIGS. 4-5). In yet other embodiments, the specific core calculates a weighted score based on one or more of the results of the assignment step of method 600 and the assignment steps of methods 400-500 (in FIGS. 4-5). For example, when the scheduler reassigns a thread from the core to which it was first assigned to another core, the scheduler uses the QoS value in addition to the weight and classification indicating the dynamic behavior of the thread. Based on the dynamic thread behavior in heterogeneous computing architectures, various combinations and methods are possible and contemplated to determine the final assignment of the thread.

[0055] It should be noted that one or more of the above-described embodiments include software. In such embodiments, the program instructions implementing the method and / or mechanism are carried or stored on a computer-readable medium. A number of types of media configured to store program instructions are available, including hard disks, floppy (registered trademark) disks, CD-ROMs, DVDs, flash memories, programmable ROMs (Programmable ROM, PROM), random access memories (random access memory, RAM), and various other forms of volatile or non-volatile storage devices. Generally speaking, computer-accessible storage media includes any storage media that can be accessed by a computer during use to provide instructions and / or data to the computer. For example, computer-accessible storage media includes magnetic or optical media, such as disks (fixed or removable), tapes, CD-ROMs, DVD-ROMs, CD-Rs, CD-RWs, DVD-Rs, DVD-RWs, or storage media such as Blu-Ray (registered trademark). Storage media further includes volatile or non-volatile memory media such as RAM (e.g., synchronous dynamic RAM (synchronous dynamic RAM, SDRAM), double data rate (DDR, DDR2, DDR3, etc.) SDRAM, low power DDR (LPDDR2, etc.) SDRAM, Rambus DRAM (Rambus DRAM, RDRAM), static RAM (SRAM), etc.), ROM, flash memory, non-volatile memory (e.g., flash memory) accessible via a peripheral interface such as a Universal Serial Bus (USB) interface. Storage media includes microelectromechanical systems (microelectromechanical system, MEMS), as well as storage media accessible via communication media such as networks and / or wireless links.

[0056] Additionally, in various embodiments, the program instructions include an operational level description or a register-transfer level (RTL) description of the hardware functions in a high-level programming language such as C, a design language (HDL) such as Verilog or VHDL, or a database format such as the GDSII stream format (GDS II). In some cases, the description is read by a synthesis tool that synthesizes the description to generate a netlist that includes a list of gates from a synthesis library. The netlist includes a set of gates that also represent the functions of the hardware including the system. The netlist is then placed and routed to generate a data set that describes the geometric shapes to be applied to the mask. The mask is then used in various semiconductor manufacturing steps to generate a semiconductor circuit or circuitry corresponding to the system. Alternatively, the instructions on a computer-accessible storage medium are, as necessary, a netlist (with or without a synthesis library) or a data set. Additionally, the instructions are utilized for emulation by hardware-based types of emulators from vendors such as Cadence®, EVE®, and Mentor Graphics®.

[0057] Although the above embodiments have been described in considerable detail, numerous variations and modifications will become apparent to those skilled in the art upon a full understanding of the above disclosure. The following claims are intended to be construed to embrace all such variations and modifications.

Claims

1. An apparatus, comprising a scheduler, wherein the scheduler is configured to: receive an indicator of the thread dynamic behavior of a predetermined thread assigned to a first core among a plurality of cores; reassign the predetermined thread to a second core different from the first core among the plurality of cores based at least in part on the indicator of the thread dynamic behavior of the predetermined thread; and perform the above operations. An apparatus.

2. The apparatus according to claim 1, wherein the scheduler is configured to examine a ranking mapping indicating a defined match between the plurality of cores and a plurality of classifications of thread dynamic behavior. The apparatus of claim 1.

3. The apparatus according to claim 2, wherein the scheduler is configured to reassign the predetermined thread from the first core to the second core in response to a determination by the ranking mapping that the second core has a defined match with the indicator of the thread dynamic behavior. The apparatus of claim 2.

4. The apparatus according to claim 2, wherein the scheduler is configured to receive a notification indicating that the ranking mapping has been updated. The apparatus of claim 2.

5. The apparatus according to claim 1, wherein the scheduler is configured to execute an operating system scheduler of a computing system using a multi-core heterogeneous computing architecture. The apparatus of claim 1.

6. The first core executes a thread with a first microarchitecture, and the second core executes a thread with a second microarchitecture different from the first microarchitecture. The apparatus of claim 5.

7. The apparatus according to claim 1, wherein the indicator of the thread dynamic behavior is based on a hardware performance counter of the first core. The apparatus of claim 1.

8. A method, comprising: a plurality of cores execute one or more applications; a predetermined core receives an indicator of the thread dynamic behavior of a predetermined thread of the one or more applications assigned to a first core among the plurality of cores; and the predetermined core reassigns the predetermined thread to a second core different from the first core among the plurality of cores based at least in part on the indicator of the thread dynamic behavior of the predetermined thread. A method.

9. The method according to claim 8, wherein the predetermined core checks a ranking mapping indicating a defined match between the plurality of cores and a plurality of classifications of thread dynamic behavior. The method of claim 8.

10. The method according to claim 9, wherein the predetermined core reassigns the predetermined thread from the first core to the second core in response to a determination by the ranking mapping that the second core has a defined match with an indicator of the thread dynamic behavior. The method of claim 9.

11. The method according to claim 9, wherein the predetermined core receives a notification indicating that the ranking mapping has been updated. The method of claim 9.

12. The method according to claim 8, wherein the predetermined core executes an operating system scheduler of a computing system using a multi-core heterogeneous computing architecture. The method of claim 8.

13. The method according to claim 12, wherein the first core executes a thread with a first microarchitecture, and the second core executes a thread with a second microarchitecture different from the first microarchitecture. The method according to claim 12, wherein the first core executes a thread with a first microarchitecture, and the second core executes a thread with a second microarchitecture different from the first microarchitecture. The method of claim 12.

14. The method according to claim 8, wherein the indicator of the thread dynamic behavior is based on a hardware performance counter of the first core. The method of claim 8.

15. A computing system, comprising: a memory configured to store one or more applications of a workload; a plurality of cores configured to execute the one or more applications, wherein a predetermined core of the plurality of cores: receives an indicator of thread dynamic behavior of a predetermined thread assigned to a first core of the plurality of cores; and reassigns the predetermined thread to a second core different from the first core of the plurality of cores based at least in part on the indicator of the thread dynamic behavior of the predetermined thread. The computing system of claim 15, wherein the predetermined core is configured to check a ranking mapping indicating a defined match between the plurality of cores and a plurality of classifications of thread dynamic behavior. The computing system of claim 15.

16. The computing system of claim 15. The computing system of claim 15, wherein the predetermined core is configured to check a ranking mapping indicating a defined match between the plurality of cores and a plurality of classifications of thread dynamic behavior. The computing system of claim 15.

17. The predetermined core is configured to reassign the predetermined thread from the first core to the second core in response to the ranking mapping determining that the second core has a defined match with the indicator of the thread dynamic behavior. The computing system of claim 16. **Claim 18** The predetermined core is configured to receive a notification indicating that the ranking mapping has been updated. The computing system of claim 16. **Claim 19** The predetermined core is configured to execute an operating system scheduler of a computing system using a multi-core heterogeneous computing architecture. The computing system of claim 15. **Claim 20** The first core executes a thread with a first microarchitecture. The second core executes a thread with a second microarchitecture different from the first microarchitecture. The computing system of claim 19.