Hierarchical asymmetric kernel property detection
By determining and reporting thread count and power efficiency asymmetry at the hierarchy level in a processing system, the resource management inefficiency problem caused by core attribute asymmetry is solved, more efficient thread scheduling and resource management are achieved, and the overall efficiency of the processing system is improved.
Patent Information
- Application Number
- CN202280076409.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-11-19
- Filing Date
- 2022-11-16
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2042-11-16
AI Technical Summary
The asymmetry of core properties in a processing system causes the operating system to inefficiently or incorrectly manage processor resources, affecting application processing efficiency and time, and potentially increasing memory and processing resource usage.
By determining thread count and power efficiency asymmetries within a hierarchical level of a processing device, reporting the asymmetry using a relatively small number of thread identifiers, an operating system defines thread identifiers to efficiently manage resources, reducing resource read time and memory requirements.
This achieves more efficient thread scheduling and resource management, reduces processing time and memory resource consumption, and improves the overall efficiency of the processing system.
Smart Images

Figure CN118696298B_ABST
Abstract
Description
Background Art
[0001] Within a processing system, an operating system uses processor topology information for one or more processors to execute software tasks associated with an application. For example, the operating system uses processor topology information to perform multiple processor resource management practices, such as task and thread scheduling, for software tasks associated with an application. The processor topology information for a processor identifies the hierarchical arrangement of hardware and software resources within the processing system for executing the software tasks of the application. However, kernel attribute asymmetry within the processing system may cause the operating system to inefficiently or incorrectly manage the hardware and software resources of the processor, which may have a negative impact on the processing efficiency and processing time of the application. In addition, such asymmetry may cause the operating system to use additional memory and processing resources to compensate for the asymmetry. BRIEF DESCRIPTION OF THE DRAWINGS
[0002] The present disclosure may be better understood, and its numerous features and advantages made apparent to those skilled in the art by referencing the accompanying drawings.The use of the same reference numbers in different drawings indicates similar or identical items.
[0003] Figure 1 is a block diagram of a processing system for hierarchical asymmetric kernel property detection, according to some embodiments.
[0004] Figure 2 is a block diagram of a processing device configured to determine an asymmetry of one or more kernel attributes in one or more hierarchical levels, according to some embodiments.
[0005] Figure 3 is a flow chart illustrating a method for determining thread count asymmetry across one or more hierarchical levels, according to some embodiments.
[0006] Figure 4 is a flow chart illustrating a method for determining power efficiency asymmetry of one or more hierarchical levels according to some embodiments. DETAILED DESCRIPTION
[0007] The techniques and systems described herein address determining thread count asymmetry within a topology of a processing device and reporting the identified asymmetry via a relatively small number of thread identifiers (e.g., via a single thread identifier for a given enumerated instance of a processing device). Thus, the described techniques and systems support efficient reporting of asymmetry, allowing for more efficient asymmetry management, such as more efficient scheduling of threads executing at a processing device.
[0008] For purposes of illustration, as used herein, a "topology" of a processing device includes the arrangement of the hardware and software resources of the processing device into one or more hierarchical levels. As used herein, a "hierarchical level" includes one or more portions of a processing device that include similar hardware or software resources of the processing device (also referred to herein as "enumerated instances"). For example, a processing device includes a die hierarchy level that includes one or more dies. As another example, a processing device includes a core complex hierarchy level that includes one or more core complexes. Each level is "hierarchical" in that a hierarchical level that includes a larger portion of the processing device (e.g., die level) is a "higher" level than a hierarchical level that includes a smaller portion of the processing device (e.g., core complex level, core level). An enumerated instance of a hierarchical level includes one or more hardware or software resources from other (e.g., lower) hierarchical levels. For example, a die at the die level includes one or more core complexes from the core complex level.
[0009] In some embodiments, the processing device is configured to determine thread count asymmetry for one or more hierarchical levels of the processor based on the number of threads for each enumeration instance of the hierarchical level (i.e., how many threads each enumeration instance has within the hierarchical level). When one or more cores of a processing device in a processing system are downgraded (e.g., disabled), thread count asymmetry may sometimes appear between the enumeration instances of the hierarchical level at one or more hierarchical levels of the processing device. That is, the number of threads for each enumeration instance within the hierarchical level will vary depending on the downgraded cores. In order to provide an indication of these thread count asymmetries to an operating system that manages one or more applications running or attempting to run on the processing device, the operating system of the processing system performs a discovery operation to determine whether any thread count asymmetry exists at any hierarchical level of the processing device. Based on the discovery operation, the operating system determines whether one or more thread count asymmetries exist for each hierarchical level of the processing device. That is, the operating system determines whether each hierarchical level is symmetric or asymmetric with respect to thread count.
[0010] In response to determining thread count asymmetry at a hierarchy level, the operating system defines a thread identifier for each enumeration instance at the hierarchy level to determine an indication of the asymmetry. To define the thread identifiers, the operating system only reads one thread identifier for each enumeration instance from a register. Because each thread identifier includes multiple threads of its enumeration instance, only one thread identifier from each enumeration instance needs to be read to determine the indication of the thread count asymmetry. In this manner, the operating system does not need to read every thread identifier at the hierarchy level, thereby reducing processing time and resources required to determine the indication of the asymmetry.
[0011] The techniques and systems described herein address determining power efficiency asymmetry at one or more hierarchical levels of a processing device based on the number of cores in each enumerated instance operating in various operating modes (e.g., a power efficiency mode and a performance mode). When the number of cores of a processing device operating in a power efficiency mode is different from the number of cores in the performance mode, power efficiency asymmetry may sometimes manifest at one or more hierarchical levels of the processing device between the enumerated instances of the hierarchical level. The processing device determines such power efficiency asymmetry at one or more hierarchical levels of the processing device by, for example, comparing the number of cores operating in a first operating mode or a second operating mode in each enumerated instance of the hierarchical level. In order to provide an indication of such power efficiency asymmetry to an application running on or attempting to run on the processing device, an operating system of the processing system performs a discovery operation to determine whether any power efficiency asymmetry exists at any hierarchical level of the processing device. For example, the operating system provides instructions to the processing device causing the processing device to store data representing power efficiency asymmetry at various hierarchical levels in one or more registers. The operating system then reads the data in the registers to determine whether any power efficiency asymmetry exists at one or more hierarchical levels.
[0012] Figure 1 is a block diagram of a processing system 100 for asymmetric kernel attribute detection according to some embodiments. The processing system 100 includes or has access to a memory 106 or other storage component implemented using non-transitory computer-readable media, such as dynamic random access memory (DRAM). However, in various embodiments, the memory 106 is implemented using other types of memory, including, for example, static random access memory (SRAM), non-volatile RAM, etc. According to various embodiments, the memory 106 includes external memory implemented external to the processing units implemented in the processing system 100. The processing system 100 also includes a bus 112 to support communications between entities implemented in the processing system 100, such as memory 106. Some embodiments of the processing system 100 include other buses, bridges, switches, routers, etc., which are not explicitly described in detail. Figure 1 Shown in.
[0013] In various embodiments, the techniques described herein are used with any of a variety of parallel processors (e.g., vector processors, graphics processing units (GPUs), general-purpose GPUs (GPGPUs), non-scalar processors, highly parallel processors, artificial intelligence (AI) processors, inference engines, machine learning processors, other multi-threaded processing units, etc.), scalar processors, serial processors, or any combination thereof. Figure 1An example of a parallel processor, specifically a graphics processing unit (GPU) 114, is shown according to some embodiments. The GPU 114 renders an image for presentation on a display 120. For example, the GPU 114 renders an object to produce pixel values that are provided to the display 120, which uses the pixel values to display an image representing the rendered object. The GPU 114 implements multiple processor cores 116-1 through 116-N that execute instructions concurrently or in parallel. According to various embodiments, one or more of the processor cores 116 operate as SIMD units that perform the same operation on different sets of data. Although in Figure 1 In the exemplary embodiment shown, three cores (116-1, 116-2, 116-N) are shown, representing N cores, but the number of processor cores 116 implemented in GPU 114 is a matter of design choice. Thus, in other embodiments, GPU 114 may include any number of cores 116. Some embodiments of GPU 114 are used for general-purpose computing. GPU 114 executes instructions (such as program code 108 stored in memory 106) and stores information (such as the results of the executed instructions) in memory 106.
[0014] The processing system 100 also includes a central processing unit (CPU) 102, which is connected to a bus 112 and thus communicates with a GPU 114 and a memory 106 via the bus 112. The CPU 102 implements a plurality of processor cores 104-1 through 104-N that execute instructions concurrently or in parallel. In various embodiments, one or more processor cores 104 operate as SIMD units that perform the same operation on different data sets. Although in Figure 1 In the exemplary embodiment shown, three cores (104-1, 104-2, 104-M) are presented, representing M cores, but the number of processor cores 104 implemented in the CPU 102 is a matter of design choice. Thus, in other embodiments, the CPU 102 may include any number of cores 104. In some embodiments, the CPU 102 and the GPU 114 have equal numbers of cores 104, 116, while in other embodiments, the CPU 102 and the GPU 114 have different numbers of cores 104, 116. The processor cores 104 execute instructions (such as program code 110 stored in the memory 106), and the CPU 102 stores information (such as the results of the executed instructions) in the memory 106. The CPU 102 is also capable of initiating graphics processing by issuing draw calls to the GPU 114. In various embodiments, the CPU 102 implements multiple processor cores that execute instructions concurrently or in parallel (for clarity). Figure 1 not shown).
[0015] The input / output (I / O) engine 118 includes hardware and software that handles input or output operations associated with the display 120, as well as other elements of the processing system 100, such as a keyboard, mouse, printer, external disk, etc. The I / O engine 118 is coupled to the bus 112, allowing the I / O engine 118 to communicate with the memory 106, the GPU 114, or the CPU 102. In the illustrated embodiment, the I / O engine 118 reads information stored on an external storage component 122, which is implemented using a non-transitory computer-readable medium, such as a compact disc (CD), a digital video disc (DVD), etc. The I / O engine 118 can also write information to the external storage component 122, such as processing results of the GPU 114 or the CPU 102.
[0016] In various embodiments, memory 106 includes one or more operating systems 124, each of which includes software configured to manage the hardware and software resources of system 100. Operating systems 124 interact with the hardware and software resources of system 100 so that one or more applications (not shown for clarity) can access the hardware and software resources of system 100. For example, operating system 124 completes one or more system calls on behalf of one or more applications, interrupts one or more applications, performs one or more hardware functions (e.g., memory allocation, input, output) for one or more applications, or any combination thereof, to name a few. In various embodiments, operating system 124 is configured to perform one or more discovery operations to determine one or more hardware and software resources of system 100. Such hardware and software resources include, for example, processing devices (e.g., CPU 102, GPU 114), device hierarchies, processing cores (e.g., cores 104, 116), threads, sockets, dies, complexes, or any combination thereof, to name a few. According to various embodiments, operating system 124 is also configured to perform one or more discovery operations to determine one or more hardware and software resource attributes of the hardware and software resources of system 100. These attributes include, for example, the number of resources (e.g., thread count), the operating mode of the resource (e.g., active, inactive, power efficiency mode, performance mode), the power efficiency of the resource, or any combination thereof, to name a few. In various embodiments, one or more hardware and software resources of system 100 include one or more hierarchical levels. For example, one or more portions of a processing device are arranged in one or more hierarchical levels. In various embodiments, the hierarchical levels include one or more similar enumerated instances (i.e., similar portions) of a device, such as a core, a core complex, a die, and a socket, to name a few. As an example, a processing device (e.g., CPU 102, GPU 114) includes four or more hierarchical levels, such as a core level, a core complex level, a die level, and a socket level, to name a few. In various embodiments, each enumerated instance of a hierarchical level includes one or more portions of a device from a different hierarchical level (e.g., a lower hierarchical level). As an example, the die level includes one or more core complexes, each core complex including one or more cores. According to various embodiments, the discovery operations include instructions for a processing device to store data representing one or more hardware and software resources, hardware and software resource attributes, hierarchy levels, or any combination thereof in one or more memory registers.
[0017] In accordance with various embodiments, the operating system 124 performs one or more discovery operations to determine one or more asymmetries of one or more hierarchical levels of the hardware resources of the system 100. As used herein, "asymmetry" includes one or more enumerated instances (e.g., cores, core complexes, dies) within a hierarchical level having one or more different hardware attributes (e.g., operating modes, thread counts). For example, at a core complex level comprising two core complexes, the asymmetry includes a different thread count (i.e., number of threads) between the two core complexes. As another example, at a die level comprising two dies, the asymmetry includes a first number of cores operating in a power efficiency mode within a first die and a second different number of cores operating in a power efficiency mode within a second die. In various embodiments, the discovery operations include instructions causing the processing device to load data representing the one or more asymmetries and the one or more hierarchical levels into one or more registers. The operating system then determines the one or more asymmetries of the one or more hierarchical levels by reading at least a portion of the data stored in the registers.
[0018] In response to determining one or more asymmetries, the operating system 124 is configured to define a thread identifier for each enumeration instance (e.g., each discrete portion) at a level of the hierarchy. For example, in response to determining one or more asymmetries at the kernel complex level, the operating system 124 is configured to send instructions to cause the processing device to store data representing a thread identifier for each thread in the level of the hierarchy. As used herein, a "thread identifier" includes, for example, data indicating a unique keyword that identifies a thread, an enumeration instance at each level of the hierarchy that includes the thread, the number of threads within each enumeration instance that includes the thread, or any combination thereof. For example, each thread identifier indicates the number of other threads within the enumeration instance. In various embodiments, the operating system 124 sends instructions to cause the processing device to generate and store a thread identifier for each thread at a level of the hierarchy based on one or more displacement values stored in one or more registers. For example, the operating system 124 sends instructions to cause the processing device to store data representing a unique key (e.g., an APICID) identifying each thread in a first register and to store a value indicating a bit shift in a second register that, when applied to the unique key identifying the thread, returns a unique key identifying a higher-level topology, e.g., a core, a core complex, a die, a socket, or any combination thereof. According to various embodiments, the operating system 124 is configured to read only one thread identifier for each enumeration instance at a hierarchy level to determine the representation of asymmetry at the hierarchy level. For example, based on the shift value, the operating system determines one thread identifier for each enumeration instance to read within the register. In this manner, the operating system 124 determines the representation of asymmetry (e.g., how the asymmetry affects the threads) without having to read thread attributes for each thread at the hierarchy level, thereby reducing the processing time required to determine the representation of asymmetry. That is, because each thread identifier indicates the number of threads for each enumeration instance, the operating system 124 only needs to read one thread identifier for each enumeration instance, thereby reducing the processing time required to determine the representation of asymmetry. According to various embodiments, the operating system 124 then allows one or more applications, based on the indication of asymmetry, to access at least a portion of the hardware of the system 100. For example, the operating system 124, based on the indication of asymmetry, allows access to one or more threads of the system 100. As an example, the operating system 124 schedules a software task for the application based on the thread count indicated in the indication of asymmetry.
[0019] Now refer to Figure 2 , presents a block diagram of a processing device 200 having one or more hardware hierarchy levels and one or more asymmetries. In various embodiments, the processing device 200 implements Figure 1 For example, the processing device 200 may be connected to the processing system 100 described in Figure 1In various embodiments, the processing device 200 includes a first hierarchical level that includes, for example, one or more cores 234, each of which is similar to or the same as cores 104, 116. For example, in Figure 2 In the exemplary embodiment shown, the processing device 200 includes a first hierarchical level (eg, a core level) that includes core 0 234-1 through core 15 234-16. Figure 2 In the exemplary embodiment of the present invention, the first hierarchical level of the processing device 200 is presented as having 16 cores (234-1, 234-2, 234-3, 234-4, 234-5, 234-6, 234-7, 234-8, 234-9, 234-10, 234-11, 234-12, 234-13, 234-14, and 234-15), but in other embodiments, the first hierarchical level of the processing device 200 can have any number of cores, for example, 4, 8, 32, 64, 128, or 256 cores, to name a few.
[0020] According to various embodiments, the processing device 200 includes a second hierarchical level (e.g., a core complex level) that includes, for example, one or more core complexes (CCXs). As used herein, a "core complex" includes one or more pairs of cores 234 and one or more caches (not shown for clarity). In various embodiments, the cores 234 within a CCX are connected to each other via one or more caches. That is, the cores 234 within a CCX share one or more caches of the CCX. For example, in Figure 2 In the exemplary embodiment of the present invention, the processing device 200 includes a second hierarchical structure level (e.g., CCX level) including CCX0 232-1 (including cores 234-1 through 234-4), CCX1 232-2 (including cores 234-5 through 234-8), CCX2 (including cores 234-9 through 234-12), and CCX3 232-4 (including cores 234-13 through 234-16). Figure 2 In the exemplary embodiment of FIG. 5 , the second hierarchical level of the processing device 200 is presented as having four CCXs, each having four cores, but in other embodiments, the second hierarchical level of the processing device 200 may include any number of CCXs, each having any corresponding number of cores.
[0021] According to various embodiments, the processing device 200 includes a third hierarchical level (e.g., die level) including, for example, one or more dies (e.g., core complex dies (CCDs)). Each die or CCD 230 includes circuitry including one or more CCXs 232, each CCX including one or more cores 234. For example, in Figure 2 In the exemplary embodiment of FIG, the processing device 200 includes a third hierarchical structure level (e.g., CCD or die level) including CCD0 230-1 (including CCX0 232-1 and CCX1 232-2) and CCD1 230-2 (including CCX2 232-3 and CCX3 232-4). Figure 2 In the exemplary embodiment of FIG, the third hierarchical level of the processing device 200 is presented as having two CCDs, each having two CCXs. However, in other embodiments, the third hierarchical level of the processing device 200 may include any number of CCDs, each having any corresponding number of CCXs. In various embodiments, one or more CCXs 232 within a CCD 230 are connected to each other via one or more caches, data structures, or any combination thereof. For example, within CCD0 23-1, CCX0 232-1 and CCX1 232-2 are connected via a data structure.
[0022] In various embodiments, the processing device 200 includes a fourth hierarchical level (e.g., a slot level) that includes, for example, one or more slots. As used herein, a "slot" includes an interface, such as one or more pins, between the processing device 200 and one or more other circuits (e.g., a motherboard). According to various embodiments, a slot includes one or more CCDs 230 connected to each other through one or more caches, data structures, or any combination thereof. For example, in Figure 2 In the exemplary embodiment of FIG, the processing device 200 includes a fourth hierarchical structure level (e.g., a slot level) including a slot (not shown for clarity) including CCD0 230-1 and CCD1 230-2 connected by a data structure (not shown for clarity). Figure 2 In the exemplary embodiment of the present invention, the fourth hierarchical structure level of the processing device 200 is presented as having one slot including two CCDs, but in other embodiments, the fourth hierarchical structure level of the processing device 200 can include any number of slots, each slot including any number of CCDs.
[0023] According to various embodiments, the processing device 200 is configured to downgrade one or more cores 234 of the processing device 200. As used herein, "downgrading" includes disabling one or more cores 234, for example, by fuse disabling, software disabling, or both, to produce disabled cores. According to various embodiments, in response to one or more cores 234 being downgraded, thread count asymmetry at one or more hierarchical levels of the processing device 200 (i.e., different thread counts between enumeration instances at the hierarchical level) sometimes emerges. For example, by downgrading one or more cores 234, one or more enumeration instances (e.g., core, CCX, CCD) at one or more hierarchical levels (e.g., core level, CCX level, CDD level, socket level) may have a different thread count than another enumeration instance at the same hierarchical level. One of ordinary skill in the art will understand that as the number of downgraded cores in an enumeration instance increases, the number of threads in the enumeration instance decreases. As an example, in Figure 2 In the exemplary embodiment of FIG, in response to de-core processing core 2 234-3 and core 3 234-4, CCX0 232-1 will have a different number of threads than CCX1 232-2. In other words, because each CCX has a different number of threads, there is thread count asymmetry at the CCX hierarchy level. As another example, in Figure 2 In the exemplary embodiment, in response to core de-core processing core 4 234-5, core 12 234-13, core 13 234-14, and core 14 234-15, there is thread count asymmetry at the CCD level because there are different numbers of threads per CCD, and there is thread count asymmetry at the CCX level because there are different numbers of threads per CCX.
[0024] In various embodiments, the processing device 200 is configured to change the operating mode of one or more cores 234, for example, between a power efficiency mode and a performance mode. As used herein, "power efficiency mode" includes one or more instructions for operating a core 234 so as to minimize the power consumption of the core (i.e., maximize power efficiency), while "performance mode" as used herein includes one or more instructions for operating a core 234 so as to maximize the computing performance of the core. In some embodiments, one or more cores are designed to operate only in a higher power efficiency mode or a performance mode. In response to changing one or more cores from a power efficiency mode to a performance mode, from a performance mode to a power efficiency mode, or both, a power efficiency asymmetry at one or more hierarchical levels sometimes manifests. Additionally, in response to the processing device including a first number of cores designed to operate only in power efficiency mode and a second number of cores designed to operate only in performance mode, a power efficiency asymmetry at one or more hierarchical levels sometimes manifests. For example, in response to the number of cores in performance mode in an enumeration instance (e.g., core, CCX, CCD) being different from the number of cores in performance mode in another enumeration instance at the same hierarchical level, a power efficiency asymmetry at that hierarchical level manifests. Similarly, in response to the number of cores in power efficiency mode within an enumeration instance (e.g., core, CCX, CCD) being different than the number of cores in power efficiency mode within another enumeration instance at the same hierarchy level, power efficiency asymmetry manifests at that hierarchy level. As an example, in response to CCX0 232-1 having four cores (234-1 through 234-4) in power efficiency mode and CCX1 232-2 having four cores (234-5 through 234-8) in performance mode, power efficiency asymmetry will manifest at the CCX hierarchy level because CCX0 and CCX1 have asymmetric power efficiencies.
[0025] In various embodiments, the processing device 200 is configured to determine one or more hardware and software resources, hardware and software resource attributes, asymmetries (e.g., thread count asymmetry, power efficiency asymmetry), or any combination thereof, at one or more hierarchical levels of the processing device 200. For example, the processing device 200 uses microcode to determine one or more thread counts, thread count asymmetries, power efficiency asymmetries, or any combination thereof. According to various embodiments, the processing device 200 is configured to determine the one or more hardware and software resources, hardware and software resource attributes, asymmetries, or any combination thereof during boot-up of the processing device 200. As an example, at boot-up, the processing device 200 is configured to determine the number of threads for each enumerated instance (e.g., core, CCX, CCD) in the processing device 200. As another example, the processing device 200 is configured to determine the one or more thread count asymmetries at the hierarchical level by executing one or more instructions, microcode, or any combination thereof to compare the number of threads for each enumerated instance (e.g., core, CCX, CCD, socket) to determine whether the number of threads for each enumerated instance differs within the hierarchical level. That is, whether one or more enumeration instances at a hierarchy level have a different number of threads than one or more other enumeration instances at that hierarchy level. In another example, the processing device 200 is configured to determine, at boot time, one or more power efficiency asymmetries at one or more hierarchy levels using microcode. According to various embodiments, the processing device 200 is configured to store one or more determined hardware and software resources, hardware and software resource attributes, asymmetries, or any combination thereof in a memory including, for example, CMOS memory, flash memory, programmable read-only memory (PROM), electrically erasable PROM (EEPROM), RAM, cache, or any combination thereof, to name a few.
[0026] According to various embodiments, processing device 200 includes or is connected to a memory 206 that is similar or identical to memory 106. In various embodiments, memory 206 includes one or more operating systems 224 that are similar or identical to operating system 124, and one or more registers 226-1 through 226-N. Figure 2 In the illustrated embodiment, three registers (226-1, 226-2, 226-N) are presented to represent N registers, but in other embodiments, memory 206 may include any number of registers. According to various embodiments, one or more operating systems 224 each include one or more cores 236 configured to interface or interact with processing device 200. For example, operating system 224 includes core 236 configured to interact with processing device 200 on behalf of one or more applications.
[0027] In various embodiments, the operating system 224 is configured to perform one or more discovery operations 228, such as a CPUID operation, a read model specific register (RDMSR) operation, a read operation on a table stored in memory 206, or any combination thereof, to determine one or more hardware resources, hardware attributes, asymmetries, or any combination thereof of the processing device 200. In various embodiments, the one or more discovery operations 228 include one or more leaves of a discovery operation, such as one or more CPUID leaves. According to various embodiments, the one or more discovery operations 228 include instructions for the processing device 200 to store data representing requested hardware and software resources, hardware and software resource attributes, asymmetries (e.g., thread count asymmetry, power efficiency asymmetry), or any combination thereof in one or more bits of register 226. For example, the discovery operations 228 include instructions for the processing device 200 to generate, store, or generate and store data representing thread count asymmetry at a level of the hierarchy. In various embodiments, the hardware and software resources, hardware and software resource attributes, asymmetries, or any combination thereof requested in the discovery operations 228 are determined by data stored in one or more bits of register 226. That is, the data stored in one or more bits of register 226 determines what hardware and software resources, hardware and software resource attributes, asymmetries, or any combination thereof, the processing device 200 will identify by storing data in register 226. For example, based on one or more bits in first register 226-2, discovery operation 228 includes instructions for the processing device 200 to store data representing the power efficiency asymmetry of the hierarchy level in second register 226-1. According to various embodiments, one or more registers 226 store one or more displacement values, such as one or more displacement values. Based on the displacement values stored in register 226, discovery operation 228 includes one or more instructions for the processing device 200 to store one or more thread identifiers 238 for each enumeration instance of the hierarchy level in one or more bits of register 226 based on the displacement values. For example, the displacement values indicate which bits of register 226 the processing device 200 is to store the thread identifiers 238. As an example, the discovery operation 228 includes one or more instructions for the processing device to store one or more thread identifiers 238 (e.g., APICIDS) in the second register 226-2 based on the displacement value stored in the first register 226-1 indicating a displacement value of 4. The stored thread identifiers 238 may be shifted by 16 bits to provide a unique key for a higher level topology (e.g., a core, a core complex, a die, or a socket).
[0028] In various embodiments, the operating system 224 is configured to determine one or more hardware resources, hardware attributes, asymmetries, or any combination thereof, of the processing device 200 by reading data from one or more bits stored in the register 226. For example, the operating system 224 determines one or more hardware and software resource attributes of the processing device 200 by reading data in one or more bits of the register 226-1 stored by the processing device 200 during the discovery operation 228. According to various embodiments, the operating system 224 provides the one or more determined hardware and software resources, hardware and software resource attributes, asymmetries, or any combination thereof to one or more applications, such that the applications run on at least a portion of the processing device 200. For example, the operating system 224 schedules software tasks for the applications using the one or more hardware and software resources of the processing device 200.
[0029] According to various embodiments, in response to determining one or more thread count asymmetries at one or more hierarchical levels of the processing device 200, the operating system 224 is configured to determine an indication of thread count asymmetry at each hierarchical level of the processing device 200. The indication of thread count asymmetry includes, for example, the number of threads per enumeration instance at the hierarchical level. Determining the indication of thread count asymmetry at each hierarchical level includes, for example, reading data from one or more registers 226 indicating thread identifiers 238 that identify the number of threads per enumeration instance at each hierarchical level. By way of example, thread identifiers 238 indicating the number of threads per CCX at the CCX hierarchy level and the number of threads per CCD at the CCD hierarchy level are read from one or more registers 226. To determine the indication of asymmetry, the operating system 224 is configured to read data from registers 226 indicating only one thread identifier 238 for each enumeration instance in the hierarchical level. For example, for the CCX hierarchy level, the operating system 224 reads data from registers 226 indicating only one thread identifier 238 per CCX. According to various embodiments, the operating system 224 is configured to read data representing one thread identifier 238 for each enumeration instance in a hierarchy level from registers 226 based on a shift value stored in one or more registers 226. For example, based on a shift value indicating a shift of 4 (e.g., a logical right shift operation of 4 results in a thread identifier that is an even multiple of 16), the operating system 224 determines that the first thread identifier 238 for each corresponding enumeration instance is stored at the shift position within each thread's unique key (e.g., at the shift position within the APICID). In other words, the first thread identifier 238 for each corresponding enumeration instance is found in the thread's unique key (e.g., APICID) 0, 16, 32, 48, etc. in registers 226. In this manner, the operating system 224 only reads data representing one thread identifier 238 for each enumeration instance in a hierarchy level from registers 226, rather than for each thread identifier 238 at that hierarchy level. Consequently, the processing time and memory resources required to determine the representation of asymmetry are reduced. In various embodiments, operating system 224 uses the asymmetric representation to provide one or more applications with access to hardware and software resources of processing device 200, such that the applications perform one or more operations, calculations, or instructions using one or more hardware resources of processing device 200. For example, operating system 224 schedules software tasks for the applications based on the thread count indicated in the asymmetric representation.
[0030] Now refer to Figure 3, presents a flowchart of an exemplary method 300 for determining thread count asymmetry at one or more hierarchical levels. At step 305, similar to or identical to processing device 200, the processing device de-cores one or more cores of the processing device to produce disabled cores. For example, one or more cores of the processing device are disabled by fuses, software, or both. At step 310, the processing device determines one or more thread count asymmetries at one or more hierarchical levels of the processing device. For example, at boot time, the processing device determines that one or more corresponding thread counts of an enumeration instance (e.g., core, CCX, CCD) at a hierarchical level device (e.g., core level, CCX, CCD, socket level) are different from one or more corresponding thread counts of another enumeration instance at the hierarchical level. That is, the processing device determines that one or more enumeration instances at the hierarchical level have a different number of threads than one or more other enumeration instances at the hierarchical level. In various embodiments, the processing device uses microcode to determine the number of threads for each enumeration instance in the hierarchical level and compares the thread counts of each enumeration instance to determine the one or more thread count asymmetries. Furthermore, at step 310, the processing device determines, for each thread at each hierarchy level, a thread identifier that is similar to or the same as thread identifier 238. According to various embodiments, the processing device stores the determined asymmetry and thread count in a memory, such as CMOS memory, flash memory, PROM, EEPROM, RAM, cache, or any combination thereof, to name a few.
[0031] At step 315, one or more operating systems similar to or the same as operating system 224 perform one or more discovery operations similar to or the same as discovery operation 228. In various embodiments, the discovery operations each include instructions for the processing device to generate, store, or generate and store the requested data in one or more registers similar to or the same as registers 226. For example, the discovery operations include instructions for the processing device to load (i.e., report) data representing any determined asymmetry, thread count, and thread identifier into one or more bits of a register. For example, the discovery operations include instructions for the processing device to load a corresponding bit for each hierarchical level, the bit indicating the presence or absence of thread count asymmetry at the corresponding hierarchical level (i.e., the bit indicating whether the corresponding hierarchical level is symmetric or asymmetric with respect to thread count). In various embodiments, the discovery operations include instructions to store data in one or more registers based on one or more values stored in the registers (e.g., displacement values stored in the registers). As an example, the discovery operation includes an instruction for the processing device to store the thread identifier (e.g., APICID) at the corresponding hierarchy level in a first register based on a shift value stored in a second register, the shift value indicating a shift of 4 (e.g., a logical right shift operation of 4 results in the thread identifier being an even multiple of 16). Based on the shift of 4, the processing device indicates that the first thread identifier of each enumerated instance of the hierarchical level will be in the place of the thread identifier (e.g., APICID 0, 16, 32, 48, etc.) in the first register. At step 320, the operating system determines whether one or more thread count asymmetries for one or more hierarchy levels of the processing device are reported (i.e., loaded) into the register by the processing device. That is, whether each hierarchy level is symmetric or asymmetric with respect to thread count. For example, the operating system reads one or more bits in the operating system to determine whether the processing device reported one or more asymmetries. In response to determining that no asymmetry exists at any hierarchy level of the processing device (i.e., the hierarchy levels are symmetric with respect to thread count), the system moves to step 325 and ends the discovery operation. In response to determining that one or more asymmetries are reported for one or more hierarchy levels, the system moves to step 330.
[0032] At step 330, in response to determining one or more asymmetries at one or more hierarchical levels, the operating system defines a thread identifier for each enumeration instance (e.g., kernel, CCX, CCD) in the hierarchical level reporting the asymmetry. For example, in response to determining the asymmetry at the CCD hierarchical level, the operating system determines a thread identifier for each CCX in the CCD hierarchical level. In various embodiments, the operating system is configured to determine the thread identifier by reading one or more bits indicating the thread identifier stored in a register. According to various embodiments, the operating system is configured to define the thread identifier by reading only the bits indicating the first thread identifier of each enumeration instance at the hierarchical level stored in the register. The operating system determines the position of the first thread identifier for each enumeration instance based on one or more shift values stored in the register. As an example, based on a shift of 4 (e.g., a logical right shift operation of 4 results in a thread identifier that is an even multiple of 16), the operating system determines that the first thread identifier for each enumeration instance at the hierarchical level will be at thread identifier (e.g., APICID) 0, 16, 32, 48, etc. in the register. After determining the location of the first thread identifier of each enumeration instance in the register, the operating system reads only the first thread identifier of each enumeration instance to determine the indication of asymmetry (e.g., the thread count of each enumeration instance). Because the first thread identifier of each enumeration indicates the number of threads in its respective enumeration instance, the operating system determines the thread count of each enumeration instance at a level of the hierarchy without reading each thread identifier at a level of the hierarchy. Thus, processing time and memory resources used to determine the thread count of each enumeration instance are reduced.
[0033] At step 335, the operating system determines an indication of asymmetry (e.g., how many threads per hierarchy level or how many threads per enumeration instance), and provides one or more applications with access to one or more hardware and software resources of the processing device based on the indication of asymmetry. For example, the operating system schedules one or more tasks for the one or more applications based on the thread count indicated in the indication of asymmetry. The applications are configured to perform or implement one or more operations using the hardware and software resources of the processing device based on the indication of asymmetry. For example, the applications perform the one or more operations using one or more threads of the processing device scheduled by the operating system.
[0034] Now refer to Figure 4, presents a flowchart of an exemplary method 400 for determining power efficiency asymmetries at one or more hierarchical levels of the processing device. At step 405, a processing device similar to or identical to processing device 200 boots up, waits for a period of time after booting up, and changes the operating mode of one or more cores of the processing device, or any combination thereof. For example, the processing device changes one or more cores operating in a power efficiency mode to a performance mode, changes one or more cores operating in a performance mode to a power efficiency mode, or both. As another example, the processing device boots up with a first number of cores operating in a power efficiency mode and a second, different number of cores operating in a performance mode. At step 410, the processing device determines one or more power efficiency asymmetries at one or more hierarchical levels of the processing device. For example, at boot time, the processing device determines that the number of cores having a first operating mode (e.g., performance mode, power efficiency mode, density mode) within an enumeration instance (e.g., core, CCX, CCD) of a hierarchical level is different from the number of cores having the same operating mode within a different enumeration instance of the hierarchical level. That is, the processing device determines that one or more enumeration instances of a hierarchy level have a different number of cores in a first operating mode (eg, performance mode, power efficiency mode, density mode) than one or more other enumeration instances of the hierarchy level.
[0035] At step 415, one or more operating systems similar to or the same as operating system 224 perform one or more discovery operations similar to or the same as discovery operation 228. In various embodiments, the discovery operations each include instructions for the processing device to generate, store, or generate and store the requested data in one or more registers similar to or the same as register 226. For example, the discovery operations include instructions for the processing device to load (i.e., report) data representing any determined power efficiency asymmetry into one or more bits of a register. For example, the discovery operations include instructions for the processing device to load a corresponding bit for each hierarchical level, the bit indicating whether there is a power efficiency asymmetry at the corresponding hierarchical level (i.e., whether the corresponding hierarchical level is symmetric or asymmetric with respect to power efficiency). At step 420, the operating system determines whether one or more power efficiency asymmetries for one or more hierarchical levels of the processing device are reported (i.e., loaded) into the register by the processing device. For example, the operating system reads one or more bits in the operating system to determine whether the processing device has reported one or more power efficiency asymmetries. In response to determining that no power efficiency asymmetry exists at any hierarchy level of the processing device (i.e., all hierarchy levels are symmetric with respect to power efficiency), the system moves to step 425 and ends the discovery operation. In response to determining that one or more power efficiency asymmetries are reported for one or more hierarchy levels, the system moves to step 430.
[0036] At step 430, in response to determining one or more power efficiency asymmetries for one or more hierarchical levels, the operating system determines a representation of the asymmetry. For example, the operating system determines an operating mode of one or more cores within an asymmetric hierarchical level. According to various embodiments, the operating system provides access to one or more hardware and software resources of the processing device based on the one or more determined representations of the power efficiency asymmetry. For example, the operating system schedules one or more software tasks for one or more applications based on the one or more operating modes of the corresponding cores indicated in the representation of the power efficiency asymmetry. In various embodiments, the application is configured to perform or implement one or more operations using the hardware of the processing device based on the representation of the power efficiency asymmetry. For example, based on the needs of the application, the application uses one or more cores of the processing device in a first mode (e.g., power efficiency mode) and uses one or more cores of the processing device in a second mode (performance mode) to perform one or more operations.
[0037] As disclosed herein, in some embodiments, a method includes: in response to determining thread count asymmetry at a hierarchical level of a processing device, defining a thread identifier for each enumeration instance at the hierarchical level to generate a representation of the asymmetry at the hierarchical level; and scheduling one or more tasks based on the representation of the asymmetry at the hierarchical level. In one aspect, the method includes: determining a second thread count asymmetry at a second hierarchical level of the processing device, wherein the second hierarchical level includes one or more dies that include the enumeration instance at the hierarchical level; and defining only one thread identifier for each die within the second hierarchical level of the processing device. In another aspect, only one thread identifier is defined for each enumeration instance within the hierarchical level of the processing device to generate the representation of the asymmetry at the hierarchical level.
[0038] In one aspect, the method includes determining thread count symmetry at a second hierarchical level of a processing device. In another aspect, the method includes accessing a first portion of a register and a second portion of a register, the first portion of the register being configured to store a thread identifier associated with a first enumeration instance at a hierarchical level, the second portion of the register being configured to store a thread identifier associated with a second enumeration instance at a hierarchical level. In yet another aspect, the thread identifier for each enumeration instance is based on a displacement value stored in the register. In yet another aspect, the method includes disabling a first core of the processing device to generate a disabled core, wherein the thread count asymmetry is based on the disabled core. In another aspect, the disabled core is disabled by a fuse. In yet another aspect, the disabled core is disabled by software.
[0039] In some embodiments, a system includes: a memory; and a processing device including one or more processing cores configured to: determine thread count asymmetry at a hierarchical level of the processing device; and execute tasks of an application based on the thread count asymmetry. In one aspect, the one or more processing cores are further configured to: determine thread counts for each enumeration instance at the hierarchical level. In another aspect, the one or more processing cores are further configured to: determine thread count symmetry at a second hierarchical level of the processing device. In yet another aspect, the one or more processing cores are further configured to: determine a thread identifier for each enumeration instance at the hierarchical level based on bits stored in a register.
[0040] In one aspect, the one or more processing cores are further configured to disable a processing core of the processing device to generate a disabled core, wherein the thread count asymmetry is based on the disabled core. In another aspect, the hierarchical level comprises a die level of the processing device. In yet another aspect, the hierarchical level comprises a core complex level of the processing device.
[0041] In some embodiments, a system includes: a memory; and a processing device including one or more processing cores, the processing cores configured to: determine a power efficiency asymmetry at a hierarchical level of the processing device based on the number of cores operating in the operating mode in the hierarchical level; and execute tasks of an application based on the power efficiency asymmetry at the hierarchical level of the processing device. In another aspect, the one or more processing cores are further configured to: change the processing cores of the processing device from an operating mode to a second operating mode. In yet another aspect, the one or more processing cores are further configured to: determine the power efficiency asymmetry at the hierarchical level of the processing device further based on the number of cores in the hierarchical level operating in the second operating mode. In another aspect, the operating mode includes a power efficiency mode, and the second operating mode includes a performance mode.
[0042] In some embodiments, the above-described apparatus and techniques are implemented in a system (such as the one or more integrated circuit (IC) devices (also referred to as integrated circuit packages or microchips)) that includes one or more integrated circuit (IC) devices (also referred to as integrated circuit packages or microchips). Figures 1 to 4The system described herein is implemented in a computer system). Electronic design automation (EDA) and computer-aided design (CAD) software tools can be used in the design and manufacture of these IC devices. These design tools are typically represented by one or more software programs. The one or more software programs include code that can be executed by a computer system to manipulate the computer system to operate on the code representing the circuits of one or more IC devices in order to perform at least a portion of the process for designing or adjusting a manufacturing system to manufacture the circuits. The code may include instructions, data, or a combination of instructions and data. The software instructions representing the design tools or manufacturing tools are typically stored in a computer-readable storage medium that can be accessed by the computing system. Similarly, the code representing one or more stages of the design or manufacture of the IC device may be stored in and accessed from the same computer-readable storage medium or different computer-readable storage media.
[0043] Computer-readable storage media may include any non-transitory storage medium or combination of non-transitory storage media that can be accessed by a computer system during use to provide instructions and / or data to the computer system. Such storage media may include, but are not limited to, optical media (e.g., compact discs (CDs), digital versatile discs (DVDs), Blu-ray discs), magnetic media (e.g., floppy disks, magnetic tapes, or magnetic hard drives), volatile memory (e.g., random access memory (RAM) or cache), non-volatile memory (e.g., read-only memory (ROM) or flash memory), or microelectromechanical system (MEMS)-based storage media. Computer-readable storage media may be embedded in a computing system (e.g., system RAM or ROM), fixedly attached to a computing system (e.g., a magnetic hard drive), removably attached to a computing system (e.g., an optical disc or flash memory based on a universal serial bus (USB)), or coupled to a computer system via a wired or wireless network (e.g., a network accessible storage device (NAS)).
[0044] In some embodiments, certain aspects of the above-mentioned technology can be implemented by one or more processors of a processing system that executes software. The software includes one or more sets of executable instructions that are stored in or otherwise tangibly embodied on a non-transient computer-readable storage medium. The software may include instructions and certain data that, when executed by one or more processors, manipulate one or more processors to perform one or more aspects of the above-mentioned technology. The non-transient computer-readable storage medium may include, for example, a magnetic or optical disk storage device, a solid-state storage device such as a flash memory, a cache memory, a random access memory (RAM), or other one or more non-volatile memory devices. The executable instructions stored on the non-transient computer-readable storage medium may be source code, assembly language code, object code, or other instruction formats that are interpreted by one or more processors or that can be executed in other ways.
[0045] It should be noted that not all activities or elements described above in the general description are required, a portion of a particular activity or device may not be required, and one or more additional activities may be performed, or elements may be included in addition to those described. Furthermore, the order in which the activities are listed is not necessarily the order in which they are performed. In addition, these concepts have been described with reference to specific embodiments. However, it is understood by those skilled in the art that various modifications and changes may be made without departing from the scope of the present disclosure as set forth in the claims below. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive, and all such modifications are intended to be included within the scope of the present disclosure.
[0046] The preposition "or" used in the context of "at least one of A, B, or C" is used herein to mean "including or." That is, in the above and similar contexts, "or" is used to mean "at least one of or any combination thereof." For example, "at least one of A, B, and C" is used to mean "at least one of A, B, and C or any combination thereof."
[0047] Benefits, other advantages and solutions to problems have been described above with respect to specific embodiments. However, benefits, advantages, solutions to problems, and any features that may cause any benefit, advantage, or solution to appear or become more pronounced should not be construed as key, required, or essential features of any or all of the claims. Furthermore, the specific embodiments disclosed above are illustrative only, as the disclosed subject matter may be modified and practiced in different but equivalent manners that will be apparent to those skilled in the art having the benefit of the teachings herein. No limitation is intended to the details of construction or design shown herein, except as described in the claims below. It is therefore apparent that the specific embodiments disclosed above may be changed or modified, and all such variations are considered to be within the scope of the disclosed subject matter. Accordingly, the protection sought herein is as set forth in the claims below.
Claims
1. A method comprising: responsive to determining a thread count asymmetry at a hierarchical level of a processing device, defining a thread identifier for each enumeration instance at the hierarchical level to produce a representation of the asymmetry at the hierarchical level; as well as One or more tasks are scheduled based on the representation of the asymmetry at the hierarchy level.
2. The method according to claim 1, further comprising: determining a second thread count asymmetry at a second hierarchical level of the processing device, wherein the second hierarchical level includes one or more dies that include enumeration instances at the hierarchical level; as well as Only one thread identifier is defined for each die within the second hierarchical level of the processing device.
3. The method according to claim 1, wherein Only one thread identifier is defined for each enumeration instance within the hierarchical level of the processing device to produce the representation of the asymmetry at the hierarchical level.
4. The method according to claim 1, further comprising: Thread count symmetry at a second hierarchical level of the processing device is determined.
5. The method according to claim 1, further comprising: Accessing a first portion of a register configured to store a thread identifier associated with a first enumeration instance at a level of the hierarchy and a second portion of the register configured to store a thread identifier associated with a second enumeration instance at a level of the hierarchy.
6. The method according to claim 1, wherein The thread identifier for each enumeration instance is based on a displacement value stored in a register.
7. The method according to claim 1, further comprising: A first core of the processing device is disabled to produce a disabled core, wherein the thread count asymmetry is based on the disabled core.
8. The method according to claim 7, wherein: The disabled core is disabled by a fuse.
9. The method according to claim 7, wherein: The disabled core is disabled by software.
10. A system comprising: Memory; and A processing device, the processing device comprising one or more processing cores, the processing cores being configured to: determining thread count asymmetry at a hierarchy level of the processing device; and Tasks of the application program are executed based on the thread count asymmetry.
11. The system of claim 10, wherein the one or more processing cores are further configured to: A thread count is determined for each enumeration instance at the hierarchy level.
12. The system of claim 10, wherein the one or more processing cores are further configured to: Thread count symmetry at a second hierarchical level of the processing device is determined.
13. The system of claim 10, wherein the one or more processing cores are further configured to: A thread identifier for each enumeration instance of the hierarchy level is determined based on bits stored in a register.
14. The system of claim 10, wherein the one or more processing cores are further configured to: Processing cores of the processing device are disabled to produce disabled cores, wherein the thread count asymmetry is based on the disabled cores.
15. The system according to claim 10, wherein: The hierarchical levels include a die level of the processing device.
16. The system according to claim 10, wherein: The hierarchical levels include a core complex level of the processing device.
17. A system comprising: Memory; and A processing device, the processing device comprising one or more processing cores, the processing cores being configured to: determining a power efficiency asymmetry of the processing device at a hierarchical level based on a number of cores operating in an operating mode in the hierarchical level; as well as Tasks of an application are executed based on the power efficiency asymmetry at the hierarchical level of the processing device.
18. The system of claim 17, wherein the one or more processing cores are further configured to: A processing core of the processing device is changed from the operating mode to a second operating mode.
19. The system of claim 18, wherein the one or more processing cores are further configured to: The power efficiency asymmetry at the hierarchical level of the processing device is determined further based on a number of cores in the hierarchical level operating in the second operating mode.
20. The system of claim 18, wherein: The operating mode comprises a power efficiency mode, and the second operating mode comprises a performance mode.
Citation Information
Patent Citations
Virtual machine and / or multi-level scheduling support on systems with asymmetric processor cores
CN102402458A
System and method for dymanic ordering in a network processor
CN1759379A