Memory bandwidth control in the core

JP7909352B2Active Publication Date: 2026-08-21INTEL CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2022027270
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-03-27
Filing Date
2022-02-24
Publication Date
2026-08-21
Estimated Expiration
2042-02-24

Smart Images

  • Figure 0007909352000007
    Figure 0007909352000007
  • Figure 0007909352000008
    Figure 0007909352000008
  • Figure 0007909352000009
    Figure 0007909352000009
Patent Text Reader

Abstract

To provide a device and a system, for controlling a bandwidth in a core.SOLUTION: In a system, a core includes a local memory bandwidth monitor for each thread. The local bandwidth monitor of each thread at least allocates a bandwidth, for a memory request originating from the thread according to a class of service level stored in a field of quality of service (QoS) model-specific register (MSR). The class of service level includes the memory bandwidth monitor specified by a class of service field in a platform quality of service MSR, and execution resources to support execution of at least one thread of the core.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The field of the present invention generally relates to computer architecture, and more specifically to the allocation of shared resources.

Background Art

[0002] Processor cores in a multi-core processor can use shared system resources such as a cache (e.g., a last-level cache or LLC), system memory, input / output (I / O) devices, and an interconnect. The quality of service provided to an application may be degraded and / or become unpredictable due to competition with those or other shared resources.

[0003] Some processors include technologies such as the Resource Director Technology (RDT) by Intel Corporation, which enables visibility and / or control of how shared resources such as LLC and memory bandwidth are being used by different applications running on the processor. For example, such technologies can provide system software to allocate different amounts of resources to different applications and / or monitor resource utilization to temporarily prevent access to resources beyond the allocated amount by low-priority applications.

Brief Description of the Drawings

[0004] Various embodiments according to the present disclosure are described in relation to the drawings.

[0005] [Figure 1] A block diagram of a system that supports memory bandwidth per thread is shown.

[0006] [Figure 2] Embodiments of the IA32_PQR_ASSOC MSR and the IA32_Qos_Core_BW_Thrtl_N MSR are shown.

[0007] [Figure 3] This example shows the mapping of MSRs exposed to software to microarchitectural resources.

[0008] [Figure 4] This illustrates a flow of an exemplary method, including changing memory bandwidth within the core.

[0009] [Figure 5] An exemplary embodiment of the system is shown.

[0010] [Figure 6] A block diagram of an embodiment of processor 600 is shown, which may have more than one core, an integrated memory controller, and integrated graphics.

[0011] [Figure 7(A)] This block diagram shows both an exemplary in-order pipeline and an exemplary out-of-order issue / execution pipeline with register renaming, according to embodiments of the present invention.

[0012] [Figure 7(B)] This block diagram shows exemplary embodiments of both an in-order architecture core and an out-of-order issue / execution architecture core, which are exemplary register renaming, to be included in a processor according to embodiments of the present invention.

[0013] [Figure 8] The following are embodiments of the execution unit circuit, such as the execution unit circuit 762 shown in Figure 7(B).

[0014] [Figure 9] This is a block diagram of the register architecture 900 according to several embodiments.

[0015] [Figure 10] Shows an embodiment of the command format.

[0016] [Figure 11] Shows an embodiment of the addressing field 1005.

[0017] [Figure 12] Shows an embodiment of the first prefix 1001(A).

[0018] [Figure 13(A)] Shows an embodiment of how the R, X, and B fields of the first prefix 1001(A) are used. [Figure 13(B)] Shows an embodiment of how the R, X, and B fields of the first prefix 1001(A) are used. [Figure 13(C)] Shows an embodiment of how the R, X, and B fields of the first prefix 1001(A) are used. [Figure 13(D)] Shows an embodiment of how the R, X, and B fields of the first prefix 1001(A) are used.

[0019] [Figure 14(A)] Shows an embodiment of the second prefix 1001(B). [Figure 14(B)] Shows an embodiment of the second prefix 1001(B).

[0020] [Figure 15] Shows an embodiment of the third prefix 1001(C).

[0021] [Figure 16] Shows a block diagram contrasting the use of a software instruction converter that converts binary instructions in a source instruction set to binary instructions in a target instruction set according to an embodiment of the present invention.

Mode for Carrying Out the Invention

[0022] This disclosure relates to a method, apparatus, system for adjusting the memory bandwidth of a core, and a non-temporary computer-readable storage medium.

[0023] Numerous specific details are provided in the following description. However, it should be understood that embodiments may be carried out without those specific details. In other examples, well-known circuits, structures, and techniques are not shown in detail so as not to obscure the understanding of this description.

[0024] References in this specification to "one embodiment," "one example," or "exemplary embodiment" indicate that the embodiment described may include a particular function, structure, or characteristic, but not all embodiments may necessarily include such a particular function, structure, or characteristic. Furthermore, such language does not necessarily refer to the same embodiment. Moreover, if a particular function, structure, or characteristic is described in relation to one embodiment, it is considered to be within the knowledge of those skilled in the art to make such feature, structure, or characteristic effective in relation to other embodiments, whether or not it is explicitly stated.

[0025] As used herein and in the claims, unless otherwise specified, the use of ordinal numbers such as “first,” “second,” “third,” etc., to describe an element merely indicates that a particular instance of an element or a different instance of a similar element is referred to as such, and is not intended to suggest that the element described in this manner must be in a particular order, temporally, spatially, in ranking, or in any other manner. Also, as used in describing embodiments of the invention, the letter “ / ” between terms may mean that what is being described may include the first term and / or the second term (and / or any other additional term), or may be implemented by and / or in accordance with the first term and / or the second term (and / or any other additional term).

[0026] Furthermore, terms such as “bit,” “flag,” “field,” “entry,” and “indicator” may be used to describe whether a storage location of any type or content in a register, table, database, or other data structure is implemented in hardware or software, but this does not mean that embodiments of the present invention are limited to any particular type of storage location or the number or other elements of bits within any particular storage location. The term “clear” may be used to indicate that a zero boolean value is being stored in a storage location, or otherwise being made to be stored there, and the term “set” may be used to indicate that a one boolean value, all ones, or any other specified value is being stored in a storage location, or otherwise being made to be stored there. However, since any logic convention may be used within embodiments of the present invention, these terms do not limit embodiments of the present invention to any particular logic convention.

[0027] In this specification and in the drawings, the term “thread” and / or blocks labeled “thread” may mean and / or represent applications, software threads, processes, virtual machines, containers, etc., that are executed, carried out, processed, formed, allocated, etc., on, by, and / or to the core.

[0028] The term “core” may mean any processor or execution core as described and / or shown herein and / or as known in the art. The term “uncore” may mean any circuit, logic, subsystem, etc. (e.g., uncore, system agent, etc.) that is in / on the processor or system-on-chip (SoC) but not in the core, as described and / or shown herein and / or as known in the art. However, the use of the terms core and uncore in the description and drawings is not limited to the location of any circuit, hardware, structure, etc., as the location of circuits, hardware, structures, etc. may differ in various embodiments. For example, a model-specific register (MSR) 104 (as defined below) may represent one or more registers, one or more of which may be in the core and one or more in the uncore, etc.

[0029] The term “Quality of Service” (or QoS) may be used to mean or include any degree of quality of service, as referred herein and / or known in the Art, relating to, an individual thread, a group of threads (including all threads), performance, predictability, etc. The term “Memory Bandwidth Allocation” (or MBA) may be used to refer to, or be the use of, a technique for allocating memory bandwidth, and / or an amount of memory bandwidth that is allocated, made available, or made to be allocated.

[0030] Embodiments of the present invention may be used to allocate shared resources such as cache and memory in a computer system. For example, embodiments may perform MBA with improved operation and precision, use MBA to provide increased throughput and higher efficiency compared to already known approaches, and / or provide efficient sharing of cache. The use of embodiments may reduce the "noisy neighbor" problem, where QoS with respect to threads is poor and, in some cases, affects other threads to an unacceptable degree.

[0031] In these and other embodiments, granular throttling as described above may be provided for configuration purposes and may be applied using a control mechanism that approximates the granularity based on time, credit number, etc. In embodiments, rate limit settings (e.g., throttle level, delay value) may be applied to a thread or core via a configuration or MSR that can be configured by system software to map the thread or core to a class of service (CLOS) and the CLOS to a rate limit setting. For example, throttling may be applied via a first MSR that maps the thread to a CLOS (e.g., IA32_PQR_ASSOC, where PQR represents platform quality of service) and via a second MSR that maps the CLOS to a delay value (e.g., IA32_L2_QoS_Ext_Thrtl_n).

[0032] Embodiments may provide mapping threads to any number in CLOS (e.g., 8, 15, etc.) distinct from CLOS identifiers (CLOSID). For example, one or more control registers (e.g., programmable by the Basic Input / Output System (BIOS) for power-up calibration and / or system software) may contain multiple bits (e.g., 4, 8) to specify one of the numbers corresponding to a delay value (e.g., MBEDelay). For example, four 32-bit control registers may be provided to accommodate 16 CLOSIDs and 8-bit MBEDelay values. In embodiments, a default minimum delay value may be used as a non-throttling delay and may be programmed by microcode.

[0033] This eliminates the need to directly define an upper limit on the maximum bandwidth from each logical processor. Furthermore, the bandwidth from each logical processor may be bursty, affecting latency while still meeting external bandwidth requirements.

[0034] Embodiments may provide better QoS than existing technologies where the speed of allocation determination and adjustment may be limited by the speed at which the system software operates, while maintaining compatibility with existing technologies (e.g., in terms of architecture). Embodiments may do this by using a dynamic hardware controller that can react to changing bandwidth conditions faster (e.g., at the microsecond level) than approaches that use a strict bandwidth control mechanism within the core or per-core circuitry. In embodiments, the use of dynamic hardware control of the MBA may allow software that primarily uses LLCs to experience increased throughput with respect to a given throttling level (as described below), resulting in increased system throughput through fine-grained interleaving of high-priority and low-priority requests from threads. In embodiments, the hardware may provide dynamic monitoring of bandwidth and fine-grained calibration of control, which may result in greater throughput and application performance, especially for applications involving the use of different levels of LLCs, where bandwidth demand / usage may intermittently exceed thresholds, compared to previous methods that use a control mechanism based on average bandwidth usage / demand with coarser calibration over longer periods.

[0035] Figure 1 shows a block diagram of a system in which per-thread memory bandwidth is supported. As shown, core 100 includes two threads (thread 0 110(A) and thread 1 110(B)) whose memory bandwidth may be limited. Note that the system may include any number of cores of any architecture (for example, one embodiment may include a processor in a heterogeneous environment or a system having cores of different architectures), and each core may include any number of threads (for example, one embodiment may include a first core having and / or supporting a first number of threads, and a second core having and / or supporting a second number of threads (which may be different from the first)).

[0036] In some embodiments, the bandwidth-limited memory is the cache (e.g., LLC, or Level 3 (L3) cache) and / or memory bandwidth. In embodiments, the shared cache may be manufactured on the same substrate (e.g., semiconductor chip or die, SoC, etc.), and the memory may be on one or more separate substrates and / or in one or more separate packages separate from the package containing the shared cache, however, any configuration and / or integration of shared resources (e.g., cache and / or memory) and users (e.g., cores and / or threads) on or within a substrate, chiplet, multichip module, package, etc., is possible in various embodiments. In some embodiments, the bandwidth-limited memory is the main memory.

[0037] Core 100 includes at least two types of MSRs. As noted above, for each logical processor, MSR170 (e.g., IA32_PQR_ASSOC MSR) specifies the Class of Active Services (CLOS). The software can control the memory bandwidth per core by programming the MBA delay value (percentage of throttling) to the Quality of Service MSR (e.g., IA32_L2_QoS_Ext_BW_Thrtl_CLOS(n)_MSR for traffic to external memory), as noted.

[0038] Each logical processor obtains a memory bandwidth target signaled via a Memory Bandwidth Execution (MBE) level from a Memory Bandwidth Monitor (per thread) 150, ranging from 0 to 15 (0 being unthrottled, 15 being fully throttled). In IA32_L2_QoS_Ext_BW_Thrtl_CLOS(n)_MSR, the MBE level is based on a software-programmed delay value (as a percentage of throttling). The Memory Bandwidth Monitor (per thread) 150 is responsible for adapting the MBE level to constitute memory traffic, LLC hit / miss rate, etc., and takes CLOS from the IA32_PQR_ASSOC MSR.

[0039] The per-thread quality of service bandwidth (MSR) 175 (e.g., IA32_Qos_Core_BW_Thrtl_N, where N is the number of threads) includes a field (in some embodiments, an 8-bit field, however, other sizes such as 4-bit and 16-bit fields may also be used) that specifies the throttling level for a given CLOS. This MSR 175 allows software to communicate memory bandwidth QoS requests for applications running on the logical processor. The programmed value is used by the logical processor to manage intrinsic microarchitectural resources such as queue size and service rate control. Thread scoping enables the migration of virtual machines within a virtual environment.

[0040] The CLOS field of the IA32_PQR_ASSOC MSR170 is used to index to the MSR175, which provides a memory bandwidth level. This level is used by each logic processor to control bandwidth over the interconnect. A reset value of CLOS[i].Level=0 indicates unthrottled bandwidth. In some embodiments, this field may be programmed from 0 to 15. Any value outside this range will result in a general protection violation. A higher value of CLOS[i].Level means greater bandwidth throttling, and a lower number indicates less throttling.

[0041] If the IA32_QoS_Core_BW_Thrtl_N MSR175 has a throttled value, the defined MBE level referenced by the core is at its maximum. (The CLOS.Level programmed for IA32_Qos_Core_BW_Thrtl_N is the uncore MBE level.)

[0042] The local memory bandwidth monitor and / or allocator (per thread) 115 of the bus interface unit 110 handles bandwidth throttling for threads (e.g., request rate across interconnect 160, which may be an on-die interconnect or an interconnect coupled to an off-die device). The local memory bandwidth monitor and / or allocator (per thread) 115 also specifies the number of entries in local queue 120 (which interacts with threads) and / or external queue 130 (which interacts with memory bandwidth monitor (per thread) 150, interconnect 160, and / or memory, cache, etc.). Note that bandwidth throttling may be linear or nonlinear.

[0043] It should be noted that monitors 115 and 150 are combined in some embodiments. That is, in some embodiments, the presence of the local bandwidth monitor 150 is orthogonal to the presence of the bandwidth monitor 150 (i.e., the system can function with or without monitor 150). Monitor 150 would be effective if it is present at the lower bandwidth level determined by monitor 115. If monitor 115 is present and monitor 150 is not, then monitor 115 is the sole bandwidth controller.

[0044] In embodiments, a rate limiter may limit the use of a resource (e.g., memory bandwidth) by a corresponding core and / or thread by restricting access to the resource by the core / thread, for example, based on time, or based on a crediting scheme. In embodiments, throttling techniques may be used to restrict or prevent access during one or more first periods within a second (larger than the first) period, and to allow or provide access during the remainder of the second period. Embodiments may provide various granularities to which access can be restricted / blocked, for example, an embodiment may provide 10% granularity throttling so that a rate limiter can perform throttling to reduce the MBA to 90%, 80%, 70%, etc., of full utilization.

[0045] In embodiments, for example, in an embodiment where cores are connected via a mesh interconnect through which messages can be managed or controlled using a credit scheme, the credit scheme may be used to limit the rate at which cores can pass messages, such as memory access requests to a memory controller. In these, and / or other embodiments, as is true for any circuitry included in the embodiments, the circuitry performing rate limiting may be integrated with, or with, other circuitry of the processor, such as circuitry in or within the interface between cores and meshes, which connects to the integrated memory controller (iMC) (for example, indirectly through such interfaces associated with other cores) but is conceptually represented as a separate block in the drawings.

[0046] In some embodiments, at least one memory bandwidth monitor 150 and / or local memory bandwidth monitor 1115 may include a rate selector that includes hardware and / or software providing a monitoring function (further described below) that determines whether an associated core / thread is overutilizing memory bandwidth and hardware and / or software that provides rate-setting capabilities to set or adjust rate limits for any core / thread that is overusing or consuming less bandwidth than allocated. For example, if a measurement from the monitoring function indicates that a memory bandwidth request is higher than a given memory bandwidth request, a first MBA rate setting may be selected, which is limited and otherwise slower than a second MBA rate setting (e.g., unlimited, unthrottled) that may be selected and / or used.

[0047] In some embodiments, the rate selector may be part of a feedback loop that includes inputs to the rate selector from a point downstream of the rate limiter (i.e., from further away from the source). For example, the rate selector may receive and / or have input from an interface between the LLC (e.g., an L3 or L4 cache) and memory.

[0048] In some embodiments, the rate selector may include a hardware controller (further described below) located within and / or dedicated to the core, which receives information from a caching agent located within and / or dedicated to the core. In embodiments, the rate selector may include a hardware controller that can be enabled / disabled (e.g., by programming an MSR such as MBA_CFG) so that the rate can be selected either by the hardware controller (further described below) or by a software controller (e.g., based on a feedback loop as described below and shown in Figure 1). The use of a hardware controller may be desirable for uses (e.g., data centers) and / or for any other reason, where it can benefit from faster response times (e.g., in the order of microseconds instead of hundreds of milliseconds or seconds, which may be required for software for system-level sampling of Thread Resource Monitoring Identifiers (RMIDs)). The use of a software controller may be desirable for uses (e.g., Internet of Things devices) and / or for any other reason, where it would not benefit from hardware control (e.g., as they may require simple and deterministic bandwidth capping) with respect to the programming ability of prior art that does not include a hardware controller.

[0049] In some embodiments, for each thread, the rate limiter receives an input from the rate selector and / or via a feedback loop, which determines that the corresponding thread should be limited (and, in embodiments, a value of the limited rate to be applied). The decision may be based on monitoring (or measurement, etc.) of requests and / or usage of shared resources (e.g., an intradie interconnect (IDI) or memory bandwidth), as described below and / or elsewhere in this specification.

[0050] For example, a core may be configured to limit itself based on the number of IDI requests per thread, per period, defined in the uncore (e.g., by a rate selector). Within or with respect to a mid-level cache (MLC, e.g., L2 cache 121 or 122), time may be divided into fixed-length windows. A throttling circuit / logic in the MLC cluster (e.g., a rate limiter) determines which microoperation (up) assigned throttling level applies to each thread, and a throttling circuit / logic in the out-of-order (OOO) cluster (e.g., a uop allocator) may apply to that throttling.

[0051] Figure 2 shows embodiments of IA32_PQR_ASSOC MSR and IA32_QoS_Core_BW_Thrtl_N MSRs. As shown, one of the fields of IA32_PQR_ASSOC MSR201 is the CLOS value. This value serves as the index to IA32_QoS_Core_BW_Thrtl_N MSR203. For example, if CLOS=4 in IA32_PQR_ASSOC MSR201, then the CLOS[4] level field of IA32_QoS_Core_BW_Thrtl_N MSR203 is indexed. The value stored in that field is used to map one or more request rates and / or queue thresholds.

[0052] Figure 3 shows an example of mapping a software-exposed MSR to microarchitectural resources. As shown, the software-exposed MSR301 includes the IA32_PQR_ASSOC MSR (shown here per logic) and the IA32_QoS_Core_BW_Thrtl_N register (per thread) (note that a logical processor has multiple threads).

[0053] As shown, CLOS from IA32_PQR_ASSOC MSRs provides an index in the CLOS field in IA32_QoS_Core_BW_Thrtl_N MSRs. The CLOS field has values ​​from 0 to 15, which then map to different levels of QoS. For example, if the CLOS field has a value of 4, the request threshold is 16, the external queue entry threshold per bank is 15, and the local queue entry threshold per bank is 6.

[0054] In this example, IA32_QoS_Core_BW_Thrtl_0 corresponds to CLOS values ​​0-7, and IA32_QoS_Core_BW_Thrtl_1 corresponds to CLOS values ​​8-15.

[0055] The table below further shows an exemplary CLOS field-level mapping for requesting bandwidth for queue entries and linear characteristics. [Table 1]

[0056] The table below further shows an exemplary CLOS field-level mapping for requesting bandwidth for queue entries and nonlinear characteristics. [Table 2]

[0057] In some embodiments, the number of levels and CLOS supported for a logical processor are enumerated in the CPUID leaf as follows: [Table 3]

[0058] Figure 4 shows a flow of an exemplary method involving changing the memory bandwidth within the core. In step 401, the software writes at least a CLOS value for the first software thread to the PQR MSR (e.g., IA32_PQR_ASSOC MSR). For example, a write MSR (WRMSR) instruction may be used to write the CLOS value. This CLOS value is used to index the QoS MSR (e.g., IA32_QoS_Core_BW_Thrtl_N MSR) for the software thread in order to obtain a throttle value used to determine the bandwidth level, etc.

[0059] The bandwidth level and / or queue for the first thread are updated in 403 based on the stored throttle value.

[0060] In error 405, a memory request was sent from the core (and monitored by the bandwidth monitor), and a response was given.

[0061] Furthermore, in 407, feedback is provided to the core regarding the bandwidth required for threads based on software allocation and bandwidth monitoring.

[0062] In 409, a context switch occurs in which the first software thread is swapped for the second software thread. The context switch may include stored state, but it includes writing CLOS to the PQR MSR that should be used for the second thread.

[0063] In 411, the bandwidth level and / or queue for the second software thread are updated based on the stored throttle value.

[0064] The above embodiments may be embodied in several different types of architectures and systems, examples of which are detailed below. Exemplary Computer Architectures

[0065] A description of exemplary computer architectures is detailed below. Other system designs and configurations known in the art, such as laptops, desktops, handheld PCs, portable information terminals, engineering workstations, servers, network devices, network hubs, switches, embedded processors, digital signal processors (DSPs), graphics devices, video game devices, set-top boxes, microcontrollers, mobile phones, portable media players, handheld devices, and various other electronic devices, are also suitable. Generally, a wide variety of systems or electronic devices capable of incorporating the processors and / or other execution logic disclosed herein are generally suitable.

[0066] Figure 5 shows an exemplary system embodiment. The multiprocessor system 500 is a point-to-point interconnect system and includes multiple processors, including a first processor 570 and a second processor 580, coupled via a point-to-point interconnect 550. In some embodiments, the first processor 570 and the second processor 580 are homogeneous. In some embodiments, the first processor 570 and the second processor 580 are heterogeneous.

[0067] Processors 570 and 580 are shown including integrated memory controller (IMC) unit circuits 572 and 582, respectively. Processor 570 also includes point-to-point (PP) interfaces 576 and 578 as part of its multiple interconnect controller units. Similarly, the second processor 580 includes PP interfaces 586 and 588. Processors 570 and 580 may exchange information using the PP interface circuits 578 and 588 via the point-to-point (PP) interconnect 550. IMCs 572 and 582 connect processors 570 and 580 to their respective memories, namely memories 532 and 534, which may be part of the main memory locally attached to each processor.

[0068] Processors 570 and 580 may exchange information with chipset 590 via individual PP interconnects 552 and 554 using point-to-point interface circuits 576, 594, 586, and 598, respectively. Chipset 590 may optionally exchange information with coprocessor 538 via high-performance interface 592. In some embodiments, coprocessor 538 is a dedicated processor such as a high-throughput MIC processor, a network or communications processor, a compression engine, a graphics processor, a GPGPU, or an embedded processor.

[0069] A shared cache (not shown) is located outside of either or both of processors 570 and 580, and is further connected to multiple processors via a PP interconnect, so that when a processor is placed in low-power mode, the local cache information of either or both processors may be stored in the shared cache.

[0070] The chipset 590 may be coupled to a first interconnect 516 via interface 596. In some embodiments, the first interconnect 516 may be an interconnect such as a Peripheral Component Interconnection (PCI) interconnect, a PCI Express interconnect, or another I / O interconnect. In some embodiments, one of the interconnects may be coupled to a power control unit (PCU) 517, which may include circuitry, software, and / or firmware for performing power management operations related to the processors 570, 580, and / or coprocessor 538. The PCU 517 provides control information to a voltage regulator so that it causes the voltage regulator to generate an appropriate regulated voltage. The PCU 517 also provides control information to control the generated operating voltage. In various embodiments, the PCU 517 may include various power management logic units (circuitry) to perform hardware-based power management. Such power management may be fully processor-controlled (e.g., triggered by various processor hardware and workload and / or power, thermal, or other processor constraints), and / or power management may be performed in response to an external source (such as a platform or power management source or system software).

[0071] The PCU517 is shown as existing as logic separate from the processor 570 and / or processor 580. In other cases, the PCU517 may run on one or more cores (not shown) of the processor 570 or 580. In some cases, the PCU517 may be implemented as a microcontroller (dedicated or general purpose), or as other control logic configured to run its own dedicated power management code, which may be referred to as P-code. In yet another embodiment, the power management operations to be performed by the PCU517 may be implemented externally on the processor, such as in a separate power management integrated circuit (PMIC) or other components outside the processor. In yet another embodiment, the power management operations to be performed by the PCU517 may be implemented within the BIOS or other system software.

[0072] Various I / O devices 514 may be coupled to the first interconnect 516, along with an interconnect (bus) bridge 518 that couples the first interconnect 516 to the second interconnect 520. In some embodiments, one or more additional processors 515, such as multiple coprocessors, multiple high-throughput MIC processors, multiple accelerators for a GPGPU (e.g., multiple graphics accelerators or multiple digital signal processing (DSP) units), multiple field-programmable gate arrays (FPGAs), or any other processor, are coupled to the first interconnect 516. In some embodiments, the second interconnect 520 may be a low-pin-count (LPC) interconnect. Various devices may be coupled to the second interconnect 520, including, for example, a keyboard and / or mouse 522, a communication device 527, and a storage unit circuit 528. In some embodiments, the storage unit circuit 528 may be a disk drive or other mass storage device that may contain instructions / code and data 530. Furthermore, the audio I / O 524 may be coupled to the second interconnect 520. It should be noted that architectures other than the point-to-point architecture described above are possible. For example, instead of the point-to-point architecture, a system such as the multiprocessor system 500 may implement a multidrop interconnect or other such architecture. Exemplary core architectures, processors, and computer architectures.

[0073] Processor cores can be implemented in different processors in different ways and for different purposes. For example, such core implementations may include 1) general-purpose in-order cores for general-purpose computing, 2) high-performance general-purpose out-of-order cores for general-purpose computing, and 3) dedicated cores primarily for graphics and / or scientific (throughput) computing. Different processor implementations may include 1) CPUs containing one or more general-purpose in-order cores and / or one or more general-purpose out-of-order cores for general-purpose computing, and 2) coprocessors containing one or more dedicated cores primarily for graphics and / or scientific (throughput) computing. Such different processors result in different computer system architectures, which may include: 1) a coprocessor on a separate chip of the CPU; 2) a coprocessor on a separate die in the same package as the CPU; 3) a coprocessor on the same die as the CPU (in this case, such a coprocessor may be referred to as dedicated logic, or a dedicated core, such as integrated graphics and / or scientific (throughput) logic); and 4) a system-on-a-chip on the same die as the described CPU (which may be referred to as an application core or application processor), which may include the aforementioned coprocessor and additional functionality. An exemplary core architecture is described below, followed by descriptions of exemplary processors and computer architectures.

[0074] Figure 6 shows a block diagram of an embodiment of processor 600, which may have more than one core, an integrated memory controller, and integrated graphics. The solid box shows processor 600 having a single core 602A, a system agent 610, and a set of one or more interconnect controller unit circuits 616, while optional additional dashed boxes show alternative example processor 600 having multiple multicores 602(A)-(N), a set of one or more integrated memory controller unit circuits 614 within the system agent unit circuit 610, dedicated logic 608, and a set of one or more interconnect controller unit circuits 616. Note that processor 600 may be one of the processors 570 or 580 or coprocessors 538 or 515 in Figure 5.

[0075] Accordingly, different implementations of the processor 600 may include: 1) a CPU (not shown, but may include one or more cores) having dedicated logic 608 which is integrated graphics and / or scientific (throughput) logic, and cores 602(A)-(N) which are one or more general-purpose cores (e.g., a general-purpose in-order core, a general-purpose out-of-order core, or a combination of the two); 2) a coprocessor having cores 602(A)-(N) which are a number of dedicated cores primarily intended for graphics and / or scientific (throughput); and 3) a coprocessor having cores 602(A)-(N) which are a number of general-purpose in-order cores. Accordingly, the processor 600 may be a general-purpose processor, coprocessor, or dedicated processor, such as a network or communications processor, a compression engine, a graphics processor, a GPGPU (general-purpose graphics processing unit circuit), a high-throughput multi-integrated core (MIC) coprocessor (including 30 or more cores), an embedded processor, etc. The processor may be implemented on one or more chips. The processor 600 may be part of one or more substrates and / or may be mounted on them using any of many processing technologies, such as BiCMOS, CMOS, or NMOS.

[0076] The memory hierarchy includes multiple cores 602(A)-(N) coupled to a set of multiple integrated memory controller unit circuits 614, a set of one or more shared cache unit circuits 606, and one or more levels of cache unit circuits 604(A)-(N) in external memory (not shown). The set of one or more shared cache unit circuits 606 may include one or more intermediate-level caches such as Level 2 (L2), Level 3 (L3), Level 4 (L4), etc., or other levels of cache such as Last Level Cache (LLC), and / or combinations thereof. In some embodiments, a ring-based interconnect network circuit 612 interconnects dedicated logic 608 (e.g., integrated graphics logic), the set of shared cache unit circuits 606, and system agent unit circuits 610, while alternative embodiments use any number of known techniques for interconnecting such units. In some embodiments, coherency is maintained between one or more shared cache unit circuits 606 and the cores 602(A)-(N).

[0077] In some embodiments, one or more cores 602(A)-(N) can be multithreaded. The system agent unit circuit 610 includes its components for coordinating and operating the cores 602(A)-(N). The system agent unit circuit 610 may include, for example, a power control unit (PCU) circuit and / or a display unit circuit (not shown). The PCU may be, or include, the logic and / or components necessary to coordinate the power state of the cores 602(A)-(N) and dedicated logic 608 (e.g., integrated graphics logic). The display unit circuit is for driving one or more externally connected displays.

[0078] Multiple cores 602(A)-(N) may be homogeneous or heterogeneous in terms of their architectural instruction sets. That is, two or more cores among the 602(A)-(N) may be capable of executing the same instruction set, while the other cores may be capable of executing only that instruction set or a subset of a different instruction set. Exemplary core architecture in-order and out-of-order core block diagrams.

[0079] Figure 7(A) is a block diagram illustrating exemplary in-order pipelines and exemplary register renaming, as well as both out-of-order issue / execution pipelines, according to embodiments of the present invention. Figure 7(B) is a block diagram illustrating exemplary embodiments of in-order architecture cores and exemplary register renaming, as well as both out-of-order issue / execution architecture cores, included in a processor according to embodiments of the present invention. In Figures 7(A) to 7(B), solid boxes indicate in-order pipelines and in-order cores, while optional dashed boxes indicate register renaming, out-of-order issue / execution pipelines and cores. Out-of-order embodiments are described assuming that in-order embodiments are a subset of out-of-order embodiments.

[0080] In Figure 7(A), the processor pipeline 700 includes a fetch stage 702, an optional length decode stage 704, a decode stage 706, an optional allocation stage 708, an optional renaming stage 710, a scheduling (also known as dispatch or issue) stage 712, an optional register read / memory read stage 714, an execution stage 716, a write-back / memory write stage 718, an optional exception handling stage 722, and an optional engagement stage 724. One or more operations may be performed in each of those processor pipeline stages. For example, during the fetch stage 702, one or more instructions may be fetched from instruction memory; during the decode stage 706, one or more fetched instructions may be decoded, an address using a transferred register port (e.g., a load-store unit (LSU) address) may be generated, and a branch transfer (e.g., an immediate offset or link register (LR)) may be performed. In one embodiment, the decode stage 706 and the register read / memory read stage 714 may be combined into a single pipeline stage. In one embodiment, during the execution stage 716, the decoded instruction may be executed, LSU addresses / data piped to an Advanced Microcontroller Bus (AHB) interface may be executed, multiplication and addition may be performed, arithmetic operations with branch results may be performed, and so on.

[0081] As an example, an exemplary register renaming out-of-order issue / execution core architecture may implement pipeline 700 as follows: 1) Instruction fetch 738 executes fetch and length decode stages 702 and 704. 2) Decode unit circuit 740 executes the decode stage 706. 3) Rename / allocate unit circuit 752 executes the allocation stage 708 and the renaming stage 710. 4) Scheduler unit circuit 756 executes the schedule stage 712. 5) Physical register file unit circuit 758 and memory unit circuit 770 execute the register read / memory read stage 714. Execution cluster 760 executes the execution stage 716. 6) Memory unit circuit 770 and physical register file unit circuit 758 execute the write-back / memory write stage 718. 7) Various units (unit circuits) may be involved in the exception handling stage 722. 8) The retirement unit circuit 754 and the physical register file unit circuit 758 execute the commit stage 724.

[0082] Figure 7(B) shows a processor core 790 including a front-end unit circuit 730 coupled to an execution engine unit circuit 750, both coupled to a memory unit circuit 770. The core 790 may be a reduced instruction set computing (RISC) core, a composite instruction set computing (CISC) core, a very long instruction word (VLIW) core, or a hybrid or alternative core type. Another option is that the core 790 may be a dedicated core, such as a network or communications core, a compression engine, a coprocessor core, a general-purpose computing graphics processing unit (GPGPU) core, or a graphics core.

[0083] The front-end unit circuit 730 includes a branch prediction unit circuit 732 which is coupled to the instruction cache unit circuit 734. The instruction cache unit circuit 734 is coupled to an instruction translation index buffer (TLB) 736. The TLB 736 is coupled to an instruction fetch unit circuit 738. The instruction fetch unit circuit 738 is coupled to a decode unit circuit 740. In one embodiment, the instruction cache unit circuit 734 is included in the memory unit circuit 770 rather than the front-end unit circuit 730. The decode unit circuit 740 (or decoder) decodes an instruction and may produce as output one or more microoperations, microcode entry points, microinstructions, other instructions, or other control signals that are decoded from the original instruction, otherwise reflect them, or are derived from them. The decode unit circuit 740 may further include an address generation unit circuit (AGU, not shown). In one embodiment, the AGU may generate an LSU address using the transferred register port and may further perform branch transfers (e.g., immediate offset branch transfer, LR register branch transfer, etc.). The decode unit circuit 740 may be implemented using a variety of different mechanisms. Examples of suitable mechanisms include, but are not limited to, lookup tables, hardware implementations, programmable logic arrays (PLAs), and microcode read-only memory (ROM). In one embodiment, the core 790 includes a microcode ROM (not shown) or other medium for storing the microcode of a particular macro instruction (e.g., in the decode unit circuit 740, or otherwise in the front-end unit circuit 730). In one embodiment, the decode unit circuit 740 includes a micro-operation (micro-op) or arithmetic cache (not shown) to hold / cache decoded arithmetic, microtags, or microoperations generated during decoding or other stages of the processor pipeline 700. The decode unit circuit 740 may be coupled with a rename / assign unit circuit 752 in the execution engine unit circuit 750.

[0084] The execution engine unit circuit 750 includes a rename / assign unit circuit 752 coupled with a retirement unit circuit 754 and a set of one or more scheduler circuits 756. The scheduler circuits 756 represent any number of different schedulers, including multiple reservation stations, a central instruction window, etc. In some embodiments, the scheduler circuits 756 include arithmetic logic unit (ALU) scheduler / scheduling circuits, ALU queues, arithmetic generation unit (AGU) scheduler / scheduling circuits, AGU queues, etc. The scheduler circuits 756 are coupled with physical register file circuits 758. Each of the multiple physical register file circuits 758 represents a different one that stores one or more different data types, such as scalar integers, scalar floating-point numbers, packed integers, packed floating-point numbers, vector integers, vector floating-point numbers, status (e.g., an instruction pointer, which is the address of the next instruction to be executed). In one embodiment, the physical register file unit circuit 758 includes a vector register unit circuit, a write mask register unit circuit, and a scalar register unit circuit. These register units may provide architecture vector registers, vector mask registers, general-purpose registers, etc. The physical register file unit circuit 758 is superimposed by a retirement unit circuit 754 (also known as a retirement queue or retirement queue) to show various ways in which register renaming and out-of-order execution may be implemented (e.g., using a reorder buffer (ROB) and a retirement register file, using a future file, a history buffer, and a retirement register file, using a register map and a pool of multiple registers). The retirement unit circuit 754 and the physical register file unit circuit 758 are coupled to an execution cluster 760. The execution cluster 760 includes one or more sets of execution unit circuits 762 and one or more sets of memory access circuits 764.The execution unit circuit 762 may perform various arithmetic, logic, floating-point, or other types of operations (e.g., shift, addition, subtraction, multiplication) on various types of data (e.g., scalar floating-point, packed integer, packed floating-point, vector integer, vector floating-point). Some embodiments may include many execution units or execution unit circuits dedicated to multiple specific functions or multiple sets of multiple functions, while other embodiments may include only one execution unit circuit or multiple execution units / execution unit circuits that perform all functions. In some embodiments, the scheduler circuit 756, the physical register file unit circuit 758, and the execution cluster 760 are shown as being multiple, depending on the case. If separate pipelines are used, one or more of these pipelines may be out-of-order issue / execution, while the rest may be in-order.

[0085] In some embodiments, the execution engine unit circuit 750 may perform load / store unit (LSU) address / data, pipelined to an advanced microcontroller bus (AHB) interface (not shown), address phase, write-back, data phase load, storage, and branch.

[0086] A set of memory access circuits 764 are coupled to a memory unit circuit 770. The memory unit circuit 770 includes a data TLB unit circuit 772 coupled to a data cache circuit 774, and the data cache circuit 774 is coupled to a level 2 (L2) cache circuit 776. In one exemplary embodiment, the memory access unit circuit 764 may include a load unit circuit, a store address unit circuit, and a store data unit circuit, each of which is coupled to the data TLB circuit 772 in the memory unit circuit 770. The instruction cache circuit 734 is further coupled to a level 2 (L2) cache unit circuit 776 in the memory unit circuit 770. In one embodiment, the instruction cache unit 734 and the data cache unit 774 are combined to an L2 cache unit circuit 776, a level 3 (L3) cache unit circuit (not shown), and / or a single instruction and data cache (not shown) in main memory. The L2 cache unit circuit 776 is coupled to one or more other levels of caches and ultimately to main memory.

[0087] The Core 790 may support one or more instruction sets, including the instructions described herein (e.g., the x86 instruction set (with newer versions added and several extensions), the MIPS instruction set, the ARM instruction set (with optional additional extensions such as NEON). In one embodiment, the Core 790 includes logic to support packed data instruction set extensions (e.g., AVX1, AVX2), thereby enabling operations used by many multimedia applications to be performed using packed data. Exemplary execution unit circuit

[0088] Figure 8 shows an embodiment of an execution unit circuit, such as the execution unit circuit 762 in Figure 7(B). As shown, the execution unit circuit 762 may include one or more ALU circuits 801, a vector / SIMD unit circuit 803, a load / storage unit circuit 805, and / or a branch / jump unit circuit 807. The ALU circuit 801 performs integer arithmetic and / or Boolean operations. The vector / SIMD unit circuit 803 performs vector / SIMD operations on packed data (such as SIMD / vector registers). The load / storage unit circuit 805 executes load and store instructions, loading data from memory into registers or storing data from registers into memory. The load / storage unit circuit 805 also generates addresses. The branch / jump unit circuit 807, depending on the instruction, causes branches or jumps to memory addresses. The floating-point unit (FPU) circuit 809 performs floating-point operations. The width of the execution unit circuit 762 varies depending on the embodiment and can range from 16 bits to 1,024 bits. In some embodiments, two or more smaller execution units are logically combined to form a larger execution unit (for example, two 128-bit execution units are logically combined to form a 256-bit execution unit). Exemplary register architecture

[0089] Figure 9 is a block diagram of a register architecture 900 according to several embodiments. As shown, there are vector / SIMD registers 910 ranging in width from 128 bits to 1,024 bits. In some embodiments, the vector / SIMD registers 910 are physically 512 bits, and depending on the mapping, only some of the lower bits are used. For example, in some embodiments, the vector / SIMD register 910 is a 512-bit ZMM register, with the lower 256 bits used by the YMM register and the lower 128 bits used by the XMM register. Thus, there is a register overlay. In some embodiments, the vector length field is selected from the maximum length and one or more other shorter lengths, each of which is half the length of the aforementioned length. Multiple scalar operations are operations performed at the lowest data element positions in the ZMM / YMM / XMM registers. The higher data element positions are either left in the same state as those prior to the instruction, or zeroed out depending on the embodiment.

[0090] In some embodiments, the register architecture 900 includes write mask / predicate registers 915. For example, in some embodiments, there are eight write mask / predicate registers (sometimes referred to as k0 through k7) of sizes 16 bits, 32 bits, 64 bits, or 128 bits, respectively. The write mask / predicate registers 915 enable merging (e.g., enabling any set of elements of a destination to be protected from updating during the execution of any operation) and / or zeroing (e.g., a zeroing vector mask enables any set of elements of a destination to be zeroed out during the execution of any operation). In some embodiments, each data element position in a given write mask / predicate register 915 corresponds to a data element position in a destination. In other embodiments, the write mask / predicate registers 915 are scalable and consist of a set number of enable bits for a given vector element (e.g., 8 enable bits for every 64-bit vector element).

[0091] The register architecture 900 includes several general-purpose registers 925. These registers may be 16-bit, 32-bit, 64-bit, etc., and may be used for scalar operations. In some embodiments, these registers are referred to by the names RAX, RBX, RCX, RDX, RBP, RSI, RDI, RSP, and R8 through R15.

[0092] In some embodiments, the register architecture 900 includes a scalar floating-point register 945 used for scalar floating-point operations on 32 / 64 / 80-bit floating-point data using x87 instruction set extensions or MMX registers, to perform operations on 64-bit packed integer data and to hold operands for some operations performed between MMX and XMM registers.

[0093] One or more flag registers 940 (e.g., EFLAGS, RFLAGS, etc.) store state and control information related to operations, comparisons, and system operation. For example, one or more flag registers 940 may store condition code information such as carry, parity, auxiliary carry, zero, sign, and overflow. In some embodiments, one or more flag registers 940 are referred to as program status and control registers.

[0094] The segment register 920 contains segment points for use in access memory. In some embodiments, these registers are referred to by the names CS, DS, SS, ES, FS, and GS.

[0095] Machine-specific registers (MSRs) 935 control and report on processor performance. Most MSRs 935 handle system-related functions but are not accessible to application programs. Machine check registers 960 consist of error reporting MSRs used to detect and report control, status, and hardware errors.

[0096] One or more instruction pointer registers 930 store instruction pointer values. Control registers 955 (e.g., CR0-CR4) determine the processor's operating mode (e.g., processors 570, 580, 538, 515, and / or 600) and the characteristics of the task currently being performed. Debug registers 950 control and enable monitoring of the debugging operation of the processor or core.

[0097] The memory management register 965 specifies the location of data structures used for protected mode memory management. These registers may include the GDTR, IDRT, task register, and LDTR register.

[0098] Alternative embodiments of the present invention may use broader or narrower registers. Furthermore, alternative embodiments of the present invention may use more, fewer, or different register files and registers. Instruction Set

[0099] An instruction set architecture (ISA) may include one or more instruction formats. A given instruction format may, among other things, define various fields (e.g., bit numbers, bit positions) that specify the operation to be performed (e.g., an opcode) and the operands on which that operation should be performed, and / or other data fields (e.g., a mask). Some instruction formats may be further broken down through the definition of multiple instruction templates (or multiple subformats). For example, an instruction template of a given instruction format may be defined to have a different subset of the fields of the instruction format (the included fields are usually in the same order, but have fewer included fields and therefore have at least some different bit positions), and / or be defined so that the given fields are interpreted differently. Thus, each instruction in an ISA is expressed using a given instruction format (and a given one of the instruction templates of that instruction format, if defined), and includes fields for specifying the operation and operands. For example, an exemplary ADD instruction has an instruction format that includes a specific opcode, as well as an opcode field that specifies that opcode and operand fields that select operands (source 1 / destination, and source 2). The occurrence of this ADD instruction in the instruction stream results in the operand field, which selects a specific operand, having a concrete content. Example instruction format

[0100] The embodiments of the instructions described herein may be embodied in different formats. Furthermore, exemplary systems, architectures, and pipelines are detailed below. The embodiments of the instructions may be executed on, but are not limited to, such systems, architectures, and pipelines.

[0101] Figure 10 shows an embodiment of the instruction format. As shown, the instruction may include, but is not limited to, multiple setting elements, including one or more fields for one or more prefixes 1001, an opcode 1003, addressing information 1005 (e.g., register identifier, memory addressing information), a displacement value 1007, and / or an immediate value 1009. Note that some instructions may utilize some or all of the fields in the format, while others may only utilize the field for the opcode 1003. In some embodiments, the order shown is the order in which those fields should be encoded; however, it should be understood that in other embodiments, those fields may be encoded in a different order, combination, etc.

[0102] The prefix field 1001 modifies the instruction if used. In some embodiments, one or more prefixes are used to repeat string instructions (e.g., 0xF0, 0xF2, 0xF3, etc.), to provide section overrides (e.g., 0x2E, 0x36, 0x3E, 0x26, 0x64, 0x65, 0x2E, 0x3E, etc.), to perform bus lock operations, and / or to modify operands (e.g., 0x66) and address sizes (e.g., 0x67). Certain instructions require mandatory prefixes (e.g., 0x66, 0xF2, 0xF3, etc.). Some of these prefixes may be considered “legacy” prefixes. Other prefixes, one or more of which are detailed herein, exhibit and / or provide further capabilities, such as specifying certain registers, etc. Other prefixes typically follow the “legacy” prefixes.

[0103] The opcode field 1003 is used to define, at least partially, the operations to be performed when the instruction is decoded. In some embodiments, the primary opcode encoded in the opcode field 1003 is 1, 2, or 3 bytes long. In other embodiments, the primary opcode may be of a different length. An additional 3-bit opcode field may be encoded in another field, if applicable.

[0104] The addressing field 1005 is used to address one or more operands of an instruction, such as locations in memory or one or more registers. Figure 11 shows an embodiment of the addressing field 1005. In this explanatory diagram, an optional ModR / M byte 1102 and an optional scale, index, base (SIB) byte 1104 are shown. The ModR / M byte 1102 and the SIB byte 1104 are used to encode instructions with up to two operands, each of which is a direct register or valid memory address. Note that each of these fields is optional, and not all instructions will contain one or more of these fields. The MOD R / M byte 1102 includes the MOD field 1142, the register field 1144, and the R / M field 1146.

[0105] As described above, the contents of the MOD field 1142 distinguish between memory access mode and non-memory access mode. In some embodiments, if the MOD field 1142 has the value of b11, register direct addressing mode is used; otherwise, register indirect addressing is used.

[0106] Register field 1144 may encode either a destination register operand or a source register operand, or an opcode extension, but cannot be used to encode any instruction operand. The contents of register index field 1144 specify the location of the source or destination operand (either in multiple registers or in memory), either directly or through address generation. In some embodiments, register field 1144 is supplemented with additional bits from a prefix (e.g., prefix 1001) to enable larger addressing.

[0107] The R / M field 1146 may be used to encode an instruction operand that references a memory address, or it may be used to encode either a destination register operand or a source register operand. Note that the R / M field 1146 may be combined with the MOD field 1142 to define an addressing mode in some embodiments.

[0108] SIB byte 1104 contains a scale field 1152, an index field 1154, and a base field 1156, which are to be used for address generation. The scale field 1152 indicates the scaling factor. The index field 1154 specifies the index register to be used. In some embodiments, the index field 1154 is supplemented with extra bits from a prefix (e.g., prefix 1001) to enable larger addressing. The base field 1156 specifies the base register to be used. In some embodiments, the base field 1156 is supplemented with extra bits from a prefix (e.g., prefix 1001) to enable larger addressing. In practice, the contents of the scale field 1152 allow scaling of the contents of the index field 1154 for memory address generation (e.g., using an index + base of 2 to the power of scale for address generation).

[0109] Some addressing formats utilize displacement values ​​to generate memory addresses. For example, memory addresses may be generated according to (2 to the power of scale)*index + base + displacement, index*scale + displacement, r / m + displacement, instruction pointer (RIP / EIP) + displacement, register + displacement, etc. The displacement may be a value of 1 byte, 2 bytes, 4 bytes, etc. In some embodiments, the displacement field 1007 provides this value. Furthermore, in some embodiments, the displacement coefficient utilization is encoded in the MOD field of the addressing field 1005, which indicates a compressed displacement scheme for calculating the displacement value by multiplying disp8 with a scaling coefficient N determined based on the instruction vector length, the value of b bits, and the input element size. The displacement value is stored in the displacement field 1007.

[0110] In some embodiments, the immediate value field 1009 specifies the immediate value of the instruction. The immediate value may be encoded as a 1-byte, 2-byte, 4-byte, or the like.

[0111] Figure 12 shows an embodiment of the first prefix 1001(A). In some embodiments, the first prefix 1001(A) is an embodiment of the REX prefix. Instructions using this prefix may specify general-purpose registers, 64-bit packed data registers (e.g., single instruction, multiple data (SIMD) registers, or vector registers), and / or control registers and debug registers (e.g., CR8-CR15 and DR8-DR15).

[0112] Instructions using the first prefix 1001(A) may specify up to three registers using 3-bit fields that depend on the following formats: 1) Using the reg field 1144 and R / M field 1146 of Mod R / M byte 1102; 2) Using Mod R / M byte 1102 together with SIB byte 1104, including using the reg field 1144, base field 1156, and index field 1154; 3) Using the register fields of the opcode.

[0113] In the first prefix 1001(A), bit positions 7:4 are set to 0100. Bit position 3(W) may be used to determine the operand size, but not only the operand width. Therefore, if W=0, the operand size is determined by the code segment descriptor (CS.D), and if W=1, the operand size is 64 bits.

[0114] MOD R / M reg field 1144 and MOD R / MR / M field 1146 can address only 8 registers each individually, but with the addition of another bit, 16(2 4 Note that this allows the registers of ) to be addressed.

[0115] In the first prefix 1001(A), bit position 2(R) may be an extension of the MOD R / M reg field 1144, which may be used to modify the Mod R / M reg field 1144 if the field encodes a general-purpose register, a 64-bit packed data register (e.g., an SSE register), or a control or debug register. R is ignored if the Mod R / M byte 1102 specifies another register or defines an extended opcode.

[0116] The X bit at bit position 1(X) may be modified in the SIB byte index field 1154.

[0117] Bit position B(B)B may be modified by modifying the base of the Mod R / MR / M field 1146 or the SIB byte base field 1156, or by modifying the opcode register field used to access a general-purpose register (e.g., general-purpose register 925).

[0118] Figures 13(A) to 13(D) illustrate embodiments of how the R, X, and B fields of the first prefix 1001(A) are used. Figure 13(A) shows the R and B from the first prefix 1001(A) used to extend the reg field 1144 and R / M field 1146 of the MOD R / M byte 1102 when the SIB byte 1104 is not used for memory addressing. Figure 13(B) shows the R and B from the first prefix 1001(A) used to extend the reg field 1144 and R / M field 1146 of the MOD R / M byte 1102 when the SIB byte 1104 is not used (register-to-register addressing). Figure 13(C) shows R, X, and B from the first prefix 1001(A) used to extend the reg field 1144, index field 1154, and base field 1156 of the MOD R / M byte 1102 when the SIB byte 1104 is used for memory addressing. Figure 13(D) shows B from the first prefix 1001(A) used to extend the reg field 1144 of the MOD R / M byte 1102 when the register is encoded in opcode 1003.

[0119] Figures 14(A) to 14(B) illustrate embodiments of the second prefix 1001(B). In some embodiments, the second prefix 1001(B) is an embodiment of the VEX prefix. The encoding of the second prefix 1001(B) allows instructions to have more than two operands and allows SIMD vector registers (e.g., vector / SIMD register 910) to be longer than 64 bits (e.g., 128 bits and 256 bits). The use of the second prefix 1001(B) provides a three-operand (or more) syntax. For example, previous two-operand instructions performed operations such as A=A+B, which overwrote the source operand. The use of the second prefix 1001(B) allows operands to perform non-destructive operations such as A=B+C.

[0120] In some embodiments, the second prefix 1001(B) has two forms: a 2-byte form and a 3-byte form. The 2-byte second prefix 1001(B) is mainly used for 128-bit, scalar, and some 256-bit instructions, while the 3-byte second prefix 1001(B) provides a compact alternative to the first prefix 1001(A) and 3-byte opcode instructions.

[0121] Figure 14(A) shows an embodiment of the second prefix 1001(B) in a two-byte format. In one example, the format field 1401 (byte 0 1403) contains the value C5H. In one example, byte 1 1405 contains the value "R" of bit [7]. This value is the complement of the same value in the first prefix 1001(A). Bit [2] is used to specify the length (L) of the vector (a value of 0 is a scalar or a 128-bit vector, and a value of 1 is a 256-bit vector). Bits [1:0] provide opcodes that are extensional equivalents to several legacy prefixes (e.g., 00=no prefix, 01=66H, 10=F3H, and 11=F2H). Bits [6:3], shown as vvvv, may be used as follows: 1) Encodes the first source register operand, specified in inverted (one's complement) form and valid for instructions with two or more source operands. 2) Encodes the destination register operand, specified in one's complement form for a particular vector shift. Or, 3) Does not encode any operands, and the field is reserved and should contain some value such as 1111b.

[0122] Instructions using this prefix may encode an instruction operand that references a memory address using Mod R / MR / M field 1146, or they may encode either a destination register operand or a source register operand.

[0123] Instructions using this prefix may use the Mod R / M reg field 1144 to encode either a destination register operand or a source register operand, or they may be treated as opcode extensions and may not be used to encode any instruction operand.

[0124] For the instruction syntax vvvv which supports four operands, Mod R / MR / M field 1146 and Mod R / M reg field 1144 encode three or the four operands. Bits [7:4] of the immediate value 1009 are then used to encode a third source register operand.

[0125] Figure 14(B) shows an embodiment of the second prefix 1001(B) in 3-byte format. In one example, the format field 1411 (byte 0 1413) contains the value C4H. Byte 1 1415 contains, in bits [7:5], the complements of the same value of the first prefix 1001(A): "R", "X", and "B". Bits [4:0] (shown as mmmmm) of Byte 1 1415 contain content that optionally encodes one or more suggested leading opcode bytes. For example, 00001 suggests the 0FH leading opcode, 00010 suggests the 0F38H leading opcode, 00011 suggests the leading 0F3AH opcode, and so on.

[0126] Bits [7] of byte 2 1417 are used similarly to the W of the first prefix 1001(A), including helping to determine the size of the promoteable operand. Bit [2] is used to specify the length (L) of the vector (a value of 0 is a scalar or a 128-bit vector, and a value of 1 is a 256-bit vector). Bits [1:0] provide opcodes for several legacy prefixes and their extensional equivalents (e.g., 00=no prefix, 01=66H, 10=F3H, and 11=F2H). Bits [6:3], denoted as vvvv, may be used as follows: 1) to encode a first source register operand, specified in inverted (one's complement) form and valid for instructions with two or more source operands; 2) to encode a destination register operand, specified in one's complement form for a particular vector shift. Alternatively, 3) Do not encode any operands, and the field should be reserved and contain some value such as 1111b.

[0127] Instructions using this prefix may encode an instruction operand that references a memory address using Mod R / MR / M field 1146, or they may encode either a destination register operand or a source register operand.

[0128] Instructions using this prefix may use the Mod R / M reg field 1144 to encode either a destination register operand or a source register operand, or they may be treated as opcode extensions and may not be used to encode any instruction operand.

[0129] For the instruction syntax vvvv which supports four operands, Mod R / MR / M field 1146 and Mod R / M reg field 1144 encode three or the four operands. Bits [7:4] of the immediate value 1009 are then used to encode a third source register operand.

[0130] Figure 15 shows an embodiment of the third prefix 1001(C). In some embodiments, the first prefix 1001(A) is an embodiment of the EVEX prefix. The third prefix 1001(C) is a 4-byte prefix.

[0131] The third prefix 1001(C) can encode 32 vector registers (e.g., 128-bit, 256-bit, and 512-bit registers) in 64-bit mode. In some embodiments, instructions that utilize write masks / operational masks (see the register descriptions in previous figures such as Figure 9) or predications utilize this prefix. Operational mask registers enable conditional processing or selection control. An operational mask instruction has operational mask registers as its source / destination operands, processes the contents of the operational mask registers as a single value, and is encoded using the second prefix 1001(B).

[0132] The third prefix 1001(C) can encode features specific to an instruction class (for example, packed instructions with "load + op" semantics may support embedded broadcast functionality, floating-point instructions with rounding semantics may support static rounding functionality, and floating-point instructions with non-rounding arithmetic semantics may support "all exception suppression" functionality).

[0133] The first byte of the third prefix 1001(C) is the format field 1511, which in one example has a value of 62H. The subsequent bytes are referred to as payload bytes 1515-1519 and collectively form 24-bit values ​​of P[23:0] that provide specific functionality in the form of one or more fields (as detailed herein).

[0134] In some embodiments, P[1:0] of payload byte 1519 is identical to the two lower mmmm bits. P[3:2] is reserved in some embodiments. Bit P[4](R') enables access to the upper 16 vector register set when combined with P[7] and Mod R / M reg field 1144. P[6] may also provide access to the upper 16 vector registers when SIB type addressing is not required. P[7:5] consists of R, X, and B, which are operand designation modification bits for vector registers, general-purpose registers, and memory addressing, and when combined with Mod R / M register field 1144 and Mod R / MR / M field 1146, enables access to the next set of eight registers beyond the lower eight registers. P[9:8] provides several legacy prefixes and extensional equivalent opcodes (e.g., 00=no prefix, 01=66H, 10=F3H, and 11=F2H). P

[10] is a fixed value of 1 in some embodiments. P[14:11], denoted as vvvv, may be used as follows: 1) Encodes a first source register operand, specified in inverted (one's complement) form and valid for instructions with two or more source operands; 2) Encodes a destination register operand, specified in one's complement form for a particular vector shift; or 3) Does not encode any operand, and the field is reserved and should contain some value such as 1111b.

[0135] P

[15] is similar to X in the first prefix 1001(A) and the second prefix 1011(B), and may function as an opcode extension bit or operand size promotion.

[0136] P[18:16] specifies the index of the register in the opmask (write mask) register (e.g., write mask / predicate register 915). In one embodiment of the present invention, a particular value aaa=000 has special behavior that suggests that a non-opmask is used for a particular instruction (this can be implemented in various ways, including the use of a hardwired opmask for all 1s, or the use of hardware that bypasses masking hardware). When merging, the vector mask allows any set of elements in the destination to be protected from updates during the execution of any operation (specified by basic and extended operations). In another embodiment, the old value of each element of the destination is maintained if the corresponding mask bit is 0. In contrast, when zeroing, the vector mask allows any set of elements in the destination to be zeroed out during the execution of any operation (specified by basic and extended operations). In one embodiment, if the corresponding mask bit has a value of 0, the elements of the destination are set to 0. A subset of this functionality is the ability to control the vector length of the operation being performed (i.e., the range of elements being modified from the first to the last element), however, the elements being modified do not need to be contiguous. Thus, the opmask field enables partial vector operations, including load, store, arithmetic, and logic. Embodiments of the present invention have been described in which the contents of the opmask field select one of many opmask registers containing the opmask to be used (thus indirectly identifying the masking to be performed), but alternative embodiments may, instead of or in addition to this, allow the contents of the mask write field to directly specify the masking to be performed.

[0137] P

[19] can be combined with P[14:11] to encode a second source vector register in a non-destructive source syntax that allows access to the upper 16 vector registers using P

[19] . P

[20] encodes several functions that vary across different instruction classes and can affect the meaning of the vector length / rounding control field (P[22:21]). P

[23] indicates support for merge / write masking (e.g., when set to 0) or support for zeroing and merge / write masking (e.g., when set to 1).

[0138] Exemplary embodiments of register encoding in instructions using the third prefix 1001(C) are detailed in the following table. Table 1: 32 Register Support in 64-bit Mode [Table 4] Table 2: Encoding register specifiers in 32-bit mode [Table 5] Table 3: Operative Mask Register Specifier Encoding [Table 6]

[0139] Program code may be applied to input instructions for performing the functions described herein and generating output information. The output information may be applied to one or more output devices in a known manner. For the purposes of this application, the processing system includes any system having a processor such as a digital signal processor (DSP), a microcontroller, an application-specific integrated circuit (ASIC), or a microprocessor.

[0140] The program code may be implemented in a high-level procedural or object-oriented programming language to communicate with the processing system. The program code may also be implemented in assembly language or machine language, if desired. In practice, the mechanisms described herein are not limited to any particular programming language. In any case, the language may be a compiled language or an interpreted language.

[0141] Embodiments of the mechanisms disclosed herein may be implemented in hardware, software, firmware, or a combination of such implementation methods. Embodiments of the present invention may be implemented as a computer program or program code that runs on a programmable system comprising at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0142] One or more aspects of at least one embodiment may be implemented by representative instructions stored on a machine-readable medium that represent various logics within the processor, and when read by a machine, the instructions cause the machine to generate logic for performing the techniques described herein. Such representations, known as "IP cores," may be stored on a tangible machine-readable medium and supplied to various customers or manufacturing facilities for loading into manufacturing equipment that actually creates the logic or processor.

[0143] Such machine-readable storage media may include, but are not limited to, articles of non-temporary, tangible structures manufactured or formed by machines or devices, including storage media such as hard disks, floppy disks, optical disks, compact disk read-only memory (CD-ROM), rewritable compact disks (CD-RW), and magneto-optical disks, random access memory (RAM) such as read-only memory (ROM), dynamic random access memory (DRAM), static random access memory (SRAM), semiconductor devices such as erasable programmable read-only memory (EPROM), flash memory, electrically erasable programmable read-only memory (EEPROM), and phase-change memory (PCM), magnetic or optical cards, or other types of media suitable for storing electronic instructions.

[0144] Accordingly, embodiments of the present invention include instructions such as a hardware description language (HDL) that define structures, circuits, devices, processors, and / or system functions described herein, or also include non-temporary tangible machine-readable media containing design data. Such embodiments may also be referred to as program products [including emulation (binary conversion, code morphing, etc.)].

[0145] In some cases, instruction converters may be used to translate instructions from a source instruction set to a target instruction set. For example, an instruction converter may translate, morph, emulate, or otherwise translate an instruction to one or more other instructions to be processed by the core (e.g., using static binary translation, dynamic binary translation including dynamic compilation). Instruction converters may be implemented in software, hardware, firmware, or a combination thereof. Instruction converters may be on-processor, off-processor, or partially on-processor and partially off-processor.

[0146] Figure 16 shows a block diagram in contrast to the use of a software instruction converter that converts binary instructions of a source instruction set to binary instructions of a target instruction set, according to an embodiment of the present invention. In the illustrated embodiment, the instruction converter is a software instruction converter, or the instruction converter may be implemented in software, firmware, hardware, or various combinations thereof. Figure 16 shows that a program in a high-level language 1602 may be compiled using a first ISA compiler 1604 to generate a first ISA binary code 1606 that can be inherently executed by a processor using at least one first instruction set core 1616. A processor having at least one first ISA instruction set core 1616 represents any processor capable of substantially achieving the same functionality as an Intel processor having at least one first ISA instruction set core by processing (1) a substantial portion of the instruction set of the first ISA instruction set core, or (2) a version of the object code of an application or other software targeted to run on the processor using at least one first ISA instruction set core, so as to run compatible, otherwise substantially the same as an Intel® processor using at least one first ISA instruction set core. A first ISA compiler 1604 represents a compiler capable of generating first ISA binary code 1606 (e.g., object code) that can be executed on a processor having at least one first ISA instruction set core 1616, with or without additional linkage processing. Similarly, Figure 16 shows that a program in high-level language 1602 can be compiled using an alternative instruction set compiler 1608 to generate alternative instruction set binary code 1610 that can be executed by the processor without the first ISA instruction set core 1614. The instruction converter 1612 is used to convert the first ISA binary code 1606 into code that can be executed by the processor without using the first ISA instruction set core 1614.This converted code is unlikely to be identical to the alternative instruction set binary code 1610, because creating an instruction converter that makes this possible is difficult. However, the converted code performs common operations and consists of instructions from the alternative instruction set. Thus, instruction converter 1612 represents software, firmware, hardware, or a combination thereof that makes the first ISA binary code 1606 executable on a processor or other electronic device that does not have the first ISA instruction set processor or core, through emulation, simulation, or any other process.

[0147] References such as "one embodiment," "one example," and "exemplary embodiment" indicate that the described embodiment may include a particular function, structure, or characteristic, but not all embodiments necessarily include such a particular function, structure, or characteristic. Furthermore, such language does not necessarily refer to the same embodiment. Moreover, if a particular function, structure, or characteristic is described in relation to one embodiment, it is considered within the knowledge of those skilled in the art that such features, structures, or characteristics will be affected in relation to other embodiments, whether or not they are explicitly stated.

[0148] Furthermore, in the various embodiments described above, unless otherwise specifically noted, disjunctive language such as the phrase "at least one of A, B, or C" is intended to be understood to mean any one of A, B, or C, or any combination thereof (e.g., A, B, and / or C). Therefore, disjunctive language is not intended, nor should be understood, to suggest that a given embodiment requires the presence of at least one A, at least one B, or at least one C, respectively.

[0149] Exemplary Embodiments 1. A core-local per-thread memory bandwidth monitor, wherein each thread's local bandwidth monitor allocates at least bandwidth for memory requests originating from that thread according to a service level class stored in a field of a Quality of Service (QoS) Model-Specific Register (MSR), the service level class being specified by the service field class in the platform Quality of Service MSR, and the memory bandwidth monitor, An execution resource that supports the execution of at least one thread on the above core, A memory bandwidth monitor for each thread outside the above core, which monitors memory requests from the above core and provides feedback on bandwidth based on software allocation and bandwidth monitoring, A device equipped with the following features. 2. The device described in Example 1, wherein each field of the above QoS MSR is 8 bits in size. 3. When switching context to a second thread, the above Platform Quality of Service MSR is written to update the class of the above service field, as described in one of the examples 1-2. 4. The apparatus described in Example 3, further comprising a QoS MSR for the second thread described above. 5. The apparatus described in any of Examples 1-4, further comprising a last-level cache. 6. The apparatus according to any one of Examples 1-5, further comprising a memory bandwidth monitor external to the core, which monitors memory requests from the core and provides bandwidth-related feedback based on software allocation and bandwidth monitoring. 7. The above memory request is to main memory, as described in one of the examples in 1-2. 8. The above bandwidth is adjusted in a nonlinear manner, as described in any of Examples 1-6. 9. The above bandwidth is adjusted in a linear manner, as described in any of Examples 1-8. 10. Devices listed in any of Examples 1-9 that support memory bandwidth per thread, as enumerated in the CPUID leaf. 11. The above core further, The apparatus according to any of Examples 1-10, comprising a local queue for receiving memory requests from the threads of the core, wherein the number of available entries in the local queue is configured based on the service value class of the QoS MSR. 12. The above core further, An external queue that receives memory requests from outside the core, wherein the number of available entries in the external queue is determined based on the service value class of the QoS MSR, comprising an external queue, The apparatus described in any of Examples 1-11. 13. A core-local per-thread memory bandwidth monitor, wherein each thread's local bandwidth monitor allocates at least bandwidth for memory requests originating from the thread according to a service level class stored in a field of a Quality of Service (QoS) Model-Specific Register (MSR), the service level class being specified by the service field class in the platform Quality of Service MSR, and the memory bandwidth monitor, An execution resource that supports the execution of at least one thread on the above core, The memory connected to the above core, A system that includes these features. 14. The system described in Example 13, where each field of the above QoS MSR is 8 bits in size. 15. When switching context to a second thread, the above Platform Quality of Service MSR is written to update the class of the above service field, as described in any of the systems in Examples 13-14. 16. The system described in any of Examples 13 to 15, further comprising a QoS MSR for the second thread described above. 17. The system described in any of Examples 13 to 16, wherein the above core further includes a last-level cache. 18. The system according to any one of Examples 13 to 17, further comprising a per-thread memory bandwidth monitor external to the core, which monitors memory requests from the core and provides bandwidth feedback based on software allocation and bandwidth monitoring. 19. The above bandwidth is adjusted in a nonlinear manner, as described in any of Examples 13 to 18. 20. The above bandwidth is adjusted in a linear manner, as described in any of Examples 13 to 19.

[0150] Therefore, the specification and drawings should be considered illustrative rather than restrictive. However, it will be clear that various modifications and changes may be made to them without deviating from the broader spirit and scope of the disclosure as described in the section. (Other possible items) (Item 1) A core-local per-thread memory bandwidth monitor, wherein each thread's local bandwidth monitor allocates at least bandwidth for memory requests originating from that thread according to a service level class stored in a field of a Quality of Service (QoS) Model-Specific Register (MSR), the service level class being specified by the service field class in the platform quality of service MSR, and the memory bandwidth monitor An execution resource that supports the execution of at least one thread on the above core, A device equipped with the following features. (Item 2) The device described in item 1, wherein each field of the above QoS MSR is 8 bits in size. (Item 3) When switching context to the second thread, the above-mentioned Platform Quality of Service MSR is written to update the class of the above-mentioned service field, as described in Item 1. (Item 4) The apparatus described in item 3, further comprising a QoS MSR for the second thread mentioned above. (Item 5) The device described in item 1, further equipped with a last-level cache, wherein the above memory requests are directed to the above last-level cache. (Item 6) The apparatus according to item 1, further comprising a memory bandwidth monitor for each thread outside the core, which monitors memory requests from the core and provides bandwidth feedback based on software allocation and bandwidth monitoring. (Item 7) The above memory request is for main memory, as described in item 1. (Item 8) The above bandwidth is adjusted using a nonlinear method, as described in item 1. (Item 9) The above bandwidth is adjusted in a linear manner, as described in item 1. (Item 10) Devices listed in item 1, enumerated in the CPUID leaf, that support memory bandwidth per thread. (Item 11) The above core further, The apparatus according to item 1, comprising a local queue for receiving memory requests from the threads of the core, wherein the number of available entries in the local queue is determined based on the service value class of the QoS MSR. (Item 12) The above core further, An external queue that receives memory requests from outside the core, wherein the number of available entries in the external queue is determined based on the service value class of the QoS MSR, comprising an external queue, The device described in item 1. (Item 13) A core-local per-thread memory bandwidth monitor, wherein each thread's local bandwidth monitor allocates at least bandwidth for memory requests originating from that thread according to a service level class stored in a field of a Quality of Service (QoS) Model-Specific Register (MSR), the service level class being specified by the service field class in the platform quality of service MSR, and the memory bandwidth monitor An execution resource that supports the execution of at least one thread on the above core, core including, Memory coupled to the above core, A system that includes these features. (Item 14) The system described in item 13, where each field of the above QoS MSR is 8 bits in size. (Item 15) When switching context to the second thread, the above-mentioned Platform Quality of Service MSR is written to update the class of the above-mentioned service field, as described in item 13 of the system. (Item 16) The system described in item 13 further includes a QoS MSR for the second thread mentioned above. (Item 17) The above core is further equipped with a last-level cache, The above memory request is for the last-level cache as described in item 13 of the system. (Item 18) The system according to item 13, further comprising a memory bandwidth monitor external to the above-mentioned core, which monitors memory requests from the above-mentioned core and provides feedback on bandwidth based on software allocation and bandwidth monitoring. (Item 19) The above bandwidth is adjusted in a nonlinear manner, as described in item 13. (Item 20) The above bandwidth is adjusted in a linear manner, as described in item 13 of the system.

Claims

1. A per-thread Quality of Service (QoS) model-specific register (MSR) that stores values ​​for a plurality of service level classes, a QoS MSR, Platform service quality MSR for indexing to the aforementioned QoS MSR, A core-local per-thread memory bandwidth monitor, wherein each thread's local bandwidth monitor allocates at least bandwidth for memory requests originating from that thread, according to a value for the service level class stored in the QoS MSR field, the service level class being specified by the class in the service field of the platform quality of service MSR, and the memory bandwidth monitor An execution resource that supports the execution of at least one thread of the core, A device equipped with the following features.

2. A core-local per-thread memory bandwidth monitor, wherein each thread's local bandwidth monitor allocates at least bandwidth for memory requests originating from that thread according to a service level class stored in a field of a Quality of Service (QoS) Model-Specific Register (MSR), the service level class being specified by a service field class in the platform Quality of Service MSR, and the core-local per-thread memory bandwidth monitor, An execution resource that supports the execution of at least one thread of the core, A per-thread memory bandwidth monitor outside the core monitors memory requests from the core and provides bandwidth feedback based on software allocation and bandwidth monitoring, A device equipped with the following features.

3. A core-local per-thread memory bandwidth monitor, wherein each thread's local bandwidth monitor allocates at least bandwidth for memory requests originating from that thread according to a service level class stored in a field of a Quality of Service (QoS) Model-Specific Register (MSR), the service level class being specified by the service field class in the platform Quality of Service MSR, and the memory bandwidth monitor An execution resource that supports the execution of at least one thread of the core, Equipped with, Support for per-thread memory bandwidth is enumerated in the CPUID leaf. Device.

4. A core-local per-thread memory bandwidth monitor, wherein each thread's local bandwidth monitor allocates at least bandwidth for memory requests originating from that thread according to a service level class stored in a field of a Quality of Service (QoS) Model-Specific Register (MSR), the service level class being specified by the service field class in the platform Quality of Service MSR, and the memory bandwidth monitor An execution resource that supports the execution of at least one thread of the core, Equipped with, The aforementioned core is A local queue for receiving memory requests from threads of the core, wherein the number of available entries in the local queue is configured based on the value of the service level class of the QoS MSR, comprising a local queue. Device.

5. A core-local per-thread memory bandwidth monitor, wherein each thread's local bandwidth monitor allocates at least bandwidth for memory requests originating from that thread according to a service level class stored in a field of a Quality of Service (QoS) Model-Specific Register (MSR), the service level class being specified by the service field class in the platform Quality of Service MSR, and the memory bandwidth monitor An execution resource that supports the execution of at least one thread of the core, Equipped with, The aforementioned core is An external queue for receiving memory requests from outside the core, wherein the number of available entries in the external queue is configured based on the service level class value of the QoS MSR, comprising an external queue, Device.

6. The apparatus according to any one of claims 1 to 5, wherein each field of the QoS MSR is 8 bits in size.

7. The apparatus according to any one of claims 1 to 6, wherein when the context is switched to a second thread, the platform quality of service MSR is written to update the class of the service field.

8. The apparatus according to claim 7, further comprising a QoS MSR for the second thread.

9. The apparatus according to any one of claims 1 to 8, further comprising a last-level cache, wherein the memory request is to the last-level cache.

10. The apparatus according to any one of claims 1 to 9, wherein the memory request is to main memory.

11. The apparatus according to any one of claims 1 to 10, wherein the bandwidth is adjusted in a non-linear manner with respect to the service level class.

12. The apparatus according to any one of claims 1 to 10, wherein the bandwidth is adjusted linearly with respect to the service level class.

13. A per-thread Quality of Service (QoS) model-specific register (MSR) that stores values ​​for a plurality of service level classes, a QoS MSR, Platform service quality MSR for indexing to the aforementioned QoS MSR, A core-local per-thread memory bandwidth monitor, wherein each thread's local bandwidth monitor allocates at least bandwidth for memory requests originating from that thread, according to a value for the service level class stored in the QoS MSR field, the service level class being specified by the class in the service field of the platform quality of service MSR, and the memory bandwidth monitor An execution resource that supports the execution of at least one thread of the core, Cores including, Memory coupled to the aforementioned core, A system equipped with these features.

14. A core-local per-thread memory bandwidth monitor, wherein each thread's local bandwidth monitor allocates at least bandwidth for memory requests originating from that thread according to a service level class stored in a field of a Quality of Service (QoS) Model-Specific Register (MSR), the service level class being specified by the service field class in the platform Quality of Service MSR, and the memory bandwidth monitor An execution resource that supports the execution of at least one thread of the core, core including The memory coupled to the core, A per-thread memory bandwidth monitor outside the core monitors memory requests from the core and provides bandwidth feedback based on software allocation and bandwidth monitoring. A system equipped with these features.

15. A core-local per-thread memory bandwidth monitor, wherein each thread's local bandwidth monitor allocates at least bandwidth for memory requests originating from that thread according to a service level class stored in a field of a Quality of Service (QoS) Model-Specific Register (MSR), the service level class being specified by the service field class in the platform Quality of Service MSR, and the memory bandwidth monitor An execution resource that supports the execution of at least one thread of the core, Cores including, Memory coupled to the aforementioned core, Equipped with, Support for per-thread memory bandwidth is enumerated in the CPUID leaf. system.

16. A core-local per-thread memory bandwidth monitor, wherein each thread's local bandwidth monitor allocates at least bandwidth for memory requests originating from that thread according to a service level class stored in a field of a Quality of Service (QoS) Model-Specific Register (MSR), the service level class being specified by the service field class in the platform Quality of Service MSR, and the memory bandwidth monitor An execution resource that supports the execution of at least one thread of the core, Cores including, Memory coupled to the aforementioned core, Equipped with, The aforementioned core further, A local queue for receiving memory requests from threads of the core, wherein the number of available entries in the local queue is configured based on the value of the service level class of the QoS MSR, comprising a local queue. system.

17. A core-local per-thread memory bandwidth monitor, wherein each thread's local bandwidth monitor allocates at least bandwidth for memory requests originating from that thread according to a service level class stored in a field of a Quality of Service (QoS) Model-Specific Register (MSR), the service level class being specified by the service field class in the platform Quality of Service MSR, and the memory bandwidth monitor An execution resource that supports the execution of at least one thread of the core, Cores including, Memory coupled to the aforementioned core, Equipped with, The aforementioned core further, An external queue for receiving memory requests from outside the core, wherein the number of available entries in the external queue is configured based on the service level class value of the QoS MSR, comprising an external queue, system.

18. The system according to any one of claims 13 to 17, wherein each field of the QoS MSR is 8 bits in size.

19. The system according to any one of claims 13 to 18, wherein when the context is switched to a second thread, the platform quality of service MSR is written to update the class of the service field.

20. The system according to claim 19, further comprising a QoS MSR for the second thread.

21. The system according to any one of claims 13 to 20, wherein the core further comprises a last-level cache, and the memory request is to the last-level cache.

22. The system according to any one of claims 13 to 21, wherein the bandwidth is adjusted in a non-linear manner with respect to the service level class.

23. The system according to any one of claims 13 to 21, wherein the bandwidth is adjusted linearly with respect to the service level class.

24. A per-thread Quality of Service (QoS) model-specific register (MSR) that stores values ​​for a plurality of service level classes, comprising: Platform service quality MSR for indexing to the aforementioned QoS MSR, A core-local per-thread memory bandwidth monitoring means, wherein each thread's local bandwidth monitoring means allocates at least bandwidth for memory requests originating from that thread according to a value for the service level class stored in the field of the QoS MSR, the service level class being specified by the class of the service field in the platform quality of service MSR, and the memory bandwidth monitoring means An execution means that supports the execution of at least one thread of the core, A device equipped with the following features.

25. A core-local per-thread memory bandwidth monitoring means, wherein each thread's local bandwidth monitoring means allocates at least bandwidth for memory requests originating from that thread according to a service level class stored in a field of a Quality of Service (QoS) Model-Specific Register (MSR), the service level class being specified by a service field class in the platform Quality of Service MSR, and the core-local per-thread memory bandwidth monitoring means, An execution means that supports the execution of at least one thread of the core, External thread-specific memory bandwidth monitoring means for the core, which monitors memory requests from the core and provides bandwidth feedback based on software allocation and bandwidth monitoring, A device equipped with the following features.

26. A core-local per-thread memory bandwidth monitoring means, wherein each thread's local bandwidth monitoring means allocates at least bandwidth for memory requests originating from that thread according to a service level class stored in a field of a Quality of Service (QoS) Model-Specific Register (MSR), the service level class being specified by a service field class in the platform Quality of Service MSR, and the memory bandwidth monitoring means An execution means that supports the execution of at least one thread of the core, Equipped with, Support for per-thread memory bandwidth is enumerated in the CPUID leaf. Device.

27. A core-local per-thread memory bandwidth monitoring means, wherein each thread's local bandwidth monitoring means allocates at least bandwidth for memory requests originating from that thread according to a service level class stored in a field of a Quality of Service (QoS) Model-Specific Register (MSR), the service level class being specified by a service field class in the platform Quality of Service MSR, and the memory bandwidth monitoring means An execution means that supports the execution of at least one thread of the core, Equipped with, The aforementioned core is A local queue for receiving memory requests from threads of the core, wherein the number of available entries in the local queue is configured based on the value of the service level class of the QoS MSR, comprising a local queue. Device.

28. A core-local per-thread memory bandwidth monitoring means, wherein each thread's local bandwidth monitoring means allocates at least bandwidth for memory requests originating from that thread according to a service level class stored in a field of a Quality of Service (QoS) Model-Specific Register (MSR), the service level class being specified by a service field class in the platform Quality of Service MSR, and the memory bandwidth monitoring means An execution means that supports the execution of at least one thread of the core, Equipped with, The aforementioned core is An external queue for receiving memory requests from outside the core, wherein the number of available entries in the external queue is configured based on the service level class value of the QoS MSR, comprising an external queue, Device.

29. The apparatus according to any one of claims 24 to 28, wherein each field of the QoS MSR is 8 bits in size.

30. The apparatus according to any one of claims 24 to 28, wherein when the context is switched to a second thread, the platform quality of service MSR is written to update the class of the service field.

31. The apparatus according to claim 30, further comprising a QoS MSR for the second thread.

32. The system further includes a last-level cache, and the memory request is to the last-level cache. The apparatus according to any one of claims 24 to 31.

Citation Information

Patent Citations

  • Dynamic control of memory bandwidth allocation for a processor

    US20200210332A1