EFFICIENT RANGE-BASED MEMORY WRITE-BACK TO IMPROVE HOST-TO-DEVICE COMMUNICATION FOR OPTIMAL POWER AND PERFORMANCE

DE102018002294B4Active Publication Date: 2025-07-17INTEL CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
DE102018002294
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-03-31
Filing Date
2018-03-19
Publication Date
2025-07-17
Estimated Expiration
2038-03-19

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Device comprising: a system memory (300); a plurality of hardware processor cores, at least one of which includes; a first cache (311, 312, 320, 321); a decoder circuit (330) for decoding an instruction having fields for a first memory address and a region indicator, the first memory address and the region indicator defining a contiguous region in the system memory (300), the contiguous region comprising one or more cache lines; and an execution circuit (340) for executing the decoded instruction by invalidating any instances of the one or more cache lines in the first cache (311, 312, 320, 321); wherein an invalidated instance of the one or more cache lines in the first cache (311, 312, 320, 321) is to be stored in the system memory (300) if the invalidated instance is dirty; characterized in that the instruction further includes an opcode to indicate whether the dirty invalidated instance of the one or more cache lines in the first cache (311, 312, 320, 321) is to be stored in a second cache (311, 312, 320, 321) shared by the plurality of hardware processor cores instead of the system memory.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUNDField of the invention

[0001] This invention relates generally to the field of computer processors. More particularly, the invention relates to a method and apparatus for range-based memory writeback. Description of the state of the art

[0002] An instruction set or instruction set architecture (ISA) is the part of computer architecture related to programming, including native data types, instructions, register architecture, addressing styles, memory architecture, interrupt and exception handling, and external input and output (I / O). It should be noted that the term "instruction" here generally refers to macroinstructions—that is, instructions delivered to the processor for execution—as opposed to microinstructions or micro-ops—that is, the result of a processor decoder decoding macroinstructions. The microinstructions or micro-ops can be designed to instruct an execution unit on the processor to perform operations to implement the logic associated with the macroinstruction.

[0003] The ISA is distinct from the microarchitecture, which represents the set of processor design techniques used to implement the instruction set. Processors with different microarchitectures may share a common instruction set. For example, Intel® Pentium 4 processors, Intel® Core™ processors, and processors from Advanced Micro Devices, Inc. in Sunnyvale, CA, implement nearly identical versions of the x86 instruction set (with some extensions added with newer versions), but have different internal designs. For example, the same register architecture of the ISA can be implemented in different ways on different microarchitectures using known techniques, including dedicated physical registers, one or more dynamically allocated physical registers using a register renaming mechanism (e.g.,the use of a Register Alias Table (RAT), a Reorder Buffer (ROB), and a Retirement Register File). Unless otherwise noted, the terms "register architecture," "register file," and "register" are used here to refer to what is visible to the software / programmer and the manner in which instructions specify registers. When a distinction is needed, the adjective "logical," "architectural," or "software-visible" is used to indicate registers / files in the register architecture, while different adjectives are used to name registers in a given microarchitecture (e.g., physical register, reorder buffer, retirement register, register pool).

[0004] An instruction set contains one or more instruction formats. A given instruction format defines various fields (number of bits, position of bits) to specify, among other things, the operation to be performed and the operand(s) on which the operation is to be performed. Some instruction formats are further broken down by the definition of instruction templates (or subformats). For example, the instruction templates of a given instruction format may be defined to have different subsets of the instruction format fields (the included fields are usually in the same order, but at least some have different bit positions because fewer fields are included) and / or may be defined to interpret a given field differently.A given instruction is expressed using a given instruction format (and, if defined, in a given instruction template of that instruction format) and specifies the operation and operands. An instruction stream is a particular sequence of instructions, where each instruction in the sequence is an occurrence of an instruction in an instruction format (and, if defined, in a given instruction template of that instruction format).

[0005] US 6,772,326 B1 describes an instruction for cleaning a range of addresses in a memory area specified by a start and an end parameter. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] The foregoing aspects and many of the attendant advantages of this invention will be better appreciated as the same becomes better understood by reference to the following detailed description taken in conjunction with the accompanying drawings, wherein like reference characters refer to like parts throughout the several views unless otherwise indicated: Fig. 1A-1B illustrate the potential coherence issues between host and device when using a DMA mechanism; Fig. Figure 2 is a flowchart illustrating the use of a shared buffer with a synchronization mechanism between host and device; Fig. 3 illustrates an exemplary platform on which embodiments of the invention may be implemented; Fig. 4A illustrates the contiguous range in system memory defined by the memory address and the range operands when both are memory addresses, according to one embodiment; Fig. 4B illustrates the contiguous range in system memory defined by the memory address and the range operands, where the range is an integer value in accordance with one embodiment; Fig. 5 illustrates an exemplary large data array (LDAT) testing module for debugging and testing devices that may be used to implement embodiments of the present invention; Fig. 6 is a flowchart illustrating the operations and logic for executing the range-based memory write-back instruction in accordance with one embodiment. Fig. 7 is a flowchart illustrating the operation and logic for flushing a cache line according to one embodiment. Fig. 8A is a block diagram illustrating both an example ordered pipeline and an example register renaming, unordered issue / execute pipeline according to embodiments of the invention; Fig. 8B is a block diagram illustrating both an exemplary embodiment of an ordered architecture core and an exemplary register renaming, unordered issue / execute architecture core to be integrated into a processor according to embodiments of the invention; Fig. 9 is a block diagram of a single-core processor and a multi-core processor with integrated memory controller and graphics according to embodiments of the invention; Fig. 10 illustrates a block diagram of a system in accordance with an embodiment of the present invention; Fig. 11 illustrates a block diagram of a second system in accordance with an embodiment of the present invention; Fig. 12 illustrates a block diagram of a third system in accordance with an embodiment of the present invention; Fig. 13 illustrates a block diagram of a system on a chip (SoC) in accordance with an embodiment of the present invention; and Fig. 14 illustrates a block diagram contrasting the use of a software instruction converter to convert binary instructions in a source instruction set to binary instructions in a target instruction set according to embodiments of the invention. DETAILED DESCRIPTION

[0007] The invention is defined by the independent claims. Embodiments of the method and apparatus for efficient range-based memory write-back are described herein. In the following description, numerous specific details are set forth in order to provide a thorough understanding of embodiments of the invention. However, one of ordinary skill in the art will recognize that the invention may be practiced without one or more of the specific details, or with different methods, components, materials, etc. In other instances, well-known structures, materials, or operations are not shown or described in detail to avoid obscuring aspects of the invention.

[0008] Throughout this specification, reference to "a single embodiment" or "an embodiment" means that certain features, structures, or characteristics described in connection with the embodiment are included in at least one embodiment of the present invention. Thus, appearances of the phrases "in a single embodiment" or "in an embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment. Moreover, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. For clarity, individual components in the figures included herein may be referenced by their labels in the figures rather than by a specific reference number.

[0009] Many modern computer applications require host-to-device communication through direct memory access (DMA), a mechanism that allows peripheral components in a computer system to transfer their I / O data directly to and from system memory without the need to involve the system processor. For example, network interface cards (NICs) and graphics processing units (GPUs) can retrieve data packets and blocks directly from host memory to perform their respective functions, bypassing the system processor and accelerating memory operations. Thus, the DMA mechanism can significantly increase throughput to and from a device.However, if the DMA controller controlling the DMA process does not participate in the host processor's cache coherence, which is often the case, the operating system (OS) and / or software running under the OS must ensure that cache lines currently stored in the processor cache are flushed (i.e., stored) in system memory before an outgoing DMA transfer is initiated (i.e., read from system memory). Likewise, cache lines in the processor cache must be invalidated before a memory area affected by an incoming DMA transfer is used (e.g., by writing to system memory). In addition, the OS or software must also ensure that the affected memory area is not being used by any running threads to prevent memory coherence contention.

[0010] Fig. Figures 1A-1B illustrate the potential coherence issues between host and device when using a DMA mechanism. A device initiates an asynchronous DMA read or write operation to access data directly to and from system memory instead of through the host cache. However, if the host cache for a DMA write operation has not been flushed immediately prior to the DMA write, the data transferred to system memory by the DMA operation could be overwritten with stale data that has been cached by the host in the host cache. For example, in Fig. 1A, a device (i.e., the graphics processing unit (GPU)) 108 accesses data in system memory 102 directly via a DMA controller 110, without going through the host (i.e., the central processing unit (CPU)) 104. For simplicity, the host, device, DMA controller, and system memory are shown as connected by bus 112. During a DMA write operation 114, the CPU 108 writes data to cache line 120a in system memory 102 through DMA controller 110. Following the DMA write operation 114, the host 104, either by eviction or a write-back operation, writes cache line 120b to system memory 102, overwriting cache line 120a. This creates a coherency problem because cache line 120b may not be the most recent copy of cache line 120, so the data in cache line 120a is lost.

[0011] Likewise, the data in the host cache may be more recent than the copy in system memory unless the host cache has been flushed to system memory immediately before a DMA read operation. For example, in Fig. 1B, the host cache 106 of the host 104 receives a modified copy of the cache line 120c, which has yet to be stored in system memory. If the device initiates a DMA read for the respective cache line in system memory, the device receives an unmodified cache line 120a instead of the modified version 120c.

[0012] To facilitate communication between host and device, a shared buffer with a synchronization mechanism is often used. With this handshaking function, the host predefines a shared data buffer and a flag. After the host has prepared the data to be sent to the device, the host flushes its cache to system memory and sets a "ready" flag to indicate that the data is ready for access by the device. The device that queries the "ready" flag will access the data once it determines that the flag has been set. Then, once the device has finished processing the data, the device sets a "device ready" flag to indicate to the host that it can proceed with the next round of data. As shown in Fig. As shown in Figure 2, after the host main thread has prepared the data, it notifies the device that the input data is ready (i.e., 202). This can be done by the host main thread asserting a ready flag. The device receives the notification by polling the ready flag. After the device has accessed the data, it notifies a host read thread that the output data is ready for access by the host (i.e., 204). After the host read thread has accessed the output data, the host read thread can then notify the host main thread that the shared buffer is ready for reuse (i.e., 206).

[0013] For workloads split between the host (e.g., central processing unit (CPU)) and the device (e.g., graphics processing unit (GPU)), the workloads are typically split into stages, some of which are processed by the host and some by the device. Explicit cache flushing is required at the transition between host and device processing stages. However, cache flushing requires a significant number of host processing cycles, slowing performance and consuming power. The cost is proportional to the size of the cache frame.

[0014] Existing solutions, such as cache line flush (e.g., CLFLUSH), cache line writeback (e.g., CLWB), and writeback invalidate (e.g., WBINV) instructions have their shortcomings and do not adequately address the overhead associated with cache flushing. For cache line flush and cache line writeback instructions, writing data back to system memory requires one instruction per cache line. For a data block of ~3 MB, this can consume up to 1 ms of processor time and requires tens of thousands of instructions. This incurs significant overhead, especially for large data blocks. Regarding the writeback invalidate instruction, which invalidates the entire cache and copies any dirty cache lines back to system memory, a switch to kernel code is required because such an instruction is often implemented as a privileged instruction.A kernel use-space switch itself carries significant overhead, which can wipe out any time and / or resources saved compared to not using individual cache-line instructions, such as CLFLUSH and CLWB. Furthermore, the write-back invalidate instruction invalidates the entire processor cache. This means that useful code and data currently or soon to be used by the processor are also invalidated and must be returned to the cache. This, in turn, slows performance. To address the shortcomings associated with existing solutions, a new set of region-based memory write-back instructions is described here. The new instruction set allows the processor to issue just one instruction (or a few instructions, depending on the number of memory regions) to flush shared system memory without performing a context switch.The new approach allows a significant amount of host processing cycles to be saved and redirected to other tasks, resulting in improved overall performance and user experience. Not to mention the reduction in power consumption and improved energy efficiency, which are crucial for modern computing, especially for mobile devices.

[0015] Fig. Figure 3 illustrates an exemplary processor 355 on which embodiments of the invention may be implemented. Processor 355 includes a set of general-purpose registers (GPRs) 305, a set of vector registers 306, and a set of mask registers 307. The details of a single processor core ("Core 0") are shown for simplicity in Fig. 3. However, it is understood that each Fig. 3 may have the same logic set as core 0. For example, each core may include a dedicated Level 1 (L1) cache 312 and Level 2 (L2) cache 311 for caching instructions and data according to a prescribed cache management policy. The L1 cache 312 includes a separate instruction cache 320 (IL1) for storing instructions and a separate data cache 321 (DL1) for storing data. The instructions and data stored in the various processor caches are managed at the granularity of cache lines, which may have a fixed size (e.g., 64, 128, 512 bytes in length). Each core of this exemplary embodiment has an instruction fetch unit 310 for fetching instructions from main memory 300 and / or a shared Level 3 (L3) cache 316; a decoding unit 330 for decoding the instructions (e.g.decoding program instructions into micro-operations or “µops”); an execution unit 340 for executing the instructions; and a writeback unit 350 for retiring the instructions and writing back the results.

[0016] Instruction fetch unit 310 includes various well-known components, including a next instruction pointer 303 for storing the address of the next instruction to be fetched from memory 300 (or one of the caches); an instruction translation buffer (ITLB) 304 for storing a map of recently used virtual-to-physical instruction addresses to improve address translation speed; a branch prediction unit 302 for speculatively predicting branch instruction addresses; and branch target buffers (BTBs) 301 for storing branch addresses and target addresses. The fetched instructions are then streamed to the remaining stages of the instruction pipeline, including the decode unit 330, the execution unit 340, and the write-back unit 350.The structure and function of each of these units are well known to those skilled in the art and are not described in detail here to avoid obscuring the relevant aspects of the various embodiments of the invention.

[0017] In one embodiment, decode unit 330 includes a range-based instruction decoder 331 for decoding the range-based memory write-back instructions described herein (e.g., in sequences of micro-operations in one embodiment), and execution unit 340 includes a range-based instruction execution unit 341 for executing the decoded range-based memory write-back instructions.

[0018] For region-based flushing of a processor cache, the ARFLUSH instruction is described below. According to one embodiment, the ARFLUSH instruction has the following format: ARFLUSH{S}mem_addr,range where the mem_addr operand is a memory address, the range operand is a range indicator, and S is an optional opcode. Together, the mem_addr operand and the range operand define a contiguous range in system memory. For example, in one embodiment, the mem_addr operand is a first memory address indicating the starting point of the contiguous range in system memory, and the range operand is a second memory address indicating the ending point of the contiguous range. According to one embodiment, the memory address is a linear memory address specifying a location in system memory. In other embodiments, the memory address could be an effective memory address, a virtual memory address, or a physical memory address (including a guest physical memory address). Fig. Figure 4A illustrates the contiguous range in system memory defined by the memory address and the range operands when both are memory addresses, according to one embodiment. The ARFLUSH instruction invalidates all cache lines in the processor cache that contain a memory address located in the contiguous range. The processor cache referred to here may be IL1, DL1, L2, or a combination thereof. In some embodiments, the invalidation is propagated throughout the cache coherence domain and may include cache(s) on other core(s). In one embodiment, invalidated cache lines in the processor cache that are dirty (e.g., modified) are written back to system memory. In one embodiment, this is done via a write-back operation or an eviction mechanism.

[0019] In another embodiment, the range operand is not a memory address, but instead is an integer value (i.e., "r") indicating the number of cache lines to be invalidated. According to the embodiment, the contiguous range in system memory begins at the memory address indicated by the mem_addr operand and continues for a number (i.e., "r") of cache lines specified by the range operand. In other words, the "r" number of cache lines in the contiguous range all have a memory address equal to the mem_addr operand or incrementally larger (or smaller, depending on the implementation) than the mem_addr operand. Alternatively, the range operand could indicate the number of bytes to be included in the contiguous range instead of the cache lines.The contiguous range begins at the memory address indicated by the mem_addr operand and continues for the number of bytes specified by the range operand.

[0020] Fig. Figure 4B illustrates the contiguous range in system memory defined by the memory address and the range operands, where the range is an integer value in accordance with one embodiment. In one embodiment, ARFLUSH invalidates all cache lines in the processor cache that contain a memory address equal to or incrementally greater than the mem_addr operand, for a number of cache lines indicated by the integer value in the range operand.

[0021] The optional opcode {S} indicates shared cache, according to one embodiment. In one embodiment, the ARFLUSHS instruction behaves exactly like the ARFLUSH instruction, but flushes cache lines in a contiguous range into a shared cache instead of all the way to system memory. According to one embodiment, a shared cache is a cache shared by two or more processor cores in a host processor.

[0022] According to one embodiment, the ARFLUSH instruction is pre-ordered by fencing operations, such as MFENCE, SFENCE, lock-prefix instructions, or architectural serializing instructions. In another implementation, it could be a subset of these instructions (e.g., only serializing instructions). Of course, in another embodiment, it could be more pre-ordered (e.g., as part of a TSO coherence model). The operating system and / or software running under the operating system can use these ordering instructions to ensure the desired ordering of instructions. In one embodiment, the ARFLUSH instruction can be used at all privilege levels and is subject to all permission controls and faults associated with the byte load.

[0023] For region-based writeback of cache lines in a processor cache, the ARWB instruction is described below. According to one embodiment, the ARWB instruction has the following format: ARWB{S}mem_addr,range where the mem_addr operand is a memory address, the range operand is a range indicator, and S is an optional opcode. Similar to the ARFLUSH instruction described above, the mem_addr operand and the range operand define a contiguous range in system memory. For example, in one embodiment, the mem_addr operand is a memory address that identifies the starting point of the contiguous range in system memory, and the range operand is another memory address that identifies the ending point of the contiguous range. According to one embodiment, the memory address is a linear memory address that specifies a location in system memory. In other embodiments, the memory address could be an effective memory address, a virtual memory address, or a physical memory address (including a guest physical memory address). According to the embodiment, the ARWB instruction writes all dirty (i.e.modified) cache lines that have a memory address that falls within the contiguous range, as defined by the mem_addr and range operands, from the processor cache to system memory. The processor cache referred to here may be IL1, DL1, L2, or a combination thereof. In some embodiments, the write-back is propagated anywhere in the cache coherence domain and may include cache(s) on other core(s). The cache lines in the contiguous range that are not dirty (i.e., unmodified) may be retained in the cache hierarchy.

[0024] The ARWB instruction presents a performance improvement over instructions that invalidate both dirty and clean cache lines. By not invalidating unmodified cache lines, the ARWB instruction reduces cache misses on subsequent cache accesses. In one embodiment, the hardware can choose whether to retain the unmodified cache lines at any level of the cache hierarchy or simply invalidate them.

[0025] In another embodiment, the range operand is not a memory address, but instead is an integer value (i.e., "r") indicating the number of cache lines to be invalidated. According to the embodiment, the contiguous range in system memory begins at the memory address indicated by the mem_addr operand and continues for a number (i.e., "r") of cache lines specified by the range operand. In other words, the "r" number of cache lines in the contiguous range all have a memory address equal to the mem_addr operand or incrementally larger (or smaller, depending on the implementation) than the mem_addr operand.In one embodiment, ARWB writes back all dirty cache lines in the processor cache that have a memory address equal to or incrementally greater (or less) than the mem_addr operand, for a number of cache lines as indicated by the integer value in the range operand. Alternatively, the range operand could indicate the number of bytes to include in the contiguous range instead of cache lines. The contiguous range begins at the memory address indicated by the mem_addr operand and continues for the number of bytes specified by the range operand.

[0026] The optional opcode {S} indicates shared cache, according to one embodiment. In one embodiment, the ARWBS instruction behaves exactly like the ARWB instruction. The only difference is that the ARWBS instruction writes dirty cache lines in the contiguous range to a shared cache instead of to system memory. According to one embodiment, a shared cache is a cache shared by two or more processor cores in a host processor.

[0027] According to one embodiment, ARWB is ordered only by memory. Therefore, the operating system and / or software running under the operating system may use the SFENCE, MFENCE, or lock prefix instructions to achieve the desired ordering. In one embodiment, the ARWB instruction described herein may be used at all privilege levels and is subject to all permission controls and faults associated with byte load. For uses that do not require complete data flushing and for which subsequent access to the data is expected, the ARWB instruction may be preferred over other instructions.

[0028] According to one embodiment, hardware implementations for executing the instructions described herein depend heavily on the processor architecture. In particular, processors with different levels of the cache hierarchy dictate different implementation requirements. The same applies to cache inclusion policies. For example, in a processor with L1 / L2 inclusive caches, flushing the L2 cache alone may be sufficient because any invalidated cache lines in L2 are returned to L1 via back-snooping. The hardware implementation may be simplified, according to one embodiment, based on certain assumptions that may be made regarding the operating system and / or the software running under the operating system. For example, as mentioned above, the operating system and / or the software ensures that no affected memory region is used by another running thread.

[0029] In one embodiment, the range-based memory writeback instructions leverage existing hardware to reduce implementation costs. For example, a cache shrink module or flush module for existing instructions, such as WBINV, can scan specific paths of the processor cache and eliminate dirty lines. Additionally, an existing array check register for debugging and testing devices can provide the datapath (e.g., array read / write muxing with a normal functional path) and control logic (e.g., an address-scanning finite state machine (FSM)), which can be largely shared by the range-based memory writeback instructions described herein.

[0030] Fig. Figure 5 illustrates an exemplary large data array (LDAT) testing module for device debugging and testing that can be used to implement embodiments of the present invention. Datapath 502 and control logic 504 can be reused to implement the ARFLUSH and ARWB instructions disclosed herein. In certain embodiments, a separate set of internal control registers can be defined if congestion on existing internal control registers (e.g., machine-specific registers (MSRs) such as PDAT / SDAT) is not advantageous.

[0031] According to one embodiment, a set of internal control registers is used to track cache lines whose addresses lie within the contiguous range of system memory defined by the range-based memory writeback instructions described herein. Then, each of the cache lines is flushed or written back to system memory. In one embodiment, this is done by calling a corresponding CLFLUSH / CLWB instruction described above.

[0032] However, in some cases, it may be undesirable to have microcode issue a request for each line. In such cases, according to one embodiment, a set of internal control registers is configured, and then an FSM is triggered to scan the specified address range according to the operation and logic described below. A status bit is set upon completion of the entire flow, followed by a check by a ordering instruction (e.g., MFENCE) for serialization purposes.

[0033] Fig. 6 is a flowchart illustrating the logic and operation of the range-based memory write-back instruction in accordance with one embodiment. At block 602, a current cache line (Current CL) is determined based on the mem_addr (i.e., initial memory address) operand of the instruction. At block 604, a cache line flush or cache line write-back instruction is executed for the current cache line. At block 608, the next cache line is designated as the current cache line. In one embodiment, the next cache line has a memory address equal to or incrementally greater than the memory address of the current cache line. At block 610, a decision is made as to whether the current cache line is in-range, such that the address of the current cache line falls within the contiguous range of system memory defined by the initial memory address and range indicator operands of the instruction.If the current cache line is in the range, a cache line flush or writeback instruction for the current cache line is executed at block 604. However, if the current cache line is not in the range, indicating the end of the contiguous range in system memory, the process ends.

[0034] Fig. Figure 7 is a flowchart illustrating the logic and operation for flushing a cache line according to one embodiment. At block 702, a request to flush or write back a cache line (CL) to system memory is received. In one embodiment, the cache line is the Fig. 6 that is in range. At block 704, the cache line tag is read. In one embodiment, microcode initiates a tag read operation to read the state of the cache line. Then, at block 706, a decision is made as to whether to cache the cache line in the processor cache. If the cache line is not found in the processor cache, no action is necessary, and the operation ends. Conversely, if the cache line is found in the processor cache, a further decision is made at block 708 as to whether the cache line is dirty. In one embodiment, a dirty cache line is one that has been modified but has not yet been written back to system memory. If the cache line is found to be dirty, it is evicted at block 710. According to one embodiment, eviction of the cache line is performed by the normal or existing eviction mechanism.Once the cache line has been evicted and thereby stored in system memory via a cache line write-back operation, the state of the cache line in the processor cache may be updated. This may be accomplished, for example, by changing the cache line's status label to "I" (invalidate) or "E" (exclusive). However, if it is determined back at block 708 that the cache line is not dirty (i.e., unmodified), a decision is made at block 712 as to whether to flush or invalidate the cache line. According to one embodiment, this is based on whether the request is to flush / invalidate the cache line or simply write it back to memory. As described above, the ARFLUSH instruction flushes / invalidates a cache line in the processor cache regardless of whether the cache line is dirty or not.In contrast, the ARWB instruction does not flush / invalidate cache lines that are not dirty (i.e., unmodified). If a cache line is to be flushed, the cache line's status in the processor cache is updated at block 714. As discussed above, this can be accomplished by changing the cache line's status label to "I" (invalidate) or "E" (exclusive). Alternatively, the cache line's status label can be changed to other states depending on the desired implementation. However, if the cache line is not to be flushed, no change is made to the cache line label, and the process terminates.

[0035] One embodiment of the present invention is an apparatus comprising: system memory; a plurality of hardware processor cores, each including a first cache; decoder circuitry for decoding an instruction, having first memory address and range indicator fields that collectively define a contiguous range in system memory containing one or more cache lines; and execution circuitry for executing the decoded instruction by invalidating, in the first cache, any instances of the one or more cache lines. In one embodiment, any invalidated instances of the one or more cache lines in the first cache that are dirty are stored in system memory.In some embodiments, the instruction may include an opcode to indicate whether to store the dirty invalidated instances of the one or more cache lines in the first cache, rather than system memory, in a second cache shared by the plurality of hardware processor cores. Regarding the range indicator, in some embodiments, it includes a second memory address such that the contiguous range extends from the first memory address to the second memory address. In other embodiments, the range indicator includes an indication of the number of cache lines included in the contiguous range such that each of the included cache lines has an address equal to or incrementally greater than the first memory address. In one embodiment, the first memory address is a linear memory address.

[0036] Another embodiment of the present invention is an apparatus comprising: a device including system memory; a plurality of hardware processor cores, each including a first cache; decoder circuitry for decoding an instruction, having first memory address and range indicator fields that collectively define a contiguous range in system memory that includes one or more cache lines; and execution circuitry for executing the decoded instruction by causing any dirty instances of the one or more cache lines included in the first cache to be stored in system memory.In some embodiments, the instruction may include an opcode to indicate whether to store the dirty instances of the one or more cache lines in the first cache, rather than system memory, in a second cache shared by the plurality of hardware processor cores. Regarding the range indicator, in some embodiments, it includes a second memory address such that the contiguous range extends from the first memory address to the second memory address. In other embodiments, the range indicator includes an indication of the number of cache lines included in the contiguous range such that each of the included cache lines has an address equal to or incrementally greater than the first memory address. In one embodiment, the first memory address is a linear memory address.

[0037] Another embodiment of the present invention is a method comprising: decoding an instruction having fields for a first memory address and a range indicator that together define a contiguous range in system memory containing one or more cache lines; and executing the decoded instruction by invalidating, in a first processor cache, any instances of the one or more cache lines. In some embodiments, executing a decoded instruction includes storing, in system memory, any invalidated instances of the one or more cache lines in the first processor cache that are dirty. In other embodiments, executing the decoded instruction includes storing, in a second cache shared by a plurality of hardware processor cores, any invalidated instances of the one or more cache lines from the first processor cache that are dirty.Regarding the range indicator, in some embodiments, it includes a second memory address such that the contiguous range extends from the first memory address to the second memory address. In other embodiments, the range indicator includes an indication of the number of cache lines included in the contiguous range such that each of the included cache lines has an address equal to or incrementally greater than the first memory address. In one embodiment, the first memory address is a linear memory address.

[0038] Yet another embodiment of the present invention is a method comprising: decoding an instruction having fields for a first memory address and a range indicator that together define a contiguous range in system memory containing one or more cache lines; and executing the decoded instruction by causing any dirty instances of the one or more cache lines to be stored in a first processor cache in system memory. In some embodiments, the instruction may include an opcode to indicate whether to store the dirty instances of the one or more cache lines in the first processor cache, rather than in system memory, in a second cache shared by the plurality of hardware processor cores.Regarding the range indicator, in some embodiments, it includes a second memory address such that the contiguous range extends from the first memory address to the second memory address. In other embodiments, the range indicator includes an indication of the number of cache lines included in the contiguous range such that each of the included cache lines has an address equal to or incrementally greater than the first memory address. In one embodiment, the first memory address is a linear memory address.

[0039] Fig. 8A is a block diagram illustrating both an example ordered pipeline and an example register renaming, unordered issue / execute pipeline according to embodiments of the invention. Fig. Figure 8B is a block diagram illustrating both an exemplary embodiment of an ordered architecture core and an exemplary register renaming, unordered issue / execute architecture core to be integrated into a processor according to embodiments of the invention. The solid-line boxes in the Fig. Figures 8A-B represent the ordered pipeline and the ordered core, while the optional addition of the dashed boxes represents the register renaming, the unordered issue / execution pipeline, and the core. Since the ordered aspect is a subset of the unordered aspect, the unordered aspect is described.

[0040] In Fig. 8A, a processor pipeline 800 includes a fetch stage 802, a length decode stage 804, a decode stage 806, an allocation stage 808, a rename stage 810, a scheduling stage (also known as a dispatch or issue stage) 812, a register read / memory read stage 814, an execution stage 816, a write-back / memory write stage 818, an exception handling stage 822, and a commit stage 824.

[0041] Fig. 8B shows a processor core 890 including front-end hardware 830 coupled to execution engine hardware 850, both coupled to memory hardware 870. Core 890 may be a reduced instruction set (RISC) core, a complex instruction set (CISC) core, a very long instruction word (VLIW) core, or a hybrid or alternative core type. As a further option, core 890 may be a special-purpose core, such as a networking or communications core, a compression engine, a coprocessor core, a general-purpose computing core on a graphics processing unit (GPGPU), a graphics core, or the like.

[0042] Front-end hardware 830 includes branch prediction hardware 832 coupled to instruction cache hardware 834, which is coupled to an instruction translation buffer (TLB) 836, which is coupled to instruction fetch hardware 838, which is coupled to decode hardware 840. Decode hardware 840 (or decoder) may decode instructions and generate as output one or more micro-operations, microcode entry points, microinstructions, other instructions, or other control signals that are decoded from the original instructions or that otherwise reflect or are derived from the original instructions. Decode hardware 840 may be implemented using numerous different mechanisms. Examples of suitable mechanisms include, but are not limited to, lookup tables, hardware implementations, programmable logic arrays (PLAs), microcode read-only memories (ROMs), etc.In one embodiment, core 890 includes a microcode ROM or other medium that stores microcode for specific macroinstructions (e.g., in decode hardware 840 or otherwise within front-end hardware 830). Decode hardware 840 is coupled to rename / dispatcher hardware 852 in execution engine hardware 850.

[0043] The execution engine hardware 850 includes rename / arbiter hardware 852 coupled to retirement hardware 854, and a set of one or more scheduler hardware parts 856. The scheduler hardware 856 represents any number of different schedulers, including reservation stations, central instruction windows, etc. The scheduler hardware 856 is coupled to physical register file(s) hardware 858. Each physical register file(s) hardware 858 represents one or more physical register files, others of which store one or more different data types, such as scalar integer, scalar floating point, packed integer, packed floating point, vector integer, vector floating point, status (e.g., an instruction pointer that is the address of the next instruction to be executed), etc.In one embodiment, the physical register file(s) hardware 858 includes vector register hardware, write mask register hardware, and scalar register hardware. This register hardware may provide architectural vector registers, vector mask registers, and general-purpose registers. The physical register file(s) hardware 858 is overlapped by the retirement hardware 854 to illustrate various ways in which register renaming and out-of-order execution may be implemented (e.g., using reorder buffer(s) and retirement register file(s); using future file(s), history buffer(s), and retirement register file(s); using register maps and a pool of registers, etc.). The retirement hardware 854 and the physical register file(s) hardware 858 are coupled to the execution cluster(s) 860.The execution cluster(s) 860 includes a set of one or more execution hardware parts 862 and a set of one or more memory access hardware parts 864. The execution hardware 862 can perform various operations (e.g., shift, addition, subtraction, multiplication) and on various data types (e.g., scalar floating point, packed integer, packed floating point, vector integer, vector floating point). While some embodiments may include a number of execution hardware parts dedicated to specific functions or sets of functions, other embodiments may include only one execution hardware part or multiple execution hardware parts that all perform all functions.The scheduler hardware 856, the physical register file(s) hardware 858, and the execution cluster(s) 860 are shown as possibly plural because certain embodiments create separate pipelines for certain types of data / operations (e.g., a scalar integer pipeline, a scalar floating point / packed integer / packed floating point / vector integer / vector floating point pipeline, and / or a memory access pipeline, each having its own scheduler hardware, physical register file(s) hardware, and / or execution cluster—and in the case of a separate memory access pipeline, certain embodiments are implemented in which only the execution cluster of that pipeline has the memory access hardware 864). It should also be understood that where separate pipelines are used, one or more of these pipelines may be out-of-order issue / execution and the rest may be ordered.

[0044] The set of memory access hardware 864 is coupled to memory hardware 870, which includes data TLB hardware 872 coupled to data cache hardware 874 coupled to Level 2 (L2) cache hardware 876. In an exemplary embodiment, memory access hardware 864 may include load hardware, memory address hardware, and memory data hardware, each of which is coupled to data TLB hardware 872 in memory hardware 870. Instruction cache hardware 834 is further coupled to Level 2 (L2) cache hardware 876 in memory hardware 870. The L2 cache hardware 876 is coupled with one or more levels of cache and finally with main memory.

[0045] To give an example, the exemplary register renaming, out-of-order issue / execute core architecture may implement pipeline 800 as follows: 1) instruction fetch 838 performs fetch and length decode stages 802 and 804; 2) decode hardware 840 performs decode stage 806; 3) rename / arbiter hardware 852 performs arbitration stage 808 and rename stage 810; 4) scheduler hardware 856 performs scheduling stage 812; 5) physical register file(s) hardware 858 and memory hardware 870 perform register read / memory read stage 814; execution cluster 860 performs execution stage 816; 6) the memory hardware 870 and the physical register file(s) hardware 858 perform the write-back / memory write stage 818; 7) various hardware may participate in the exception handling stage 822;and 8) the retirement hardware 854 and the physical register file(s) hardware 858 perform the commit stage 824.;

[0046] The core 890 may support one or more instruction sets (e.g., the x86 instruction set (with some extensions added with newer versions); the MIPS instruction set from MIPS Technologies of Sunnyvale, CA; the ARM instruction set (with optional additional extensions, such as NEON) from ARM Holdings of Sunnyvale, CA) that include the instruction(s) described herein. In one embodiment, the core 890 includes logic to support a packed data instruction set extension (e.g., AVX1, AVX2, and / or some form of the generic vector-friendly instruction format (U=0 and / or U=1), described below) so that the operations used by many multimedia applications can be performed using packed data.

[0047] It should be understood that the core may support multithreading (executing two or more parallel sets of operations or threads) and may do so in several ways, including time-sliced multithreading, simultaneous multithreading (where a single physical core provides a logical core for each of the threads that the physical core is concurrently processing), or a combination thereof (e.g., time-sliced fetching and decoding followed by concurrent multithreading, such as in Intel® Hyperthreading Technology).

[0048] While register renaming is described in the context of out-of-order execution, it should be understood that register renaming may be used in an ordered architecture. While the illustrated embodiment of the processor also includes separate instruction and data cache hardware 834 / 874 as well as shared L2 cache hardware 876, alternative embodiments may include a single internal cache for both instructions and data, such as an internal Level 1 (L1) cache or multiple levels of internal cache. In some embodiments, the system may include a combination of an internal cache and an external cache external to the core and / or the processor. Alternatively, the entire cache may be external to the core and / or the processor.

[0049] Fig. Figure 9 is a block diagram of a processor 900 that may have more than one core, an integrated memory controller, and integrated graphics according to embodiments of the invention. The solid-line boxes in Fig. 9 illustrate a processor 900 with a single core 902A, a system agent 910, a set of one or more bus controller hardware parts 916, while the optional addition of the dashed boxes illustrates an alternative processor 900 with multiple cores 902A-N, a set of one or more integrated memory controller hardware parts 914 in the system agent hardware 910, and special purpose logic 908.

[0050] Thus, different implementations of processor 900 may include: 1) a CPU whose special-purpose logic 908 is integrated graphics and / or scientific logic (throughput) (which may include one or more cores), and cores 902A-N that are one or more general-purpose cores (e.g., ordered general-purpose cores, unordered general-purpose cores, a combination of the two); 2) a coprocessor whose cores 902A-N are a large number of special-purpose cores primarily dedicated to graphics and / or scientific purposes (throughput); and 3) a coprocessor whose cores 902A-N are a large number of ordered general-purpose cores. Thus, processor 900 may be a general-purpose processor, a coprocessor, or a special-purpose processor, such asa network or communications processor, a compression engine, a graphics processor, a GPGPU (general-purpose graphics processing unit), a high-throughput Many Integrated Core (MIC) architecture coprocessor (30 or more cores), an embedded processor, or the like. The processor may be implemented on one or more chips. The processor 900 may be part of and / or implemented on one or more substrates, utilizing any number of process technologies, such as BiCMOS, CMOS, or NMOS.

[0051] The memory hierarchy includes one or more levels of cache within the cores, a set of one or more shared cache hardware portions 906, and external memory (not shown) coupled to the set of integrated memory controller hardware 914. The set of shared cache hardware portions 906 may include one or more mid-level caches, such as Level 2 (L2), Level 3 (L3), Level 4 (L4), or other cache levels, a last-level cache (LLC), and / or combinations thereof. While in one embodiment, ring-based interconnect hardware 912 interconnects the integrated graphics logic 908, the set of shared cache hardware portions 906, and the system agent hardware 910 / integrated memory controller hardware 914, alternative embodiments may utilize any number of known techniques for interconnecting such hardware.In one embodiment, coherency is maintained between one or more cache hardware portions 906 and cores 902-AN.

[0052] In some embodiments, one or more of the cores 902A-N are multithreaded. The system agent 910 includes these components, which coordinate and operate the cores 902A-N. The system agent hardware 910 may include, for example, a power control unit (PCU) and display hardware. The PCU may be or include logic and components needed to regulate the power state of the cores 902A-N and the integrated graphics logic 908. The display hardware is used to drive one or more externally connected displays.

[0053] Cores 902A-N may be homogeneous or heterogeneous with respect to the architectural instruction set; that is, two or more cores 902A-N may be capable of executing the same instruction set, while others may be capable of executing only a subset of that instruction set or a different instruction set. In one embodiment, cores 902A-N are heterogeneous and include both the "small" and "large" cores described below.

[0054] The Fig. 10-13 are block diagrams of example computer architectures. Other system designs and configurations known in the art as laptops, desktops, handheld PCs, personal digital assistants, engineering workstations, servers, network devices, network hubs, switches, embedded processors, digital signal processors (DSPs), graphics devices, video game devices, set-top boxes, microcontrollers, cellular phones, portable media players, handheld devices, and various other electronic devices are also suitable. In general, a vast variety of systems or electronic devices capable of accommodating a processor and / or other execution logic as disclosed herein are generally suitable.

[0055] Now on Fig. Referring to Figure 10, a block diagram of a system 1000 is shown in accordance with an embodiment of the present invention. The system 1000 may include one or more processors 1010, 1015 coupled to a controller hub 1020. In one embodiment, the controller hub 1020 includes a graphics memory controller hub (GMCH) 1090 and an input / output hub (IOH) 1050 (which may be on separate chips); the GMCH 1090 includes memory and graphics controllers coupled to a memory 1040 and a coprocessor 1045; The IOH 1050 couples input / output (I / O) devices 1060 to the GMCH 1090. Alternatively, one or both of the memory and graphics controllers are integrated within the processor (as described herein), the memory 1040 and the coprocessor 1045 are directly coupled to the processor 1010, and the controller hub 1020 in a single chip is coupled to the IOH 1050.

[0056] The optional nature of additional processors 1015 is in Fig. 10 is indicated by dashed lines. Each processor 1010, 1015 may include one or more of the processor cores described herein and may be a version of processor 900.

[0057] Memory 1040 may be, for example, dynamic random access memory (DRAM), phase change memory (PCM), or a combination of both. For at least one embodiment, controller hub 1020 communicates with processor(s) 1010, 1015 via a multi-drop bus, such as a front-side bus (FSB), a point-to-point interface, or a similar connection 1095.

[0058] In one embodiment, coprocessor 1045 is a special-purpose processor, such as a high-throughput MIC processor, a network or communications processor, a compression engine, a graphics processor, a GPGPU, an embedded processor, or the like. In one embodiment, controller hub 1020 may include an integrated graphics accelerator.

[0059] There may be a variety of differences between the physical resources 1010, 1015 with respect to a spectrum of advantageous metrics, including architectural, microarchitectural, thermal, power consumption characteristics, and the like.

[0060] In one embodiment, processor 1010 executes instructions that control data processing operations of a general type. Coprocessor instructions may be embedded within the instructions. Processor 1010 recognizes these coprocessor instructions as being of a type that should be executed by the attached coprocessor 1045. Accordingly, processor 1010 issues these coprocessor instructions (or control signals representing coprocessor instructions) to coprocessor 1045 on a coprocessor bus or other connection. Coprocessor(s) 1045 accept(s) the received coprocessor instructions and execute(s) them.

[0061] Now on Fig. Referring to Figure 11, a block diagram of a first more detailed exemplary system 1100 is shown in accordance with an embodiment of the present invention. As shown in Fig. 11, multiprocessor system 1100 is a point-to-point interconnect system including a first processor 1170 and a second processor 1180 coupled via a point-to-point interconnect 1150. Each of processors 1170 and 1180 may be a version of processor 900. In one embodiment of the invention, processors 1170 and 1180 are processors 1010 and 1015, respectively, while coprocessor 1138 is coprocessor 1045. In another embodiment, processors 1170 and 1180 are processor 1010 and coprocessor 1045, respectively.

[0062] Processors 1170 and 1180 are shown as including integrated memory controller (IMC) hardware 1172 and 1182, respectively. Processor 1170 includes point-to-point (PP) interfaces 1176 and 1178 as part of its bus controller hardware; likewise, second processor 1180 includes PP interfaces 1186 and 1188. Processors 1170, 1180 may exchange information via a point-to-point (PP) interface 1150 using interface circuits 1178, 1188. As shown in Fig. 11, the IMCs 1172 and 1182 couple the processors to the corresponding memories, namely a memory 1132 and a memory 1134, which may be parts of the main memory attached locally to the corresponding processors.

[0063] Processors 1170, 1180 may each exchange information with a chipset 1190 via individual PP interfaces 1152, 1154 using point-to-point interface circuits 1176, 1194, 1186, 1198. Chipset 1190 may optionally exchange information with coprocessor 1138 via a high-performance interface 1139. In one embodiment, coprocessor 1138 is a special-purpose processor, such as a high-throughput MIC processor, a network or communications processor, a compression engine, a graphics processor, a GPGPU, an embedded processor, or the like.

[0064] A shared cache (not shown) may be included in both processors or external to both processors but connected to the processors via a PP interconnect so that the local cache data of one or both processors can be stored in the shared cache if one processor is placed into a power saving mode.

[0065] Chipset 1190 may be coupled to a first bus 1116 via an interface 1196. In one embodiment, first bus 1116 may be a Peripheral Component Interconnect (PCI) bus or a bus such as a PCI Express bus or other third-generation I / O interconnect bus, although the scope of the present invention is not so limited.

[0066] As in Fig. 11, various I / O devices 1114 may be coupled to the first bus 1116 along with a bus bridge 1118 that couples the first bus 1116 to a second bus 1120. In one embodiment, one or more additional processors 1115, such as coprocessors, high-throughput MIC processors, GPGPUs, accelerators (such as graphics accelerators or digital signal processing (DSP) hardware), field programmable gate arrays, or any other processor, are coupled to the first bus 1116. In one embodiment, the second bus 1120 may be a low pin count (LPC) bus. Various devices may be coupled to a second bus 1120 in one embodiment, including, for example, a keyboard and / or mouse 1122, communication devices 1127, and memory hardware 1128, such as a flash memory (RAM) 1130. B. a disk drive or other mass storage device that can contain instructions / code and data 1130.Furthermore, an audio I / O interface 1124 may be coupled to the second bus 1120. Note that other architectures are possible. For example, instead of the point-to-point architecture of . Fig. 11 implement a multi-drop bus or other such architecture.

[0067] Now on Fig. Referring to Figure 12, a block diagram of a second, more detailed exemplary system 1200 is shown in accordance with an embodiment of the present invention. Like elements in the Fig. 11 and Fig. 12 bear the same reference numerals, and certain aspects of Fig. 11 are from Fig. 12 has been omitted to avoid obscuring other aspects of Fig. 12 to be avoided.

[0068] Fig. Figure 12 illustrates that processors 1170, 1180 may each include integrated memory and I / O control logic ("CL") 1172 and 1182. Thus, CLs 1172, 1182 include integrated memory controller hardware and I / O control logic. Fig. Figure 12 illustrates that memories 1132, 1134 are not only coupled to CLs 1172, 1182, but also that I / O devices 1214 are also coupled to control logic 1172, 1182. Legacy I / O devices 1215 are coupled to chipset 1190.

[0069] Now on Fig. Referring to Figure 13, a block diagram of an SoC 1300 in accordance with an embodiment of the present invention is shown. Similar elements in Fig. 9 have the same reference numerals. Furthermore, boxes with dashed lines are optional features on more advanced SoCs. Fig. 13, interconnect hardware 1302 is coupled to: an application processor 1310 including a set of one or more cores 902A-N and shared cache hardware 906; system agent hardware 910; bus controller hardware 916; integrated memory controller hardware 914; a set of one or more coprocessors 1320, which may include integrated graphics logic, an image processor, an audio processor, and a video processor; static random access memory (SRAM) hardware 1330; direct memory access (DMA) hardware 1332; and display hardware 1340 for coupling to one or more external displays. In one embodiment, coprocessor(s) 1320 includes a special-purpose processor, such as a display processor. E.g., a network or communications processor, a compression engine, a GPGPU, a high-throughput MIC processor, an embedded processor, or the like.

[0070] Embodiments of the mechanisms disclosed herein may be implemented in hardware, software, firmware, or a combination of such implementation approaches. Embodiments of the invention may be implemented as computer programs or program code executing on programmable systems comprising at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0071] Program code, such as the one in Fig. Code 1130, illustrated in Figure 11, may be applied to input instructions to perform the functions described herein and generate output data. The output data may be applied to one or more output devices in a known manner. For purposes of this application, a processing system includes any system that includes a processor, such as a digital signal processor (DSP), a microcontroller, an application-specific integrated circuit (ASIC), or a microprocessor.

[0072] The program code can be implemented in a high-level procedural or object-oriented programming language to communicate with a processing system. If desired, the program code can also be implemented in assembly or machine language. Indeed, the mechanisms described here are not limited in scope to any particular programming language. In any case, the language can be a compiled or interpreted language.

[0073] One or more aspects of at least one embodiment may be implemented by representative instructions stored on a machine-readable medium representing various logic within the processor that, when read by a machine, causes the machine to generate logic for performing the techniques described herein. Such representations, known as "IP cores," may be stored on a tangible, machine-readable medium and provided to various customers or manufacturing facilities for loading into the manufacturing machines that actually manufacture the logic or processor.

[0074] Such machine-readable storage media may include, without limitation, non-transitory, tangible arrangements of articles manufactured or formed by a machine or device, including storage media such as hard disks, any other type of disk, including floppy disks, optical disks, compact disk read-only memories (CD-ROMs), rewritable compact disks (CD-RWs), and magneto-optical disks, semiconductor devices such as read-only memories (ROMs), random access memories (RAMs) such as dynamic random access memories (DRAMs), static random access memories (SRAMs), erasable programmable read-only memories (EPROMs), flash memory, electrically erasable programmable read-only memories (EEPROMs), phase change memories (PCMs), magnetic or optical cards, or any other type of media suitable for storing electronic instructions.

[0075] Accordingly, embodiments of the invention also include non-transitory, tangible machine-readable media containing instructions or design data, such as Hardware Description Language (HDL), that define structures, circuits, devices, processors, and / or system features described herein. Such embodiments may also be referred to as program products.

[0076] In some cases, an instruction converter can be used to convert an instruction from a source instruction set to a target instruction set. For example, the instruction converter can translate (e.g., using static binary translation, dynamic binary translation, including dynamic compilation), morph, emulate, or otherwise transform an instruction into one or more other instructions for processing by the core. The instruction converter can be implemented in software, hardware, firmware, or a combination thereof. The instruction converter can be on-processor, off-processor, or partially on- and partially off-processor.

[0077] Fig. Figure 14 is a block diagram contrasting the use of a software instruction converter to convert binary instructions in a source instruction set to binary instructions in a target instruction set according to embodiments of the invention. In the illustrated embodiment, the instruction converter is a software instruction converter, although the instruction converter may alternatively be implemented in software, firmware, hardware, or various combinations thereof. Fig. 14 shows that a program in a high-level language 1402 can be compiled using an x86 compiler 1404 to produce x86 binary code 1406 that can be natively executed by a processor having at least one x86 instruction set core 1416. The processor having at least one x86 instruction set core 1416 represents any processor that can perform substantially the same functions as an Intel processor having at least one x86 instruction set core by compatibly executing (1) a substantial portion of the instruction set of the Intel x86 instruction set core or (2) object code versions of applications or other software intended to run on an Intel processor having at least one x86 instruction set core or otherwise processing to achieve substantially the same result as an Intel processor having at least one x86 instruction set core.The x86 compiler 1404 represents a compiler operable to generate x86 binary code 1406 (e.g., object code) that can be executed with or without additional linkage processing on the processor having at least one x86 instruction set core 1416. Likewise, FIG. Fig.14, the program in the high-level language 1402 can be compiled using an alternative instruction set compiler 1408 to generate alternative instruction set binary code 1410 that can be natively executed by a processor without at least one x86 instruction set core 1414 (e.g., a processor with cores executing the MIPS instruction set from MIPS Technologies in Sunnyvale, CA, USA and / or the ARM instruction set from ARM Holdings in Sunnyvale, CA, USA). The instruction converter 1412 is used to convert the x86 binary code 1406 into code that can be natively executed by the processor without an x86 instruction set core 1414. This converted code is unlikely to be the same as the alternative instruction set binary code 1410, because an instruction converter capable of doing so is difficult to manufacture; but the converted code will achieve the general function and be composed of instructions from the alternative instruction set.Thus, the instruction converter 1412 represents software, firmware, hardware, or a combination thereof that allows a processor or other electronic device that does not have an x86 instruction set processor or core to execute the x86 binary code 1406 through emulation, simulation, or any other process.

Claims

[1] Device comprising: a system memory (300); a plurality of hardware processor cores, at least one of which includes; a first cache (311, 312, 320, 321); a decoder circuit (330) for decoding an instruction having fields for a first memory address and a region indicator, the first memory address and the region indicator defining a contiguous region in the system memory (300), the contiguous region comprising one or more cache lines; and an execution circuit (340) for executing the decoded instruction by invalidating any instances of the one or more cache lines in the first cache (311, 312, 320, 321); wherein an invalidated instance of the one or more cache lines in the first cache (311, 312, 320, 321) is to be stored in the system memory (300) if the invalidated instance is dirty; characterized by , that the instruction further includes an opcode to indicate whether the dirty invalidated instance of the one or more cache lines in the first cache (311, 312, 320, 321) is to be stored in a second cache (311, 312, 320, 321) shared by the plurality of hardware processor cores instead of the system memory. [2] The apparatus of claim 1, wherein the range indicator comprises a second memory address and the contiguous range extends from the first memory address to the second memory address. [3] The apparatus of claim 1 or 2, wherein the range indicator comprises an indication of a number of bytes or cache lines to be included in the contiguous range, the included cache lines having an address equal to or incrementally greater than the first memory address. [4] The apparatus of any of claims 1-3, wherein the first memory address is a linear memory address. [5] Device comprising: a system memory (300); a plurality of hardware processor cores, at least one of which includes; a first cache (311, 312, 320, 321); a decoder circuit (330) for decoding an instruction having fields for a first memory address and a region indicator, the first memory address and the region indicator defining a contiguous region in the system memory (300), the contiguous region comprising one or more cache lines; and an execution circuit (340) for executing the decoded instruction by causing any dirty instances of the one or more cache lines in the first cache (311, 312, 320, 321) to be stored in the system memory (300); characterized by, that the instruction further includes an opcode to indicate whether a dirty instance of the one or more cache lines in the first cache (311, 312, 320, 321) is to be stored in a second cache (311, 312, 320, 321) shared by the plurality of hardware processor cores instead of the system memory. [6] The apparatus of claim 5, wherein the range indicator comprises a second memory address and the contiguous range extends from the first memory address to the second memory address. [7] The apparatus of claim 5 or 6, wherein the range indicator comprises an indication of a number of bytes or cache lines to be included in the contiguous range, the included cache (311, 312, 320, 321) lines having an address equal to or incrementally greater than the first memory address. [8] The apparatus of any of claims 5-7, wherein the first memory address is a linear memory address. [9] Procedure comprising: Decoding an instruction having fields for a first memory address and a range indicator, the first memory address and the range indicator defining a contiguous range in a system memory (300), the contiguous range comprising one or more cache lines; and executing the decoded instruction by invalidating any instances of the one or more cache lines in a first processor cache (311, 312, 320, 321); wherein executing the decoded instruction further comprises storing an invalidated instance of the one or more cache lines of the first processor cache in the system memory (300) if the invalidated instance is dirty; characterized by , that executing the decoded instruction further comprises storing an invalidated instance of the one or more cache lines of the first processor cache in a second cache (311, 312, 320, 321) if the invalidated instance is dirty, wherein the second cache (311, 312, 320, 321) is shared by a plurality of hardware processor cores. [10] The method of claim 9, wherein the range indicator comprises a second memory address and the contiguous range extends from the first memory address to the second memory address. [11] The method of claim 9 or 10, wherein the range indicator comprises an indication of a number of bytes or cache lines to be included in the contiguous range, the included cache lines having an address equal to or incrementally greater than the first memory address. [12] The method of any of claims 9-11, wherein the first memory address is a linear memory address. [13] Procedure comprising: Decoding an instruction having fields for a first memory address and a range indicator, the first memory address and the range indicator defining a contiguous range in a system memory (300), the contiguous range comprising one or more cache lines; and executing the decoded instruction by causing any dirty instances of the one or more cache lines to be stored in a first processor cache (311, 312, 320, 321) in the system memory (300); characterized by , that the instruction further includes an opcode to indicate whether a dirty instance of the one or more cache lines in the first processor cache (311, 312, 320, 321) is to be stored in a second cache (311, 312, 320, 321) shared by a plurality of hardware processor cores instead of the system memory. [14] The method of claim 13, wherein the range indicator comprises a second memory address and the contiguous range extends from the first memory address to the second memory address. [15] The method of claim 13 or 14, wherein the range indicator comprises an indication of a number of bytes or cache lines to be included in the contiguous range, the included cache lines having an address equal to or incrementally greater than the first memory address. [16] The method of any of claims 13-15, wherein the first memory address is a linear memory address.

Citation Information

Patent Citations

  • Interruptible an re-entrant cache clean range instruction

    US6772326B2