Computer system and method for pre-accessing insensitive transactional memory

DE102016219651B4Active Publication Date: 2025-07-10INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
DE102016219651
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2016-06-20
Filing Date
2016-10-11
Publication Date
2025-07-10
Estimated Expiration
2036-10-11

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A computer system (1101) for preventing a prefetch memory operation from causing a transaction to abort, the computer system (1101) comprising: one or more computer processors (804), one or more computer-readable storage media, and program instructions stored on the one or more computer-readable storage media for execution by at least one of the one or more computer processors (804), the program instructions comprising: Program instructions for receiving a pre-access request (4416) from a remote processor by a local processor; Program instructions for determining whether the prefetch request (4416) conflicts with a transaction of the local processor; Program instructions for responding to at least one of i) a determination that the local processor does not have a transaction and ii) a determination that the prefetch request (4416) does not conflict with a transaction by providing the requested prefetch data; and Program instructions for responding to a determination that the pre-access request (4416) conflicts with a transaction by suppressing processing of the pre-access request (4416), wherein the computer system (1101) comprises a transaction-bound memory system (A1104), and further wherein into the transactional memory system (A1104) non-prefetch operations (ie non-prefetch memory accesses) on a non-prefetch bus (2210) and prefetch operations (iePrefetch request memory accesses) on a bus for prefetch operations (2209), and an address decryption device (2211) decrypts the address of a non-prefetch operation and confirms the target address of a non-prefetch read operation on a read address bus (2201) and the target address of a non-prefetch write operation on a write address bus (2202), and further an address decryption device (2212) decrypts the address of a prefetch operation and confirms the target address of a prefetch read operation on a bus for a prefetch read address (2203) and the target address of a prefetch write operation on a bus for a prefetch write address (2204), and the prefetch manager (2208) is connected to an abort manager (2207) and receives all prefetch operations (2209) that are input to the transactional memory system (A1104) from the prefetch bus (2209).
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND OF THE INVENTION

[0001] The present invention relates generally to the field of computer memory management and, more particularly, to techniques for improving the efficiency of transactional memory operations.

[0002] Many computer systems use a cache to speed up data retrieval operations. A cache stores copies of data found at frequently used main memory locations. Accessing data from a cache speeds up processing because the cache is typically faster to access than main memory. If requested data is found in the cache, it is accessed from the cache. However, if requested data is not found in the cache, the data is first copied to the cache, and then accessed from the cache.

[0003] JERGER, NDE; HILL, EL; LIPASTI, MH: Friendly fire: understanding the effects of multiprocessor prefetches. In: 2006 IEEE International Symposium on Performance Analysis of Systems and Software, 2006, pp. 177-188. DOI: 10.1109 / ISPASS.2006.1620802 shows that modern processors attempt to overcome increasing memory latencies by anticipating future references and prefetching these blocks from memory. The behavior and potential negative side effects of prefetching techniques are fairly well known in uniprocessor systems. However, in a multiprocessor system, a prefetch can steal read and / or write permissions to shared blocks from other processors, leading to permission problems and an overall performance degradation. In this paper, we present a taxonomy that classifies the effects of multiprocessor prefetches.

[0004] HARRIS, Tim; LARUS, James; RAJWAR, Ravi: Transactional Memory. Synthesis Lectures on Computer Architecture. 2nd edition. Williston (USA): Morgan & Claypool Publishers, 2010. pp. i-iii, v-xi, 7-8, 30, 51-52, 147-152, 154-158. ISBN: 9781608452354. DOI:10.2200 / S00272ED1V01Y201006CAC011 shows that with the advent of multicore processors, interest in the idea of integrating transactions into the programming model for writing parallel programs has been reignited. This approach, known as transactional memory, offers an alternative and hopefully better way to coordinate concurrent threads. The ACI (atomicity, consistency, isolation) properties of transactions ensure that concurrent reads and writes of shared data do not produce inconsistent or incorrect results.At a higher level, a computation embedded in a transaction is executed atomically—either it completes successfully and the result is fully committed, or it is aborted. Furthermore, isolation ensures that the transaction produces the same result as if no other transactions were executing concurrently. Although transactions are not a panacea for parallel programming, they shift much of the burden of synchronizing and coordinating parallel computations from the programmer to a compiler, a language runtime, or the hardware. The challenge for system implementers is to build an efficient transactional memory infrastructure. This book provides an overview of the current state of the art in the design and implementation of transactional memory systems (as of Spring 2010).

[0005] HARRIS, Tim; LARUS, James; RAJWAR, Ravi: Transactional Memory. Synthesis Lectures on Computer Architecture. 2nd edition. Williston (USA): Morgan & Claypool Publishers, 2010. Title page + Imprint + Table of contents + pp. 147-204. ISBN 978-3-031-01728-5 DOI: 10.1007 / 978-3-031-01728-5 presents the most important techniques proposed for hardware implementations of transactional memory. Although STM systems are very flexible, HTM systems can offer several key advantages over them: • • HTM systems can typically run applications with lower overhead than STM systems. Therefore, they rely less on compiler optimizations for performance than STM systems. • • HTM systems can have better power and energy profiles. • HTM systems can be less invasive in an existing execution environment: For example, some HTM systems implicitly treat all memory accesses within a transaction as transactional and accommodate third-party transaction-safe libraries. • HTM systems can provide strong isolation without requiring changes to non-transactional memory accesses. • HTM systems are well suited for system languages such as C / C++ that work without dynamic compilation, garbage collection, etc.

[0006] RAJWAR, Ravi; GOODMAN, James R.: Transactional Lock-free execution of lock-based programs. In: Proceedings of the 10th international conference on architectural support for programming languages and operating systems, Oct. 6-9, 2002, San Jose, California. ACM, 2002. pp. 5-17. ISBN 1-58113-574-2 DOI: 10.1145 / 635508.605399 was written because of the difficulty of writing correct, high-performance programs. Writing multithreaded programs with shared memory requires a complex trade-off between ease of programming and performance, largely due to intricacies in coordinating access to shared data. To ensure correctness, programmers often rely on conservative locking, at the expense of performance. The resulting serialization of threads is a performance bottleneck.Locks also interact poorly with thread scheduling and errors, leading to poor system performance. We attempt to improve the trade-offs in multithreaded programming by providing architectural support for optimistic lock-free execution. In lock-free execution, shared objects are never locked when accessed by different threads. We propose Transactional Lock Removal (TLR) and show how a program that uses lock-based synchronization can be executed lock-free by hardware, even in the presence of contention, without requiring any support from the programmer or any software changes. TLR uses timestamps for contention resolution, modest hardware, and features already present in many modern computer systems. The benefits of TLR include improved programmability, stability, and performance.Programmers can take advantage of lock-free data structures, such as non-blocking behavior and wait-freeness, while using lock-protected critical sections to write programs.

[0007] IYER, Ravi: CQoS: A framework for enabling QoS in shared caches of CMP platforms. In: Proceedings of the 18th annual international conference on supercomputing, June 26 - July 1, 2004, Saint-Malo, France. ACM, 2004. pp. 257-266. ISBN 1-58113-839-3 DOI: 10.1145 / 1006209.1006246 shows that cache hierarchies have traditionally been designed for use by a single application, thread, or core. As multi-threaded (MT) and multi-core platform (CMP) architectures emerge and their workloads range from single-threaded and multi-threaded applications to complex virtual machines (VMs), a shared cache resource is used by these different units, generating heterogeneous memory access streams with different locality properties and varying memory sensitivity.Therefore, conventional cache management approaches that treat all memory accesses equally inevitably lead to inefficient memory usage and poor performance, even for applications with good locality properties. To address this problem, this paper presents a new cache management (CQoS) framework that (1) recognizes the heterogeneity in memory access flows, (2) introduces the notion of QoS to handle the varying degrees of locality and latency sensitivity, and (3) assigns and enforces priorities to flows based on latency sensitivity, degree of locality, and application performance requirements.

[0008] US 2015 / 0212 851 A1 shows that access to at least one memory location is provided by one of multiple transactions in a multi-processor transaction execution environment. This includes a computer system assigning a contention priority to a transaction; based on the occurrence of a contention with another process for a memory location, the computer system comparing the assigned contention priority of the transaction with a different priority of the other process; and based on the conflicting priority of the transaction being the higher priority, continuing the transaction; and based on the conflicting priority of the transaction being the lower priority, aborting the transaction.

[0009] US 2003 / 0 163 649 A1 shows a shared bypass bus structure for low-latency access to a coherence controller in a coherent scalable switch. In a coherent scalable switch with multiple coherent interconnect ports, distributed coherence control structures, and a crossbar interface between them, a shared bypass bus enables data transfer between the coherent interconnect ports and the coherence control structures while bypassing the crossbar interface. Some embodiments may include scalable switches to support one or more sets of processors with substantially independent snoop or cache coherence paths or arrangements.

[0010] A multi-level cache is a structure with multiple caches. For example, a data processing system may have three levels: an L1 cache, an L2 cache, and an L3 cache. Typically, L1 in a multi-level cache configuration is the smallest and has a short access time. If requested data is not found in the L1 cache, the system searches the L2 cache, which is usually larger and physically farther away than the L1 cache and thus has a longer access time. Similarly, if the data is not found in the L2 cache, the L3 cache is searched. Main memory is only accessed if the requested data is not in the L1, L2, or L3 caches. There are many different implementations of caches.

[0011] To improve performance, a computer architecture often includes prefetch instructions that can move data from main memory or a lower cache level to a higher cache level (closer to the processor) in anticipation of an access to the data. The execution of a prefetch instruction can be speculative in nature in that it can be executed before it is known that the data moved by the prefetch instruction will actually be accessed. Data that is expected to be either read or written can be prefetched, referred to as read prefetch or write prefetch, respectively. Cache coherence protocols in computers that enable data sharing, synchronization, and parallel processing often require a processor to gain ownership of the data before it is overwritten (i.e.changed). Because data may currently be under the responsibility of another processor, acquiring responsibility can be a lengthy process. The concept of responsibility is necessary to prevent data read by one processor from one memory location from being changed by another processor writing to the same memory location or a different memory location (when copies of the data are located in several different memory locations). Therefore, it is often advantageous for a processor to execute a write prefetch instruction to acquire responsibility over the data in anticipation of the data being overwritten, so that processing is not delayed once responsibility is acquired.

[0012] Because cache access time is often critical to the performance of executing code, and a cache is often busy with many operations (e.g., servicing misses), it is beneficial to reduce cache workload whenever possible. A common technique used to reduce cache workload involves aggregating multiple store operations stored in a common cache line in a cache line buffer, and then storing the contents of the cache line buffer in a cache as a single operation. This reduces a cache's workload and improves its response time, thus potentially improving the performance of executing code. Such a technique is generally implemented in a mechanism called a memory cache.

[0013] A transactional load is a type of memory operation that groups one or more load and store operations performed by a processor into a single transaction that is visible to other processors as a single operation when the transaction completes. The effects (e.g., the data) of multiple memory operations participating in the single transaction are not made visible to other processors until the transaction completes. A transactional load is a load that accesses and stores data until a transaction completes, after which the data is passed on to the processor that requested it. If the data from a transactional load is modified after the load occurs and before the transaction completes (i.e.,their memory location is overwritten), the transaction aborts. A transactional write is a write whose data is not visible until the transaction completes. A read from or overwrite to a memory location that is the target of a transactional write in a transaction aborts the transaction. Transactional memory is often useful for synchronizing work that is performed in parallel on multiple CPUs (by allowing atomic operations on any set of memory locations) and when multiple write operations must not be interrupted. Because write operations that participate in a transaction are not visible until the transaction completes, a read from a memory location that is currently being overwritten by a write in the transaction aborts the transaction. SUMMARY OF THE INVENTION

[0014] Embodiments of the present invention provide a method and computer system for preventing a prefetch memory operation from causing a transaction to be aborted. According to the invention, a local processor of a group of one or more processors receives a prefetch request from a remote processor. According to the invention, a processor of the group of one or more processors determines whether the prefetch request conflicts with a transaction of the local processor. According to the invention, a processor of the group of one or more processors responds to at least one of i) a determination that the local processor does not have a transaction and ii) a determination that the prefetch request does not conflict with a transaction by providing requested prefetch data.According to the invention, a processor of the group of one or more processors responds to a determination that the prefetch request conflicts with a transaction by suppressing processing of the prefetch request. BRIEF DESCRIPTION OF THE DIFFERENT VIEWS OF THE DRAWINGS Fig. 1 illustrates an exemplary transactional multi-core memory environment according to an illustrative embodiment; Fig. 2 illustrates an exemplary transactional multi-core memory environment according to an illustrative embodiment; Fig. 3 illustrates example components of an example CPU according to an illustrative embodiment; Fig. 4 illustrates a block diagram of a data processing system according to an embodiment of the present invention; Fig. 5 illustrates a Fig.4 illustrates a caching system according to an embodiment of the present invention; Fig. 6 illustrates a format of a read request and a response to a read request according to an embodiment of the present invention; Fig. 7 illustrates a format of a pre-access request according to an embodiment of the present invention; Fig. 8 illustrates a format of a packed prefetch request according to an embodiment of the present invention; Fig. 9 illustrates a flow chart for an operation of the Fig. 5 illustrates a transactional storage system according to an embodiment of the present invention; Fig. 10 illustrates a flow chart for an operation of the Fig. 5 illustrated pre-access manager according to an embodiment of the present invention; and Fig. Figure 11 illustrates a block diagram of a computer system embodying the Fig. 4 and Fig. 5 illustrated buffer system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0015] Detailed embodiments of the present invention are disclosed herein with reference to the accompanying drawings. It should be understood that the disclosed embodiments are merely illustrative of potential embodiments of the present invention and may take various forms. Furthermore, each example given in connection with the various embodiments is to be considered illustrative but not restrictive. Further, the figures are not necessarily to scale; some features may be exaggerated to show details of particular components. Therefore, specific structural and functional details disclosed herein are not to be interpreted as limiting, but merely as a representative basis for teaching one skilled in the art to variously utilize the present invention.

[0016] References in the specification to "one embodiment," "an embodiment," "any embodiment," an "exemplary embodiment," etc., indicate that the described embodiment may include a particular feature, structure, or characteristic, but that not every embodiment necessarily includes the particular feature, structure, or characteristic. Furthermore, such language does not necessarily refer to the same embodiment. Further, where a particular feature, structure, or characteristic is described in connection with one embodiment, it is believed that it is within the skill of one of ordinary skill in the art to classify such features, structures, or characteristics in connection with other embodiments, regardless of whether they are explicitly described.

[0017] A memory hierarchy in a modern computer often contains multiple layers of cache, with some layers dedicated to and accessible by a single processor, and other, lower, and higher layers of cache accessible by multiple processors. A cache often provides fast access to recently accessed data or to data near recently accessed data. Caches are named in order of their logical position relative to a particular processor: L1, L2, L3, etc., with the L1 cache being logically closest to the processor. This arrangement is also the order in which the caches are accessed when the processor attempts to read data. L1 is accessed first for the data, followed by L2 if the data is not found in L1, and so on.Sometimes an L0 cache is used, which is small and hard-wired into the processor, often providing 1-cycle access. When present, an L0 cache is logically closer to the processor than an L1 cache. The cache levels in a computer system, together with main memory (often a large amount of dynamic RAM), form the computer system's memory hierarchy. In the context of a memory hierarchy, the term "low" means logically farther from the processor.

[0018] Many techniques have been developed to increase the efficiency of a memory hierarchy. The efficiency of the memory hierarchy relative to a benchmark program is usually measured by the average access time during the execution of the benchmark program. A memory cache is a technique used to improve the average access time by reducing the workload of a cache in the next cache layer below that of the memory cache. This is done by accumulating stores to a cache line in a buffer and then writing the contents of the buffer to the cache line in the next cache layer in one access, thereby preventing multiple accesses that would have occurred if each separate store performed a memory operation.

[0019] In modern multi-processor computing systems, there is a desire to increase performance by processing parts of a program in parallel on the same processor, if it is multi-threaded and / or across multiple different processors, and combining or comparing results as needed at intervals during execution. This is often achieved by synchronizing multiple threads of execution on the same or different processors and / or by making results produced by one thread of execution visible to other threads of execution. Synchronization is often achieved by executing "atomic" instructions and groups of instructions. An atomic (atomic-sized) instruction appears to execute "all at once" to other threads of execution and other processors; that is, the atomic instruction can never be observed to be partially completed.Likewise, a group of instructions can be made atomic by making their effect visible to other threads and processors at once. Memory operations are often available in an atomic version because multiple threads of execution often exchange data with each other and synchronize themselves via values written to or read from memory locations known to all participating threads of execution. For example, an atomic instruction can read a memory location and, if the memory location contains a particular value, write a different value back to the same memory location in a single atomic operation. This action would be reported to other processes that might test the variable (by reading it).

[0020] A modern technique that can improve the performance of an application running on a processor is to execute some instructions speculatively when resources are available. Speculative execution is a term that refers to execution that is likely to occur in the future, but may not occur. This usually occurs when a branch instruction is detected in the instruction stream and its behavior is predicted based on past behavior, because the information that determines its actual behavior is not yet available. Instead of waiting for that information to become available, the processor can operate based on a predicted path that that branch is likely to take and execute instructions along that predicted path.Instructions executed on this predicted path are speculative instructions until it is known that they are on an actual execution path—that is, that the branch was correctly predicted. If these speculative instructions turn out to be on the wrong actual execution path (misprediction), the effects of these instructions must be suppressed—that is, eliminated, undone, and made invisible to other processors.

[0021] Instructions that have been speculatively executed but should not have been executed can pose a demanding design challenge, and are particularly challenging with respect to speculative store instructions that should not have been executed. If a store operation that should not have been executed stores data in memory, it may overwrite data that should not have been overwritten, and the written data may be read and used by the same or a different processor in an application. Computer architectures address this problem by not executing speculative store instructions, or by executing them speculatively and storing store results in a store buffer or store queue until the instruction's in-order point (which reduces performance).An alternative approach is to prevent other threads of execution from seeing the data produced by a speculative store operation until it is known to be on the correct path, and to delete the data if it is a result of a store operation on an incorrect path.

[0022] Concurrent multithreading is a technique often incorporated into modern processors that allows a single processor to run multiple applications (or multiple parts of the same application) simultaneously, with each application having its own thread of execution. The single processor fetches instructions on each thread separately and executes the instructions on shared execution units (e.g., adders, multipliers, etc.) within the processor, constantly keeping track of which instructions belong to which thread. When such threads synchronize their work with other threads, they often do so through the execution of atomic instructions or groups of instructions whose execution is made atomic.

[0023] Historically, a computer system, or processor, had only a single processor (also known as a processing unit or central processing unit). The processor included an instruction processing unit (IPU), a branch unit, a memory control unit, and the like. Such processors could only execute a single thread of a program at a time. Operating systems were designed to time-share a processor by allocating one program to run on the processor for one period of time, and then allocating another program to run on the processor for another period of time. As technology advanced, memory subsystem caches were often added to the processor, as well as complex dynamic address translation, including translation buffers (TLBs). The IPU itself was often referred to as a processor.As technology advanced, an entire processor could be packaged onto a single semiconductor chip, or microchip; such a processor was called a microprocessor. Subsequently, processors were developed that integrated multiple IPUs; such processors were often referred to as multiprocessors. Each such processor of a multiprocessor computer system (processor) may contain individual or shared caches, memory interfaces, system buses, address translation mechanisms, and the like. Virtual machine emulators and instruction set architecture (ISA) added a software layer to a processor that provided the virtual machine with multiple "virtual processors" (also called processors) by time-slicing a single IPU in a single hardware processor.As technology continued to evolve, multithreaded processors were developed, allowing a single hardware processor with a single multithreaded IPU to provide the capability for concurrent execution of threads from different programs, so that each thread of a multithreaded processor appeared to the operating system as a processor. As technology continued to evolve, it became possible to package multiple processors (each with an IPU) on a single semiconductor chip or microchip. These processors were referred to as processor cores or simply cores. Therefore, terms such as processor, central processing unit, processing unit, microprocessor, core, processor core, processor thread, and thread, for example, are often used interchangeably. Aspects of embodiments herein may be practiced by any or all processors, including those identified above, without departing from the teachings herein.Where the term "thread" or "processor thread" is used herein, it is expected that a particular advantage of the embodiment will result in a processor thread implementation. Transaction execution in Intel®-based embodiments

[0024] In "Intel® Architecture Instruction Set Extensions Programming Reference" 319433-012A, February 2012, incorporated herein by reference in its entirety, Chapter 8 teaches in part that multithreaded applications can take advantage of the increasing number of CPU cores to achieve higher performance. However, writing multithreaded applications requires programmers to understand and account for data sharing among the multiple threads. Access to shared data typically requires synchronization mechanisms. These synchronization mechanisms are used to ensure that multiple threads update shared data by serializing operations applied to the shared data, often through the use of a critical section protected by a lock.Because serialization limits concurrency, programmers try to limit the system overhead due to synchronization.

[0025] Intel® Transactional Synchronization Extensions (Intel® TSX) allow a processor to dynamically determine whether threads need to be serialized through lock-protected critical sections and to perform serialization only when necessary. This allows the processor to expose and exploit shared access hidden in an application due to dynamically unnecessary synchronization.

[0026] With Intel TSX, programmer-specified regions of code (also called "transactional regions" or just "transactions") execute transactionally. When transactional execution completes successfully, all memory operations performed within the transactional region appear concurrent to other processors. A processor only makes the memory operations of the executing transaction executing within the transactional region visible to other processors upon a successful commit—that is, when the transaction successfully completes execution. This process is often referred to as atomic commit.

[0027] Intel TSX provides two software interfaces for specifying code regions for transactional execution. Hardware Lock Elision (HLE) is a traditionally compatible instruction group extension (with the XACQUIRE and XRELEASE prefixes) for specifying transactional regions. Restricted Transactional Memory (RTM) is a new instruction group interface (with the XBEGIN, XEND, and XABORT instructions) for programmers to define transactional regions in a more flexible manner than is possible with HLE. HLE is for programmers who prefer the backward compatibility of the traditional mutual exclusion programming model and would like to run HLE-enabled software on traditional hardware, but would also like to take advantage of the new lock eliminator capabilities on HLE-enabled hardware.RTM is for programmers who prefer a flexible interface to transactional execution hardware. Intel TSX also provides an XTEST instruction. This instruction allows software to query whether the logical processor is executing transactionally in a transactional region identified by either HLE or RTM.

[0028] Because successful transactional execution ensures atomic commit, the processor optimistically executes the code region without explicit synchronization. If synchronization was unnecessary for that particular execution, the execution can be committed without cross-thread serialization. If the processor cannot make an atomic commit, optimistic execution fails. When this happens, the processor rolls back the execution, a process known as transactional abort. In transactional abort, the processor discards all updates made to the memory region used by the transaction, restores the architectural state so that it appears as if optimistic execution never occurred, and resumes execution non-transactionally.

[0029] A processor can perform a transactional abort for numerous reasons. A primary reason for a transaction to abort is due to conflicting memory accesses between the transactional logical processor and another logical processor. Such conflicting memory accesses can prevent successful transactional execution.

[0030] Memory addresses read from a transactional region form the transactional region's read set, and addresses written to the transactional region form the transactional region's write set. Intel TSX manages the read and write sets at the granularity of a cache line. A conflicting memory access occurs when another logical processor either reads a memory location that is part of the transactional region's write set or writes to a memory location that is part of either the transactional region's read or write set. A conflicting access typically means that serialization is required for that region of code.Because Intel TSX detects data conflicts at the granularity of a cache line, unrelated data locations placed within the same cache line are detected as conflicts, resulting in transactional aborts. Transactional aborts can also occur due to limited transactional resources. For example, the amount of data accessed in the range may exceed an implementation-specific capacity. In addition, some instructions and system events can cause transactional aborts. Frequent transactional aborts result in unnecessary cycles and increased inefficiency. Hardware Lock Elision

[0031] Hardware Lock Elision (HLE) provides a traditional, compatible instruction group interface for programmers to use transactional execution. HLE provides two new instruction prefix hints: XACQUIRE and XRELEASE.

[0032] With HLE, a programmer adds the XACQUIRE prefix before the instruction used to acquire the lock protecting the critical section. The processor treats the prefix as a hint to omit the write operation associated with the lock acquisition operation. Even if the lock acquisition includes a write operation to the lock, the processor does not add the lock address to the write group of the transactional region, nor does it issue any write requests to the lock. Instead, the lock address is added to the read group. The logical processor begins transactional execution. If the lock was available before the XACQUIRE-prefixed instruction, all other processors continue to see the lock as available afterward.Because the transactional logical processor has neither added the lock's address to its write group nor performed any externally visible write operations on the lock, other logical processors can read the lock without causing a data conflict. This allows other logical processors to also enter the section protected by the lock and execute it concurrently. The processor automatically detects any data conflicts that occur during transactional execution and performs a transactional abort if necessary.

[0033] Even if the omitting processor has not performed any external write operations on the lock, the hardware ensures the program order of operations on the lock. If the omitting processor itself reads the value of the lock in the critical section, it appears as if the processor has acquired the lock, meaning the read operation returns the non-omitted value. This behavior allows an HLE execution to be functionally equivalent to an execution without the HLE prefixes.

[0034] An XRELEASE prefix can be added before an instruction used to release the lock protecting the critical section. A lock release involves a write operation on the lock. If the instruction is to restore the lock's value to the value it had before the XACQUIRE-prefixed lock acquisition operation on the same lock, the processor skips the external write request associated with the lock release and does not add the lock's address to the write set. The processor then attempts to commit transactional execution.

[0035] If multiple threads with HLE execute critical sections protected by the same lock but do not perform conflicting operations on each other's data, the threads can execute concurrently without serialization. Even if the software uses lock acquisition operations on a shared lock, the hardware detects this, omits the lock, and executes the critical sections on both threads without requiring any data exchange across the lock—assuming such data exchange was dynamically unnecessary.

[0036] If the processor is unable to execute the scope transactionally, the processor executes the scope non-transactionally and without egress. HLE-aware software has the same forwarding progress guarantees as the underlying non-HLE lock-based execution. For successful HLE execution, the lock and critical section code must follow certain guidelines. These guidelines only affect performance; failure to follow these guidelines does not result in functional failure. Hardware without HLE support ignores the XACQUIRE and XRELEASE prefix hints and does not perform egress, as these prefixes correspond to the REPNE / REPE IA-32 prefixes, which are ignored in instructions where XACQUIRE and XRELEASE are valid. Most importantly, HLE is compatible with the existing lock-based programming model.Improper use of hints does not lead to functional errors, although it may reveal latent errors already present in the code.

[0037] Restricted Transactional Memory (RTM) provides a flexible software interface for transactional execution. RTM provides three new instructions—XBEGIN, XEND, and XABORT—for programmers to start, commit, and abort transactional execution.

[0038] The programmer uses the XBEGIN instruction to specify the beginning of a transactional code region and the XEND instruction to specify the end of the transactional code region. If the RTM region could not be successfully executed in a transactional manner, the XBEGIN instruction takes an operand that provides a relative offset to the reset instruction address.

[0039] A processor can abort transactional RTM execution for many reasons. In many cases, the hardware automatically detects transactional abort conditions and restarts execution from the reset instruction address with the architectural state corresponding to that which existed at the beginning of the XBEGIN instruction and the EAX register updated to describe the abort status.

[0040] The XABORT instruction allows programmers to explicitly abort an RTM region. The XABORT instruction takes an 8-bit immediate argument, which is loaded into the EAX register, making it available to software after an RTM abort. RTM instructions do not have a data storage location associated with them. While hardware offers no guarantees that an RTM region can truly be successfully committed transactionally, most transactions that follow recommended guidelines are expected to be successfully committed transactionally. However, programmers must always provide an alternative code sequence on the rollback path to guarantee forward progress. This can be as simple as acquiring a lock and executing the specified code region non-transactionally.Furthermore, a transaction that always aborts in a particular implementation may be completed transactionally in a future implementation. Therefore, programmers must ensure that the code paths for the transactional scope and the alternative code sequence are tested for functionality. Detection of HLE support

[0041] A processor supports HLE execution if CPUID.07H.EBX.HLE [bit 4] = 1. However, an application can use the HLE prefixes (XACQUIRE and XRELEASE) without checking whether the processor supports HLE. Processors without HLE support ignore these prefixes and execute the code without entering transactional execution. Detection of RTM support

[0042] A processor supports RTM execution if CPUID.07H.EBX.RTM [bit 11] = 1. An application must check whether the processor supports RTM before using the RTM instructions (XBEGIN, XEND, XABORT). These instructions generate a #UD exception if used on a processor that does not support RTM. Detection of an XTEST instruction

[0043] A processor supports the XTEST instruction if it supports either HLE or RTM. An instruction must check both of these characteristic flags before using the XTEST instruction. This instruction generates a #UD exception if used on a processor that does not support either HLE or RTM. Querying the transaction-related execution status

[0044] The XTEST instruction can be used to determine the transactional status of a transactional region specified by HLE or RTM. Note that although the HLE prefixes are ignored on processors that do not support HLE, the XTEST instruction generates a #UD exception when used on processors that do not support HLE or RTM. Requirements for HLE locks

[0045] For an HLE execution to be successfully committed transactionally, the lock must satisfy certain properties and access to the lock must follow certain policies.

[0046] An instruction prefixed with XRELEASE must restore the value of the omitted lock to the value it had before the lock was acquired. This allows the hardware to safely omit locks by not adding them to the write group. The data size and data address of the lock release instruction (prefixed with XRELEASE) must match those of the lock acquisition instruction (prefixed with XACQUIRE), and the lock must not cross a cache line boundary.

[0047] Software should not overwrite the omitted lock in a transactional HLE section with any instruction other than an XRELEASE-prefixed instruction; otherwise, such a write operation may cause a transactional abort. Furthermore, recursive locks (where a thread acquires the same lock multiple times without first releasing the lock) may also cause a transactional abort. Note that software can observe the result of the omitted lock acquisition in the critical section. Such a read operation returns the value of the lock overwrite.

[0048] The processor automatically detects violations of these policies and safely transitions to non-transactional execution without omission. Because Intel TSX detects conflicts at the granularity of a cache line, overwrite operations of data residing on the same cache line as the omitted lock can be detected as data conflicts by other logical processors omit the same lock. Transaction-based nesting

[0049] HLE and RTM support nested transactional scopes. However, a transactional abort restores the state to the operation that initiated transactional execution: either the outermost HLE-legitimate XACQUIRE-prefixed instruction or the outermost XBEGIN instruction. The processor treats all nested transactions as a single transaction. HLE nesting and omission

[0050] Programmers can nest HLE regions to an implementation-specific depth of MAX_HLE_NEST_COUNT. Each logical processor tracks the nest count internally, but this count is not available to software. An HLE-legal XACQUIRE-prefixed instruction increments the nest count, and an HLE-legal XRELEASE-prefixed instruction decrements it. The logical processor enters transactional execution when the nest count increases from zero to one. The logical processor attempts a commit only when the nest count becomes zero. A transactional abort can occur if the nest count exceeds MAX_HLE_NEST_COUNT.

[0051] In addition to supporting nested HLE locks, the processor can also omit multiple nested locks. The processor keeps track of a lock omission starting with the HLE-eligible XACQUIRE-prefixed instruction for that lock and ending with the HLE-eligible XRELEASE-prefixed instruction for the same lock. The processor can keep track of locks up to a number equal to MAX_HLE_ELIDED_LOCKS at any one time. For example, if the implementation supports a MAX_HLE_ELIDED_LOCKS value of two, and if the programmer nests three HLE-identified critical locks (by executing XACQUIRE-prefixed HLE instructions on three different locks without executing an intermediate XRELEASE-prefixed HLE instruction on any of the locks), the first two locks are omitted, but the third is not omitted (instead, it is added to the transaction's write set).However, execution still continues transactionally. If an XRELEASE is encountered for either of the two dropped locks, a subsequent lock acquired by the HLE-allowed XACQUIRE-prefixed statement will be dropped.

[0052] The processor attempts to commit HLE execution when all omitted XACQUIRE and XRELEASE pairs have been reconciled, the interleave count is zero, and the lock requests are satisfied. If execution cannot be committed atomically, execution proceeds to non-transactional, no-omission execution, as if the first instruction had no XACQUIRE prefix. RTM nesting

[0053] Programmers can nest RTM regions to an implementation-specific depth of MAX_RTM_NEST_COUNT. The logical processor tracks the nest count internally, but this count is not available to software. An XBEGIN instruction increments the nest count, and an XEND instruction increments the nest count. The logical processor attempts a commit only when the nest count becomes zero. A transaction-bound abort occurs if the nest count exceeds MAX_RTM_NEST_COUNT. Nesting of HLE and RTM

[0054] HLE and RTM provide two alternative software interfaces for a general transactional execution facility. Transactional processing behavior is implementation-specific when HLE and RTM are nested together, e.g., HLE is inside RTM or RTM is inside HLE. In all cases, however, the implementation preserves both HLE and RTM semantics. An implementation may ignore HLE hints when used in RTM regions and may cause a transactional abort when RTM instructions are used in HLE regions. In the latter case, the transition from transactional to non-transactional execution occurs seamlessly because the processor reexecutes the HLE region without actually performing an omission and then executes the RTM instructions. Abort status definition

[0055] RTM uses the EAX register to communicate the abort status to the software. After an RTM abort, the EAX register has the following definition. TABLE 1 RTM abort status definition EAX register bit position Meaning 0 Set if the abort was caused by XABORT statement 1 If set, the transaction can succeed on retry, this bit is always cleared when bit 0 is set 2 Set when another logical processor conflicts with a memory address that was part of the aborted transaction 3 Set when an internal buffer has overflowed 4 Set when a debug breakpoint is hit 5 Set if an abort occurred during the execution of a nested transaction 23:6 Reserved 31-24 XABORT argument (only valid if bit 0 is set, otherwise reserved)

[0056] The EAX abort status for RTM only indicates abort causes. It does not encode whether an abort or a commit occurred for the RTM region. The EAX value may be 0 after an RTM abort. For example, if a CPUID instruction is used in an RTM region, it causes a transactional abort and may not meet the requirements for setting any of the EAX bits. This may result in an EAX value of 0. RTM memory arrangement

[0057] A successful RTM commit causes all memory operations within the RTM region to appear as if they were executed atomically. A successfully committed RTM region, which even without any memory operations within the RTM region consists of an XBEGIN followed by an XEND, has the same ordering semantics as a LOCK-prefixed instruction.

[0058] The XBEGIN instruction has no shielding semantics. However, if an RTM execution is aborted, all memory updates within the RTM scope are discarded and become invisible to any other logical processor. RTM-capable debugger support

[0059] By default, any debug exception in an RTM region causes a transactional abort and redirects control to the reset instruction address with the architectural state restored and bit 4 set in EAX. However, to allow software debuggers to intercept execution for debug exceptions, the RTM architecture provides additional features.

[0060] If bit 11 of DR7 and bit 15 of IA32_DEBUGCTL_MSR are both 1, any RTM abort due to a debug exception (#DB) or a breakpoint exception (#BP) causes execution to roll back and restart from the XBEGIN instruction instead of the reset address. In this scenario, the EAX register is also restored to the point of the XBEGIN instruction. Programming considerations

[0061] Typical areas identified by programmers are assumed to be transactional and commit successfully. However, Intel TSX offers no such guarantee. Transactional execution can abort for many reasons. To fully utilize transactional features, programmers should follow certain guidelines to increase the likelihood that their transactional execution will commit successfully.

[0062] This section discusses various events that can cause transactional aborts. The architecture ensures that updates performed in a transaction that subsequently aborts execution are never visible. Only committed transactional executions initiate an update to the architecture state. Transactional aborts never cause functional failures; they only affect performance. Considerations based on instructions

[0063] Programmers can safely use any statement in a transaction (HLE or RTM) and can use transactions at any privilege level. However, some statements will always abort transactional execution and cause execution to seamlessly and safely transition to a non-transactional path.

[0064] Intel TSX allows most common instructions to be used in transactions without causing aborts. The following operations in a transaction typically do not cause aborts: • Operations on the instruction pointer register, the general-purpose registers (GPRs) and the status flags (CF, OF, SF, PF and ZF); and • Operations on XMM and YMM registers and the MXCSR register.

[0065] However, programmers must be careful when combining SSE and AVX operations in a transactional area. Combining SSE instructions that access XMM registers and AVX instructions that access YMM registers can cause transactions to abort. Programmers can use REP / REPNE-prefixed string operations in transactions. However, long strings can cause aborts. Furthermore, the use of CLD and STD instructions can cause aborts if they change the value of the DF flag. However, if DF is 1, the STD instruction does not cause an abort. Likewise, if DF is 0, the CLD instruction does not cause an abort.

[0066] Instructions not listed here as causing an abort when used in a transaction typically do not cause a transaction to abort (examples include, but are not limited to, MFENCE, LFENCE, SFENCE, RDTSC, RDTSCP, etc.).

[0067] The following statements abort a transactional execution on any implementation: • XABORT • CPU ID • BREAK

[0068] Additionally, in some implementations, the following instructions may always cause transactional aborts. These instructions are not generally expected to be used in typical transactional areas. However, programmers should not rely on these instructions forcing a transactional abort, as whether they cause transactional aborts depends on the implementation. • Operations on X87 and MMX architecture states. This includes all MMX and X87 instructions, including the FXRSTOR and FXSAVE instructions. • Update to non-status section of EFLAGS: CLI, STI, POPFD, POPFQ, CLTS. • Instructions that update segment registers, debug registers, and / or control registers: MOV to DS / ES / FS / GS / SS, POP DS / ES / FS / GS / SS, LDS, LES, LFS, LGS, LSS, SWAPGS, WRFSBASE, WRGSBASE, LGDT, SGDT, LIDT, SIDT, LLDT, SLDT, LTR, STR, Far CALL, Far JMP, Far RET, IRET, MOV to DRx, MOV to CR0 / CR2 / CR3 / CR4 / CR8 and LMSW. • Ring transitions: SYSENTER, SYSCALL, SYSEXIT and SYSRET. • TLB and cacheability control: CLFLUSH, INVD, WBINVD, INVLPG, INVPCID and memory instructions with a non-temporal hint (MOVNTDQA, MOVNTDQ, MOVNTI, MOVNTPD, MOVNTPS and MOVNTQ). • Processor state save operation: XSAVE, XSAVEOPT and XRSTOR. • Interrupts: INTn, INTO. • IO: IN, INS, REP INS, OUT, OUTS, REP OUTS and their variants. • VMX: VMPTRLD, VMPTRST, VMCLEAR, VMREAD, VMWRITE, VMCALL, VMLAUNCH, VMRESUME, VMXOFF, VMXON, INVEPT, and INVVPID. • SMX: GETSEC. • UD2, RSM, RDMSR, WRMSR, HLT, MONITOR, MWAIT, XSETBV, VZEROUPPER, MASKMOVQ and V / MASKMOVDQU. Runtime considerations

[0069] In addition to per-statement considerations, runtime events can cause transactional execution to abort. These can be due to data access patterns or characteristics of a microarchitectural implementation. The following list does not represent a comprehensive discussion of all abort causes.

[0070] Any error or interruption in a transaction that must be disclosed to the software is suppressed. Transactional execution is aborted, and execution proceeds to non-transactional execution as if the error or interruption never occurred. If an exception is not masked, the unmasked exception results in a transactional abort, and the state appears as if the exception never occurred.

[0071] Synchronous exception events (#DE, #OF, #NP, #SS, #GP, #BR, #UD, #AC, #XF, #PF, #NM, #TS, #MF, #DB, #BP / INT3) that occur during transactional execution can cause execution to commit non-transactionally and require non-transactional execution. These events are suppressed as if they never occurred. Because the non-transactional code path with HLE is identical to the transactional code path, these events typically reappear when the instruction that caused the exception is executed non-transactionally again, causing the associated synchronous events to be passed accordingly in non-transactional execution. Asynchronous events (NMI, SMI, INTR, IPI, PMI, etc.)) that occur during transactional execution can cause the transactional execution to abort and transition to non-transactional execution. The asynchronous events are deferred and processed after the transactional abort is processed.

[0072] Transactions only support cacheable write-back memory operations. A transaction can always be aborted if the transaction contains operations on any other memory type. This includes instruction fetches for the UC memory type.

[0073] Memory accesses in a transactional region may require the processor to set the "Accessed" and "Dirty" flags of the referenced page table entry. The behavior of how the processor handles this is implementation-specific. Some implementations may allow updates to these flags to become externally visible even if the transactional region is subsequently aborted. Some Intel TSX implementations may abort transactional execution if these flags need to be updated. Furthermore, the page table traversal operation may generate accesses to its own transactionally written but uncommitted state. Some Intel TSX implementations may abort execution of a transactional region in such situations.Nevertheless, the architecture ensures that if the transactional area aborts, the transactionally written state is not exposed by the behavior of structures such as TLBs in the architecture.

[0074] Transactional execution of self-modifying code can also cause transactional aborts. Programmers must still follow Intel's recommended guidelines for writing self-modifying or cross-modifying code, even when using HLE and RTM. While an implementation of RTM and HLE typically provides sufficient resources for execution of common transactional regions, implementation limitations and excessive sizes for transactional regions can cause transactional execution to abort and transition to non-transactional execution. The architecture provides no guarantee regarding the amount of resources available for transactional execution and does not guarantee that transactional execution will ever succeed.

[0075] Conflicting requests to a cache line accessed in a transactional region may prevent the transaction from completing successfully. For example, if a logical processor P0 reads a line A in a transactional region and another logical processor P1 writes a line A (either inside or outside a transactional region), the logical processor P0 may abort if the logical processor P1's write operation interferes with the transactional execution capability of the processor P0.

[0076] Similarly, if P0 writes a row A in a transactional region and P1 reads or writes a row A (either inside or outside a transactional region), P0 may abort if P1's access to row A impairs P0's ability to execute transactionally. Additionally, other coherence traffic may sometimes appear as conflicting requests and cause aborts. While these false conflicts can occur, they are expectedly rare. The conflict resolution policy for determining whether P0 or P1 aborts in the aforementioned scenarios is implementation-specific. Embodiments of a general transaction execution:

[0077] According to "ARCHITECTURES FOR TRANSACTIONAL MEMORY," a dissertation submitted by Austen McDonald in June 2009 to the Department of Computer Science and the Committee on Graduate Studies at Stanford University in partial fulfillment of the requirements for the degree of Doctor of Philosophy, incorporated herein by reference in its entirety, there are essentially three mechanisms needed to implement an atomic and isolated transactional space: version control, conflict detection, and conflict management.

[0078] For a transactional code region to be considered atomic, all modifications performed by that transactional code region must be saved and kept isolated from other transactions until commit time. The system achieves this by implementing a version control policy. There are two version control paradigms: eager and lazy. An eager version control system saves newly generated transactional values in place and stores previous memory values alongside them in a so-called undo log. A lazy version control system temporarily stores new values in a so-called write buffer and only copies them into memory upon commit. In both systems, caching is used to optimize the storage of new versions.

[0079] To ensure that transactions appear to be executed atomically, conflicts must be detected and resolved. Both the eager and lazy version control systems detect conflicts by implementing either an optimistic or pessimistic conflict detection policy. An optimistic system executes transactions in parallel, checking for conflicts only when a transaction commits. A pessimistic system checks for conflicts on every load and store operation. Similar to version control, conflict detection also uses the cache, marking each row as either part of the read group, part of the write group, or both. Both systems resolve conflicts by implementing a conflict management policy. There are many conflict management policies, some more suitable for optimistic and some more suitable for pessimistic conflict detection.Some example guidelines are described below.

[0080] Since every transactional storage (TM) system requires both version control and conflict detection, these options result in four different TM designs: Eager-Pessimistic (EP), Eager-Optimistic (EO), Lazy-Pessimistic (LP), and Lazy-Optimistic (LO). Table 2 briefly describes all four different TM designs.

[0081] The Fig. 1 and Fig. 2 illustrate an example of a multi-core TM environment. Fig.Figure 1 shows many TM-capable CPUs (CPU1 114a, CPU2 114b, etc.) on a microchip 100 connected to a link circuit 122 under the management of a link controller 120a, 120b. Each CPU 114a, 114b (also known as a processor) may have a partitioned cache consisting of an instruction cache 116a, 116b for caching instructions to be executed from main memory and a TM-enabled data cache 118a, 118b for caching data (operands) from memory locations to be processed by the CPU 114a, 114b (in Fig.1, each CPU 114a, 114b and its associated caches are referred to as 112a, 112b. In one implementation, caches of multiple microchips 100 are interconnected to support cache coherence between the caches of the multiple microchips 100. In one implementation, a single cache is used rather than the partitioned cache containing both instructions and data. In implementations, the CPU caches are one caching level in a hierarchical caching structure. For example, each microchip 100 may use a shared cache 124 that is shared by all CPUs on the microchip 100. In another implementation, each microchip may have access to a shared cache 124 that is shared by all processors of all microchips 100.

[0082] Fig.Figure 2 shows the details of an exemplary transactional CPU environment 112 with a CPU 114, including additions to TM support. The transactional CPU (processor) 114 may include hardware to support register checkpoints 126 and special TM registers 128. The transactional CPU cache may have the MESI bits 130, tags 140, and data 142 of a conventional cache, but also, for example, bits R 132, which indicate that a line was read by the CPU 114 when executing a transaction, and bits W 138, which indicate that a line was overwritten by the CPU 114 when executing a transaction.

[0083] A key detail for programmers in any TM system is how non-transactional accesses interact with transactions. By design, transactional accesses are isolated from each other using the mechanisms mentioned above. However, the interaction between a regular non-transactional load operation and a transaction that contains a new value for that address still needs to be considered. Furthermore, the interaction between a non-transactional store operation and a transaction that has read that address must also be examined. These are issues of database design isolation.

[0084] A TM system is said to implement strong isolation, sometimes referred to as strong atomicity, if every non-transactional load and store operation behaves like an atomic transaction. Therefore, non-transactional load operations cannot see uncommitted data, and non-transactional stores cause atomicity violations in all transactions that have read that address. A system that does not implement this is said to implement weak isolation, sometimes referred to as weak atomicity.

[0085] Strong isolation is often more desirable than weak isolation due to the relative simplicity of conceptualizing and implementing strong isolation. Furthermore, if a programmer forgets to surround some shared memory references with transactions, causing program errors, the programmer can often detect this oversight using a simple debug interface in the case of strong isolation because the programmer sees a non-transactional area that causes atomicity violations. Also, programs written in one model may perform differently in another model. Furthermore, strong isolation is often easier to support in a hardware TM than weak isolation.Because in strong isolation the coherence protocol already manages the exchange of load and store operations between processors, transactions can detect non-transactional load and store operations and act accordingly. Implementing strong isolation in software from transactional memory (TM) requires modifying non-transactional code to include read and write barriers that potentially hinder performance. Although considerable effort has been devoted to removing many unnecessary barriers, such techniques are often complex, and performance is typically far lower than that of HTMs.

[0086] Table 2 illustrates the basic design framework of transactional storage (version control and conflict detection). Eager-Pessimistic (EP)

[0087] This first TM design, described below, is known as EAGER-Pessimistic. An EP system stores its write set "in place" (hence the name "eager") and, to support rollback, stores the old values of overwritten lines in a "revocation log." Processors use the buffer bits W 138 and R 132 to track read and write sets and to detect conflicts when monitored load requests are received. Perhaps the most notable examples of EP systems in the known literature are LogTM and UTM.

[0088] Starting a transaction in an EP system is very similar to starting a transaction in other systems: tm_begin() takes a register checkpoint and initializes any status registers. An EP system also requires initialization of the undo log, the details of which depend on the log format but often include initializing a log base pointer to a region of preallocated, thread-private memory and clearing a log boundary register.

[0089] Version control: Due to the way eager version control is designed in EP, the state transitions of MESI bits 130 (cache line indicators corresponding to the code states Modified, Exclusive, Shared, and Invalid) remain unchanged in most cases. Outside of a transaction, the state transitions of MESI bits 130 remain completely unchanged. When a line is read within a transaction, the standard coherence transitions apply (S (Shared) → S, I (Invalid) → S, or I → E (Exclusive)), possibly generating a load miss but also setting bit R 132. Likewise, when a line is overwritten, the standard transitions apply (S → M, E → I, I → M), possibly generating a miss but also setting bit W 138 (Written).When a row is first overwritten, the old version of the entire row is loaded and then written to the undo log to preserve it in case the current transaction aborts. The newly written data is then stored "in place" above the old data.

[0090] Conflict detection: Pessimistic conflict detection uses coherence messages exchanged during misses or updates to check for conflicts between transactions. When a read miss occurs within a transaction, other processors receive a load request; however, they ignore the request if they do not have the required row. If the other processors have the required row non-speculatively or have the R 132 (Read) row, they demote that row to S and, in certain cases, issue a cache-to-cache transfer if they have the row in state M or E in MESI bits 130. However, if the cache has the W 138 row, a conflict is detected between the two transactions, and one or more additional actions must be taken.

[0091] Similarly, if a transaction attempts to update a row from shared to modified (on a first write operation), the transaction issues an exclusive load request, which is also used to detect conflicts. If a receiving cache has the row non-speculative, the row is invalidated, and in certain cases, a cache-to-cache transfer (M or E states) is issued. However, if the row is R 132 or W 138, a conflict is detected.

[0092] Validation: Because conflict detection is performed on every load operation, a transaction always has exclusive access to its own write group. Therefore, no additional work is required for validation.

[0093] Commit: Because eager version control stores the new version of the data elements in place, the commit process simply clears bits W 138 and R 132 and discards the undo log.

[0094] Abort: When a transaction is rolled back, the original version of each cache line must be restored to the undo log, a process known as "rolling up" or "applying" the log. This occurs during a tm_discard() and must be atomic with respect to other transactions. In particular, the write set must still be used to detect conflicts: this transaction has the only correct version of the lines in its undo log, and requesting transactions must wait for the correct version to be restored from that log. Such a protocol can be applied using a hardware state machine or a software abort handler.

[0095] Eager-Pessimistic has the following properties: Committing is simple and, because it occurs in-place, very fast. Likewise, validation is a zero-operation. Pessimistic conflict detection detects conflicts early, reducing the number of doomed transactions. For example, if two transactions participate in a write-after-read dependency (counter-dependency), this dependency will be detected immediately in pessimistic conflict detection. However, in optimistic conflict detection, such conflicts are not detected until the writer commits.

[0096] Eager-Pessimistic also has the following properties: As described above, the old value must be written to the log the first time a cache line is overwritten, incurring additional cache accesses. Aborts are expensive because they require reversing the log. A load operation must be issued for each cache line in the log, possibly all the way to main memory, before proceeding to the next line. Pessimistic conflict detection also prevents the emergence of certain serializable plans.

[0097] Furthermore, since conflicts are handled in the order in which they occur, there is a possibility of livelock, and careful conflict management mechanisms must be used to guarantee forwarding progress. Lazy-Optimistic (LO)

[0098] Another popular TM interpretation is Lazy-Optimistic (LO), which stores its write group in a "write buffer" or "recovery log" and detects conflicts at commit time (still using bits R 132 and W 138).

[0099] Version control: Just as in the EP system, the MESI protocol of the LO design is enforced outside of transactions. Within a transaction, reading a row results in standard MESI transitions but also sets bit R 132. Similarly, writing a row sets bit W 138 of the row, but the handling of MESI transitions in the LO design differs from that in the EP design. First, with lazy version control, new versions of the written data are first stored in the cache hierarchy until committed, while other transactions have access to old versions available in memory or other caches. To make the old versions available, preliminary rows (M rows) must be removed when they are first overwritten by a transaction.Second, because of the optimistic conflict detection feature, no update misses are necessary: if a transaction has a row in the S state, it can simply overwrite it and update the row to an M state without communicating the changes to other transactions, because conflict detection occurs at commit time.

[0100] Conflict detection and validation: To validate a transaction and detect conflicts, LO only discloses the addresses of speculatively modified lines to other transactions when preparing to commit. During validation, the processor sends a potentially large network packet containing all addresses in the write group. Data is not sent but remains in the commit cache and is marked tentative (M). To create this packet without scanning the cache for lines marked W, a simple bit vector called a "store buffer" with one bit per cache line is used to track these speculatively modified lines. Other transactions use this address packet to detect conflicts: if an address is found in the cache and bits R 132 and / or W 138 are set, a conflict is initiated.If the row is found but neither R 132 nor W 138 are set, the row is simply invalidated, which is similar to processing an exclusive load operation.

[0101] To support transaction atomicity, these address packets must be processed atomically, meaning no two address packets with the same addresses can exist at the same time. In an LO system, this can be achieved by simply taking a global commit token before sending the address packet. However, a two-phase commit scheme could be used, sending the address packet first, collecting responses, enforcing an ordering protocol (perhaps the oldest transaction first), and committing once all responses are satisfactory.

[0102] Commit: Once validation is complete, no special processing is required for committing: bits W 138 and R 132 and the memory buffer are simply cleared. The transaction's writes are already marked as tentative in the cache, and copies of these lines in other caches have been invalidated via the address packet. Other processors can then access the committed data via the regular coherence protocol.

[0103] Abort: The reset is also simple: since the write group is contained in the local buffers, these lines can be invalidated, then the W 138 and R 132 bits and the memory buffer are cleared. The memory buffer allows W lines to be found for invalidation without having to search the buffer.

[0104] Lazy-optimistic has the following properties: Aborts occur very quickly, require no additional load or store operations, and make only local changes. More serialized plans can exist than those found in EP, allowing a LO system to more aggressively speculate that transactions are independent, which can lead to higher performance. Finally, late detection of conflicts can increase the probability of forwarding progress.

[0105] Lazy-optimistic also has the following properties: Validation consumes global data exchange time proportional to the size of the write group. Doomed transactions can waste work because conflicts are not detected until commit time. Lazy-Pessimistic (LP)

[0106] Lazy-Pessimistic (LP) represents a third TM design option that lies somewhere between EP and LO: newly written lines are stored in a write buffer, but conflicts are detected on a per-access basis.

[0107] Version control: Version control is similar to, but not identical to, LO: when a row is read, bit R 132 is set; when a row is written, bit W 138 is set; and a memory buffer is used to track W rows in the cache. Furthermore, tentative (M) rows must be evicted when they are first overwritten by a transaction, just as in LO. However, because conflict detection is pessimistic, exclusive loads must be performed when a transaction-bound row is updated from I, S → M, which is not LO.

[0108] Conflict detection: LP's conflict detection works like EP's: it uses coherence messages to search for conflicts between transactions.

[0109] Validation: As in EP, pessimistic conflict detection ensures that a running transaction does not conflict with any other running transaction at any point, so validation is a null operation.

[0110] Commit: No special processing is required for committing: the bits W 138 and R 132 and the memory buffer are simply cleared as in LO.

[0111] Abort: The reset is also performed like that of LO: the write group is invalidated using the memory buffer, and the W and R bits and the memory buffer are cleared Eager-Optimistic (EO)

[0112] LP has the following properties: Like LO, aborts occur very quickly. Like EP, the use of pessimistic conflict detection reduces the number of doomed transactions. Like EP, some serializable plans are not allowed, and conflict detection must be performed for every cache miss.

[0113] The final combination of version control and conflict detection is eager-optimistic (EO). EO can be a less than optimal choice for HTM systems: because new transactional versions are written in place, other transactions have no choice but to detect conflicts when they occur (i.e., when cache misses occur). However, because EO waits until commit time to detect conflicts, such transactions become "zombies" that continue to execute, wasting resources, but are ultimately "doomed" to abort.

[0114] EO has proven useful in STMs and is implemented by Bartok-STM and McRT. A lazy version control STM must check its write buffer on every read operation to ensure it is reading the most recent value. Since the write buffer is not a hardware structure, this is expensive, hence eager version control's preference for in-place writing. Furthermore, since checking for conflicts is also expensive in an STM, optimistic conflict detection offers the advantage of performing this operation globally. Conflict management

[0115] The rollback of a transaction after the system decides to abort that transaction was described above. However, since a conflict involves two transactions, the issues of which transaction should be aborted, how to initiate that abort, and when to retry the aborted transaction still need to be explored. These issues are addressed by conflict management (CM), a key component of transactional memory. The following describes policies on how systems initiate aborts and the various existing procedures for managing which transactions should be aborted in a conflict. Conflict management policies

[0116] A conflict management (CM) policy is a mechanism that determines which conflicting transaction should be aborted and when the aborted transaction should be retried. For example, immediately retrying an aborted transaction often does not result in the best performance. Conversely, using a suspensory mechanism that delays retrying an aborted transaction may result in better performance. STMs were the first to address the issue of finding the best conflict management policies, and many of the policies briefly outlined below were originally developed for STMs.

[0117] CM policies utilize a number of metrics to make decisions, including transaction age, read and write group size, number of previous aborts, etc. The metric combinations for such decisions are endless, but specific combinations are roughly described below in order of increasing complexity.

[0118] When establishing a nomenclature, it is first important to note that there are two sides to a conflict: the attacker and the defender. The attacker is the transaction requesting access to a shared memory location. In pessimistic conflict detection, the attacker is the transaction issuing the load or exclusive load operation. In optimistic conflict detection, the attacker is the transaction attempting to validate. In both cases, the defender is the transaction receiving the attacker's request.

[0119] An aggressive CM policy always immediately replays either the attacker or the defender. In LO, "aggressive" means the attacker always wins, and thus "aggressive" is sometimes referred to as "commitment wins." Such a policy was used for the earliest LO systems. In the case of EP, "aggressive" can mean either "the defender wins" or "the attacker wins."

[0120] Restarting a conflicting transaction that immediately leads to another conflict must result in wasted work—that is, link bandwidth filled with cache misses. A Polite-CM policy uses exponential suspension (although a linear one could also be used) before restarting conflicting transactions. To prevent deadlock, a situation in which a process has not been allocated resources by the scheduler, exponential suspension greatly increases the probability of a successful transaction after n retries.

[0121] Another approach to conflict resolution is to randomly abort the attacker or the defender (a policy called "randomized"). Such a policy can be combined with a randomized suspension scheme to avoid unnecessary conflicts.

[0122] However, random selection when choosing a transaction to abort can lead to the aborting of transactions that have already completed "a lot of work," which can waste resources. To avoid such waste, the amount of work already completed for the transaction can be taken into account when determining which transaction to abort. One measure of work should be the age of a transaction. Other methods include Oldest, Bulk TM (global TM), Size Matters, Karma, and Polka. Oldest is a simple timestamping method that aborts the younger transaction in a conflict. Bulk TM uses this scheme. Size Matters works like Oldest, but instead of transaction age, it uses the number of words read / written as the priority, reverting to Oldest after a specified number of aborts.Karma works similarly, using the size of the write group as the priority. After a suspension for a specified period of time, the transaction then proceeds with the rollback. Aborted transactions retain their priorities after aborting (hence the name Karma). Polka works like Karma, but instead of a suspension for a specified period of time, it suspends for exponentially longer periods.

[0123] Since aborting wastes work, it's a logical argument that blocking an attacker until the defender completes their transaction would result in better performance. Unfortunately, such a simple scheme easily leads to deadlock.

[0124] Deadlock avoidance techniques can be used to solve this problem. Greedy uses two rules to avoid deadlocks. The first rule states that if a first transaction T1 has a lower priority than a second transaction T0, or if T1 is waiting for another transaction, T1 aborts if it conflicts with T0. The second rule states that if T1 has a higher priority than T0 and is not waiting, T0 waits until T1 commits, aborts, or starts waiting (in which case the first rule applies). Greedy provides some guarantees about time bounds for the execution of a group of transactions. An EP interpretation (LogTM) uses a CM policy similar to Greedy to achieve blocking with conservative deadlock avoidance.

[0125] Example MESI coherence rules provide four possible states in which a cache line of a multiprocessor cache system can be, M, E, S, and I, which are defined as follows: Modified (M): The cache line exists only in the current cache and has been modified; it was modified from the value in main memory. The cache must write the data back to main memory at some point in the future before any further read operation of the (no longer valid) main memory state is permitted. The write back changes the line to the exclusive state. Exclusive (E): The cache line exists only in the current cache, but is clean; it is consistent with main memory. It can be changed to the Shared state at any time in response to a read request. Alternatively, it can be changed to the Modified state when overwritten. Shared (S): This indicates that this cache line may be stored in other caches on the machine and is "clean"; it is consistent with main memory. The line can be discarded (changed to the Invalid state) at any time. Invalid (I): This indicates that this cache line is invalid (unused).

[0126] For each cache line, TM coherence status indicators (R 132, W 138) may be provided in addition to or encoded within the MESI coherence bits. An R 132 indicator indicates that the current transaction has read from the cache line's data, and an W 138 indicator indicates that the current transaction has overwritten data in the cache line.

[0127] In another aspect of the TM design, a system is configured to use transactional memory buffers. U.S. Patent No. 6,349,361, entitled "Methods and Apparatus for Reordering and Renaming Memory References in a Multiprocessor Computer System," filed March 31, 2000, and incorporated herein by reference in its entirety, teaches a method for reordering and renaming memory references in a multiprocessor computer system having at least a first and a second processor. The first processor has a first private cache and a first buffer, and the second processor has a second private cache and a second buffer.The method includes, for each of a plurality of gateway-initiated store requests received by the first processor to store a piece of data, the steps of exclusively acquiring a cache line containing the piece of data from the first private cache and storing the piece of data in the first buffer. After the first buffer receives a load request from the first processor to load a particular piece of data, the particular piece of data is provided to the first processor from the data stored in the first buffer based on an in-order sequence of load and store operations.After the first cache receives a load request for a particular date from the second cache, an error condition is indicated and a current state of at least one of the processors is reset to a previous state if the load request for the particular date corresponds to the date stored in the first cache.

[0128] The main implementation components of such a transaction-bound memory device are a transaction-backed register file for holding pre-transaction GR (General Purpose Register) contents, a cache directory for tracking cache lines accessed during the transaction, a memory cache for buffering memory operations until the transaction completes, and firmware routines for performing various complex functions. This section describes a detailed implementation. Embodiment with zEnterprise EC12 Enterprise Server from IBM

[0129] The IBM zEnterprise EC12 Enterprise Server introduces transactional execution (TX) in transactional memory and is described in part in a paper entitled "Transactional Memory Architecture and Implementation for IBM System z," pages 25 to 36, presented at MICRO-45, December 1-5, 2012, Vancouver, British Columbia, Canada, available from IEEE Computer Society Conference Publishing Services (CPS), which is incorporated herein by reference in its entirety.

[0130] Table 3 shows an example transaction. Transactions started with TBEGIN are not guaranteed to complete successfully with TEND, as they may encounter an abort condition on each attempted execution, e.g., due to repeated contention with other CPUs. This requires the program to support a rollback path to perform the same operation non-transactionally, e.g., by using conventional locking schemes. This represents a significant overhead for the programming and software testing teams, especially in cases where the rollback path is not automatically generated by a reliable compiler. TABLE 3 Example transaction code LHI R0,0 *Initialize repetition count = 0 Ribbon TBEGIN * Start transaction JNZ Demolition * go to abort code if CC1 = 0 LT R1, lock *Load and test reset lock JNZ Ickbzy *branch if lock is occupied ... Opera ... tion execute ... TEND *Complete transaction. Ickbzy ... TABORT ......... *Cancel if lock is busy; this *will resume after TBEGIN Demolition JO Reset *no repetition if CC = 3 AHI R0, 1 *Increase the number of repetitions CIJNL R0.6, reset *give up after 6 attempts PPA R0, TX *Any delay based on the number of repetitions ... potentially wait until the lock is released... J Ribbon * Jump back to repeat reset OBTAIN block *Use Compare & Swap ... perform operation ... RELEASE block .........

[0131] The requirement to provide a rollback path for aborted transactions of a transactional execution (TX) can be cumbersome. With many transactions operating on shared data structures, they are expected to be short, contact only a few different memory locations, and use only simple instructions. For these transactions, IBM's zEnterprise EC12 introduces the concept of constrained transactions; under normal conditions, the CPU 114 ( Fig.1) ensures that constrained transactions ultimately terminate successfully, although without imposing a strict upper limit on the number of necessary retries. A constrained transaction begins with a TBEGINC statement and ends with a regular TEND. Implementing a task as a constrained or unconstrained transaction typically results in very comparable performance, but constrained transactions simplify software development by eliminating the need for a rollback path. IBM's transactional execution architecture is further described in z / Architecture, Principles of Operation, Tenth Edition, SA22-7832-09, published by IBM in September 2012, incorporated herein by reference in its entirety.

[0132] A restricted transaction begins with the TBEGINC instruction. A transaction initiated with TBEGINC must adhere to a list of programming constraints; otherwise, the program executes a non-filterable constraint violation interrupt. Example constraints may include, but are not limited to: the transaction can execute a maximum of 32 instructions; the entire instruction text can have a maximum of 256 consecutive bytes in memory; the transaction contains only forward-pointing relative branches (i.e., no loops or subroutine calls); the transaction can access a maximum of 4 aligned octowords (an octoword consists of 32 bytes) of memory; and instruction group restriction to exclude complex instructions such as decimal or floating-point operations.The constraints are chosen to allow for many common operations, such as doubly linked list insert / delete operations, including the very powerful concept of atomic compare and swap for up to four aligned octowords. At the same time, the constraints were chosen conservatively so that future CPU implementations can ensure successful transactions without having to adjust the constraints, which would otherwise lead to software incompatibility.

[0133] TBEGINC behaves largely like XBEGIN in TSX or TBEGIN in IBM's zEC12 servers, except that the floating-point register (FPR) control and interrupt filtering fields are absent, and the controls are assumed to be zero. Upon transaction abort, the instruction address is reset directly to TBEGINC rather than to the instruction after it, reflecting the immediate retry and the lack of an abort path for constrained transactions.

[0134] Nested transactions are not allowed within constrained transactions, but if a TBEGINC occurs within an unrestricted transaction, it is treated as opening a new unrestricted nesting level, just as TBEGIN would. This can occur, for example, if an unrestricted transaction calls a subroutine that internally uses a constrained transaction.

[0135] Because interrupt filtering is implicitly disabled, any exceptions during a constrained transaction result in an interrupt in the operating system (OS). The eventual successful completion of the transaction relies on the OS's ability to swap in the maximum of four pages accessed by each constrained transaction. The OS must also ensure time slices long enough to allow the transaction to complete. TABLE 4 Example transaction code TBEGINC * start restricted transaction ... perform operation.. perform... TEND *Complete transaction

[0136] Table 4 shows the constrained-transactional implementation of the code in Table 3, assuming that the constrained transactions do not interact with other lock-based code. Therefore, no lock checking is shown, but it could be added if constrained transactions and lock-based code are mixed.

[0137] If errors occur repeatedly, software emulation is performed using millicode as part of the system firmware. Advantageously, constrained transactions have desirable properties because of the reduced effort required by programmers.

[0138] With reference to Fig.3, the IBM zEnterprise EC12 processor introduced transactional execution. The processor can execute three instructions per clock cycle; simple instructions are dispatched as individual micro-operations, and more complex instructions are broken down into multiple micro-operations. The micro-operations (uops 232b) are written to a unified issue queue 216, from where they can be issued out-of-order. Each cycle can be executed by up to two fixed-point, one floating-point, two load / store, and two branch instructions. A global completion table (GCT) 232 records each micro-operation 232b and one transaction nesting depth (TND) 232a.The GCT 232 is written in order at decryption time, tracks the execution status of each micro-operation 232b, and terminates instructions when all micro-operations 232b of the oldest instruction group have been successfully executed.

[0139] Level 1 (L1) data cache 240 is a 96 kB (kilobyte) 6-way associative cache with 256 byte cache lines and 4 cycles of use latency, coupled with a 1 MB (megabyte) 8-way associative level 2 (L2) data cache 268 with 7 cycles of use latency penalty for L1 240 misses. L1 cache 240 is the cache closest to a processor, and Ln cache is a cache at the nth level of caching. Both L1 240 and L2 268 are write-through caches. Six cores on each central processor (CP) chip share a 48 MB level 3 write-back cache, and six CP chips are connected to a 384 MB off-chip level 4 cache, packaged on a glass-ceramic multichip module (MCM).Up to 4 multichip modules (MCMs) can be connected to a coherent symmetric multiprocessor (SMP) system with up to 144 cores (not all cores are available to run customer workloads).

[0140] Coherency is managed using a variant of the MESI protocol. Cache line ownership can be read-only (shared) or exclusive; L1 240 and L2 268 are write-through caches and therefore do not contain preliminary lines. L3 272 and L4 caches (not shown) are write-back caches and track preliminary states. Each cache includes all of its connected lower-level caches.

[0141] Coherence requests are called "cross-over requests" (XI) and are sent hierarchically from higher-level to lower-level caches and between L4s. When a core misses L1 240 and L2 268 and requests the cache line from its local L3 272, the L3 272 checks whether it is responsible for the line and, if so, sends an XI to the currently responsible L2 268 / L1 240 under that L3 272 to ensure coherence before returning the cache line to the requester. If the request also misses L3 272, the L3 272 sends a request to the L4 (not shown), which enforces coherence by sending XIs to all necessary L3s under that L4 and the neighboring L4s. The L4 then replies to the requesting L3, which forwards the reply to the L2 268 / L1 240.

[0142] Note that due to the inclusivity rule of the cache hierarchy, cache lines sometimes undergo cross-queries from lower-level caches due to evictions on higher-level caches caused by associativity overflows from queries to other cache lines. These XIs can be referred to as "LRU XIs," where LRU stands for "last recently used."

[0143] Referring to yet another type of XI request, Demote (Demote) XIs transfer cache ownership from the read-only to the Exclusive state, and Exclusive XIs transfer cache ownership from the Exclusive to the Invalid state. Demote XIs and Exclusive XIs require a return response to the XI sender. The target cache can "accept" the XI or send a "reject" response if it must first clean up preliminary data before accepting the XI. The L1 240 / L2 268 caches are write-through caches, but can reject Demote XIs and Exclusive XIs if they have store operations in their store queues that must be sent to the L3 before demotion from the Exclusive state. A rejected XI is retried by the sender.Read-only XIs are sent to caches responsible for the read-only row; no response is required for such XIs because they cannot be rejected. The details of the SMP protocol are similar to those described for the IBM z10 by P. Mak, C. Walters, and G. Strait in "IBM System z10 processor cache subsystem microarchitecture," IBM Journal of Research and Development, Volume 53:1, 2009, which is incorporated herein by reference in its entirety. Transaction-bound instruction execution

[0144] Fig. Figure 3 illustrates example components of an example transactional execution environment, including a CPU and caches / components with which it interacts (such as those described in Fig. 1 and Fig.2). The instruction decryption unit 208 (IDU) keeps track of the current transaction nesting depth 212 (TND). When the IDU 208 receives a TBEGIN instruction, the nesting depth 212 is increased, and conversely, decreased for TEND instructions. The nesting depth 212 is written to the GCT 232 for each dispatched instruction. When a TBEGIN or TEND is decrypted on a speculative path that is later flushed, the nesting depth 212 of the IDU 208 is updated from the most recent entry of the GCT 232 that is not flushed. The transaction-bound state is also written to the issue queue 216 for processing by the execution units, primarily the load / store unit (LSU) 280, with an effective address calculator 236 also included in the LSU 280.The TBEGIN statement can specify a transaction diagnostic block (TDB) to record status information should the transaction abort before reaching a TEND statement.

[0145] Similar to nesting depth, IDU 208 / GCT 232 jointly track access register / floating-point register (AR / FPR) modification masks across transaction nesting. IDU 208 may issue an abort request to GCT 232 when an AR / FPR modification instruction is decoded and the modification mask blocks it. If the instruction is next in line for completion, completion is blocked, and the transaction is aborted. Other restricted instructions are similarly handled, including TBEGIN, if they are decoded while in a restricted transaction or exceed the maximum nesting depth.

[0146] An outermost TBEGIN is decomposed into several micro-operations, depending on the GR memory mask; each micro-operation 232b (including, for example, uop 0, uop 1, and uop 2) is executed by one of the two fixed-point units (FXUs) 220 to save a pair of GRs 228 to a special transaction-saving register file 224, which is later used to restore the contents of the GR 228 in the event of a transaction abort. Also generated with TBEGIN are micro-operations 232b to perform an accessibility test on the TDB, if one is specified; the address is stored in a special-purpose register for later use in the event of an abort. Upon decryption of an outermost TBEGIN, the instruction address and instruction text of the TBEGIN are also stored in special-purpose registers for later potential abort processing.

[0147] TEND and NTSTG are instructions for individual microoperations 232b; NTSTG (non-transactional store operation) is processed like a normal store operation, except that it is marked as non-transactional in the issue queue 216 so that the LSU 280 can treat it accordingly. TEND is a null operation at execution time; transaction termination occurs when TEND completes.

[0148] As mentioned, instructions within a transaction are marked as such in the issue queue 216, but otherwise execute largely unchanged; the LSU 280 performs isolation tracking as described in the next section.

[0149] Because decryption occurs in order, and because the IDU 208 tracks the current transactional state and writes it to the issue queue 216 along with each instruction from the transaction, the execution of TBEGIN, TEND, and instructions before, within, and after the transaction can occur out of order. It is even possible (though unlikely) for TEND to execute first, followed by the entire transaction, and TBEGIN to execute last. Program order is restored at exit time by the GCT 232. The length of transactions is not limited by the size of the GCT 232 because the general-purpose registers (GRs) 228 can be restored from the backup register file 224.

[0150] During execution, Program Event Recorder (PER) events are filtered based on the event suppression control, and a PER TEND event is detected if enabled. Similarly, in transactional mode, a pseudorandom generator can cause random aborts, as enabled by the transaction diagnostic control. Tracking transaction-based isolation

[0151] The load / store unit 280 tracks cache lines accessed during transactional execution and triggers an abort if an XI from another CPU (or an LRU XI) conflicts with the memory requirement. If the conflicting XI is an exclusive or demote XI, the LSU 280 refers the XI back to the L3 272 in the hopes of completing the transaction before the L3 272 retries the XI. This "bumping" is very efficient in highly contentious transactions. To prevent deadlocks when two CPUs bump each other, an XI rejection counter is implemented that triggers a transaction abort when a threshold is reached.

[0152] The L1 cache directory 240 is traditionally implemented using static random access memory (SRAM). For the implementation of transactional memory, the valid bits 244 (64 rows x 6 ways) of the directory have been shifted to normal logical latches and are supplemented with two additional bits per cache line: the TX-read bits 248 and TX-dirty bits 252.

[0153] The TX-read 248 bits are reset when a new outermost TBEGIN is decrypted (which is locked against a previous, pending transaction). The TX-read 248 bit is set at execution time of each load instruction marked as "transaction-bound" in the issue queue. Note that this can lead to over-marking when speculative load operations are performed, for example, on a mispredicted branch path. Alternatively, setting the TX-read 248 bit at load completion time was too expensive for a silicon area, as multiple load operations can complete at the same time, requiring many read ports on the load queue.

[0154] Store operations are performed in the same way as in non-transactional mode, but a transaction marker is placed in the store instruction's entry in the store operation queue (STQ) 260. At writeback time, when the data is written from the STQ 260 to the L1 240, the TX-dirty 252 bit in the L1 cache directory 240 is set for the written cache line. A store writeback to the L1 240 occurs only after the store instruction completes, and a maximum of one store writeback is performed per cycle. Before completion and the writeback, load operations can access the data from the STQ 260 via store forwarding; after the writeback, the CPU 114 ( Fig.1) access the speculatively updated data in L1 240. If the transaction completes successfully, the TX-dirty 252 bits for all cache lines are cleared, and the TX flags of unwritten memory operations are also cleared in STQ 260, effectively turning the pending memory operations into normal memory operations.

[0155] When a transaction aborts, all outstanding transaction-bound memory operations from the STQ 260 are invalidated, even those that have already completed. For all cache lines modified by the transaction in the L1 240 (i.e., those whose TX-dirty 252 bit is turned on), their valid bits are turned off, effectively and immediately removing them from the L1 240.

[0156] The architecture requires that the isolation of the transaction's read and write groups be maintained before a new instruction completes. This isolation is ensured by blocking instruction completion at appropriate times when XIs are outstanding; speculative out-of-order execution is permitted, optimistically assuming that the outstanding XIs are directed to different addresses and do not, in fact, cause transaction conflict. This design fits very naturally with the XI-to-completion locks implemented on earlier systems to ensure the strong memory ordering required by the architecture.

[0157] When the L1 240 receives an XI, the L1 240 accesses the directory to check the validity of the cross-request address in the L1 240. If the TX-read 248 bit on the cross-request line is active and the XI is not rejected, the LSU 280 triggers an abort. If a cache line with the TX-read 240 bit active is the least recently used from the L1 240, a special LRU extension vector for each of the 64 lines of the L1 240 remembers that a TX-read line was present on that line. Because there is no precise address tracking for LRU extensions, a non-rejected XI that results in a valid extension line triggers an abort by the LSU 280. Providing LRU expansion effectively increases the read memory footprint from L1 size to L2 size and causes associative aborts, provided there are no conflicts with other CPUs 114 ( Fig. 1 and Fig.2) to the non-precise LRU extension tracking.

[0158] The memory footprint is limited by the memory cache size (memory footprint is discussed in more detail below) and thus implicitly by the size of L2 268 and associativity. No LRU expansion action needs to be performed if a cache line with TX-dirty 252 is the least recently used from L1 240. Memory cache

[0159] Because in previous systems, the L1 240 and L2 268 are write-through caches, each store instruction causes a memory access to the L3 272; now with 6 cores per L3 272 and further improved performance of each core, the memory rate for the L3 272 (and to a lesser extent the L2 268) becomes problematic for certain workloads. To avoid delays in store queues, a capture store cache 264 had to be added, which combines store operations with adjacent addresses before sending them to the L3 272.

[0160] For transactional memory performance, it is acceptable to invalidate any cache line with TX-dirty 252 from L1 240 upon transaction aborts, because the cache is very close to L2 268 (7 cycles, L1 240 miss penalty) to reclaim the clean lines. It would be unacceptable for performance (and tracking area) for transactional memory operations to overwrite L2 268 before the transaction ends and subsequently invalidate all L2 268 tentative cache lines upon abort (or worse, on the shared L3 272).

[0161] The twin problems of memory bandwidth and transactional memory processing can be addressed with the capture memory cache 264. The capture memory cache 264 is a 64-entry circular queue, with each entry containing 128 bytes of data with byte-exact valid bits. In a non-transactional operation, when a store operation is received from the LSU 280, the store cache 264 checks to see if an entry exists for the same address and, if so, captures the new store operation in the existing entry. If no entry exists, a new entry is written to the queue, and when the number of free entries falls below a threshold, the oldest entries are written back to the L2 and L3 caches 268 and 272.

[0162] When a new outermost transaction begins, all existing entries in the memory cache are marked as closed, allowing new memory operations to be committed to them, and the cleanup of these entries for L2 268 and L3 272 begins. From this point on, transactional memory operations from LSU 280 STQ 260 are allocated new entries or committed to existing transactional entries. Writing these memory operations back to L2 268 and L3 272 is blocked until the transaction completes successfully; at this point, subsequent (post-transaction) memory operations can continue to be committed to existing entries until the next transaction closes these entries.

[0163] The memory cache 264 is polled for each exclusive or demo XI and causes an XI rejection when the XI is compared to an active entry. If the core fails to complete any more instructions while XIs continue to be rejected, the transaction is aborted at a certain threshold to avoid deadlocks.

[0164] The LSU 280 requests a transaction abort when the memory cache 264 overflows. The LSU 280 detects this condition when it attempts to submit a new memory operation that cannot be captured in an existing entry, and the entire memory cache 264 is filled with memory operations from the current transaction. The memory cache 264 is managed as a subset of the L2 268: while transaction-bound preliminary lines can be flushed from the L1 240, they must remain resident in the L2 268 throughout the transaction. The maximum memory footprint is thus limited to the memory cache size of 64 x 128 bytes, and it is also limited by the associativity of the L2 268. Because the L2 268 is an 8-way associative memory and has 512 lines, it is typically large enough not to cause transaction aborts.

[0165] When a transaction aborts, the memory cache 264 is notified, and all entries containing transaction-bound data are invalidated. The memory cache 264 also has a flag per double word (8 bytes) indicating whether the entry was written by an NTSTG instruction—these double words remain valid across transaction aborts. Functions implemented by Millicode

[0166] Traditionally, IBM mainframe server processors contain a firmware layer known as millicode, which performs complex functions such as certain CISC instruction execution, interrupt handling, system synchronization, and remote access. Millicode includes machine-dependent instructions and instruction set architecture (ISA) instructions, which are fetched from memory and executed similarly to application program and operating system (OS) instructions. Firmware resides in a restricted area of main memory that is inaccessible to customer programs. When the hardware detects a situation where millicode needs to be invoked, the instruction fetch unit switches to "millicode mode" and begins fetching from the appropriate memory location in the millicode memory area.Millicode can be fetched and executed in the same way as instruction set architecture (ISA) instructions and can contain ISA instructions.

[0167] For transactional memory, millicode is involved in several complex situations. Each transaction abort calls a dedicated millicode subroutine to perform the necessary abort steps. The transaction abort millicode begins by reading special-purpose registers (SPRs) containing the hardware-related internal abort reason, potential exception reasons, and the address of the aborted instruction, which the millicode then uses to save a TDB, if specified. The TBEGIN instruction test is loaded from an SPR to obtain the GR memory mask, which the millicode needs to know which DRs 238 to restore.

[0168] The CPU 114 ( Fig.1) supports a special millicode-only instruction for reading the backup GRs 224 and copying them to the main GRs 228. The TBEGIN instruction address is also loaded from an SPR to establish the new instruction address in the PSW to continue execution after the TBEGIN after the millicode abort subroutine has completed. This PSW can later be saved as the PSW of the old program if the abort was caused by an unfiltered program interrupt.

[0169] The TABORT instruction can be implemented by millicode; when IDU 208 decodes TABORT, it instructs the instruction fetch unit to branch into TABORT millicode, from where millicode branches into general-purpose abort subroutines.

[0170] The Extract Transaction Nesting Depth (ETND) instruction can also be encoded in millicode because it is not performance-critical; millicode loads the current nesting path from a special hardware register and places it in a GR 228. The PPA instruction is encoded in millicode; it performs the optimal delay based on the current abort count, which is provided by software as an operand to PPA, and also based on other hardware internal conditions.

[0171] For constrained transactions, Millicode can track the number of aborts. The counter is reset to 0 upon successful TEND completion or in the event of an interrupt in the OS (since it is unknown if or when the OS will return to the program). Depending on the current abort count, Millicode can invoke certain mechanisms to improve the chance of success for the subsequent transaction retry. Examples of these mechanisms include successively increasing random delays between retries and decreasing the amount of speculative execution to avoid the occurrence of aborts caused by speculative accesses to data that the transaction is not currently using. As a last resort, Millicode can outsource to other CPUs 114 ( Fig.1) Broadcast, stopping any conflicting work, and retrying the local transaction before releasing the other CPUs 114 to continue normal processing. Multiple CPUs 114 must be coordinated to avoid causing deadlocks, so some serialization between Millicode instances on different CPUs 114 is required.

[0172] One technique for making groups of memory instructions atomic is the use of transactional memory (TM). The most difficult challenge for a hardware TM is handling write operations. A TM system must track memory operations and assemble a write group (WS) during the transaction. The actual WS data must be buffered until the transaction completes successfully. In case of success, all memory operations become atomically visible globally in the WS (i.e., all at once), typically by writing the buffered data to a cache. Alternatively, if the transaction aborts, the buffered memory operations are discarded without modifying memory.

[0173] Load operations are somewhat simpler because they do not modify memory. The TM must track transactional load operations, creating a read group (RS) containing the data and access address of each transactional load operation. A successful transaction transfers the data from the load operations in the RS to the register file.

[0174] A primary function performed by transaction execution is to detect interference from remote requests, i.e., memory accesses that do not correspond to the current application thread.

[0175] Often, a TM monitors for contention with the read group and write group at a cache line address granularity (e.g., a 64B cache line) during a transaction. That is, RS and WS contain the addresses of the cache lines where transaction-bound data is located, not the addresses of the data in a cache line. An RS contention occurs when any data in a cache line in the read group is overwritten by another thread. A WS contention occurs when any data in a cache line in the WS is read or overwritten by another thread. When a contention occurs, the transaction aborts, undoing (i.e., nullifying) the effects of any load and store operations that participated in the transaction and affected a state of the computer.To achieve this, the TM allows load and store operations participating in a transaction to change the architectural state of a computer only after the transaction has completed successfully.

[0176] To improve performance, a computer architecture often includes prefetch instructions that can move data from main memory or a lower cache level to a higher cache level (closer to the processor) in anticipation of an access to the data. The execution of a prefetch instruction can be speculative in nature in that it can be executed before it is known that the data moved by the prefetch instruction will actually be accessed. Data that is expected to be either read or written can be prefetched, referred to as read prefetch or write prefetch, respectively. Cache coherence protocols in computers that enable data sharing, synchronization, and parallel processing often require a processor to gain ownership of the data before it is overwritten (i.e.changed). Because data may currently be under the responsibility of another processor, acquiring responsibility can be a lengthy process. The concept of ownership is necessary to prevent data read from one memory location by one processor from being changed by another processor writing to the same memory location or a different memory location (when copies of the data are located in different memory locations). Therefore, it is often advantageous for a processor to execute a write prefetch instruction to acquire ownership of the data in anticipation of the data being overwritten, so that processing is not delayed once ownership is acquired.

[0177] According to the present invention, embodiments avoid performance degradations traditionally associated with the execution of prefetch instructions in existing transactional memory systems, where prefetch instructions (i.e., prefetch requests) cause transactions to fail. In one embodiment of the present invention, the execution of prefetch instructions in a transactional memory system does not require a memory transaction in a local processor to abort due to an address conflict with a prefetch request from a remote processor. In one embodiment of the present invention, a prefetch instruction (i.e., a prefetch request) may cause a memory transaction (i.e., a transaction) to abort due to an address conflict under certain conditions (e.g.,if a priority or importance of the prefetch is higher than that of the memory transaction). In one embodiment, a prefetch request that conflicts with a transaction is evaluated, the evaluation including one or more of i) comparing a priority of the prefetch request with a priority of the transaction, ii) comparing a priority of the local processor with a priority of the remote processor, and iii) determining an importance of the prefetch request. In one embodiment, an action is taken based on the evaluation, the action including one or more of i) aborting the prefetch request, ii) quiescing the prefetch request (i.e.the pre-access request is placed in a dormant but alert state while waiting for a condition to occur before taking action), iii) delaying the pre-access request for a delay period (i.e., the pre-access request is placed in a dormant state until the delay period is over before taking action), and iv) executing the pre-access request. In one embodiment, a duration of a particular delay is determined by one or both of i) one or more conditions in the local processor and ii) one or more conditions in the remote processor. For example, in one embodiment, a delay period, i.e., a duration of a particular delay, depends on conditions in the local processor (e.g.,CPU utilization, cache utilization, and number of queued processes, as well as current instructions executed per cycle). In another example, in one embodiment, a delay period depends on conditions in a remote processor that generated the prefetch request (e.g., CPU utilization, cache utilization, and number of queued processes, as well as current instructions executed per cycle).

[0178] In one embodiment, a prefetch request is suppressed if the prefetch request conflicts with a transaction, where the meaning of the term "suppressed" includes any temporal sequence and combination of: i) not immediately processing the prefetch request (i.e., delaying the prefetch request), ii) queuing the prefetch request, and iii) aborting the prefetch request. Each of these operations constitutes prefetch request suppression in the present invention.

[0179] In one embodiment, a prefetch request operation (i.e., a prefetch operation) is a memory access performed during the execution of a prefetch request. Any performance increase due to the execution of a prefetch request is likely to be significantly less than any performance degradation due to a memory transaction abort.In one embodiment, a prefetch request participates in a prefetch protocol (which is part of the memory and cache coherence protocols described herein) that identifies a read or write memory access as a prefetch request and provides information about the prefetch request such that the prefetch request can be managed without causing a transaction abort if the memory operation associated with the prefetch request conflicts with the transaction. In one embodiment, a prefetch protocol includes a read prefetch request and a write prefetch request. In one embodiment, a prefetch request includes one or more prefetch operations.In one embodiment, the pre-access protocol includes one or more pre-access acknowledgments sent by a recipient of a pre-access request to the source of the pre-access request (i.e., the source of a pre-access request is notified of the status of the pre-access request). In one embodiment, a pre-access acknowledgment includes information about the type (i.e., status) of the pre-access request at the recipient of the pre-access request (e.g., a local processor).

[0180] In one embodiment, the prefetch request is discarded at the receiving node if it is determined that processing the prefetch request would cause the memory transaction to abort. In one embodiment, if the source of the prefetch request (e.g., a remote processor) identifies a prefetch request as particularly important, the prefetch request is queued at the receiving node and executed when it no longer causes the memory transaction to abort. In one embodiment, there is a separate queue for each remote processor, and a prefetch request from a remote processor is queued for that remote processor.In one embodiment, there is a queue for pre-accession requests for exclusive jurisdiction data and a queue for pre-accession requests for shared jurisdiction data. In one embodiment, when the pre-accession request includes a plurality of pre-accession request operations (i.e., pre-accession operations), each pre-accession operation is queued separately. In one embodiment, a priority is associated with a pre-accession request (e.g., the priority of the pre-accession source). In one embodiment, a pre-accession operation is serviced in a queue at the receiver of the pre-accession request according to the priority of the pre-accession request. In one embodiment, a priority is associated with the source of the pre-accession request and the receiver of a pre-accession request. In one embodiment, when the source of a pre-accession request (e.g.,If the prefetch request (e.g., a remote processor) has a higher priority than the priority of the prefetch recipient (e.g., a local processor), the prefetch request is served and may cause a transaction abort in the recipient if it conflicts with the transaction. In one embodiment, there is a separate queue for each prefetch request priority, and a prefetch with a priority is placed in the queue for that priority.

[0181] In one embodiment, a pre-access protocol includes a pre-access time horizon associated with each pre-access request, measured in cycles that the recipient of a pre-access request uses to determine how to manage a pre-access request. In one embodiment, a pre-access request is discarded after a pre-access time horizon associated with the pre-access request has expired if it is in a queue and has not yet been serviced. In one embodiment, when the pre-access request is discarded by the recipient of the pre-access request, a notification is sent to the source of the pre-access request that the pre-access request has been discarded. In some embodiments, the time horizon is specified in other metrics, e.g.,Number of instructions executed, number of memory instructions executed, or other metric associated with a program and / or processor.

[0182] In one embodiment, a pre-accession protocol includes a specification of a pre-accession time horizon associated with a pre-accession operation divided into time steps from the near future to the far future, where each time step is specified in a preset number of cycles that the recipient of a pre-accession request uses to manage a pre-accession request. In one embodiment, execution of a pre-accession request, passed to a recipient of the pre-accession request by the pre-accession protocol in delay information, is delayed by a number of time steps. In one embodiment, the recipient of the pre-accession request delays the pre-accession request by the specified number of time steps (i.e., a delay period or, equivalently, a delay duration), after which the pre-accession request is attempted to be executed.In one embodiment, if the prefetch request encounters a conflict with a memory transaction when attempting to execute the prefetch request after the delay has expired, the prefetch request is aborted. In one embodiment, if the prefetch request encounters a conflict with a memory transaction when attempting to execute the prefetch request after the delay has expired, the prefetch request is queued.

[0183] In one embodiment, if a prefetch request conflicts with a memory transaction and the number of time steps specified in the prefetch protocol is less than a threshold, the prefetch request is discarded; otherwise, the prefetch request is queued. In one embodiment, a prefetch request in a prefetch protocol is referred to as a short-term prefetch request or a long-term prefetch request. In one embodiment, a short-term prefetch request is active (i.e., the prefetch request can potentially be executed) for a specified number of cycles.When a short-term prefetch request is active, the short-term prefetch request is queued if the short-term prefetch request conflicts with a memory transaction, and aborted if the short-term prefetch request is no longer active. In one embodiment, a long-term prefetch request remains active until it has been executed. In one embodiment, a long-term prefetch request is queued if it conflicts with a memory transaction, and dequeued and executed if that long-term prefetch request no longer conflicts with a memory transaction.In one embodiment, a long-term pre-access request is removed from the queue and discarded when the queue is full and that long-term pre-access request is the oldest pre-access request in the queue. In one embodiment, if there is at least one entry in the queue that is not a short-term pre-access request and a particular long-term pre-access request in the queue is the oldest pre-access request in the queue, the oldest long-term pre-access request is not removed from the queue and is discarded when the queue is full. In such an embodiment and scenario, if the queue becomes full and at least one entry is not a long-term pre-access request, at least one entry in the queue that is not a long-term pre-access request is removed from the queue.In one embodiment, a long-term preempt request is not dequeued and is discarded when the queue becomes full.

[0184] Fig.4 illustrates a processor system 1101, which in one embodiment includes a processor A 1102 coupled to a cache system A 1103 and a processor B 1106 coupled to a cache system B 1107. Cache system A 1103 and cache system B 1107 are coupled to a system bus 1111 and are substantially similar. Cache system A 1103 includes a transactional memory system A 1104 and cache A 1105. Cache system B 1107 includes a transactional memory system B 1108 and cache B 1109. System memory 1110 is coupled to system bus 1110.In one embodiment, cache system A 1103 and cache system B 1107 participate in a cache coherency protocol that allows processor A 1102 and processor B to execute instructions that operate on data shared by both processors in an organized manner. For example, processor B 1106 may execute a read instruction that accesses data from cache system B 1107, and if cache system B 1107 does not have the data, cache system B 1107 asserts a cache coherency protocol request on system bus 1111. Cache system A 1103 monitors (via snooping) for cache coherency protocol requests on system bus 1111, and if cache system A 1103 has the requested data, the data is acknowledged on system bus 1111.Caching system A 1103 also compares the address of the requested data to determine if the address conflicts with an active transactional memory operation. In one embodiment, transactional memory system A 1104 in caching system A 1103 determines whether a memory address conflicts with that of an active transactional memory operation.

[0185] Fig.Figure 5 shows the transactional memory system A 1104 in greater detail. In one embodiment, the transactional memory system A 1104 includes write groups 2205, read groups 2206, an abort manager 227, and a prefetch manager 2208. The write groups 2205 contain the target addresses of the write operations (store operations) participating in a transaction. The read groups 2206 contain the target addresses of the read operations (load operations) participating in a transaction. In one embodiment, the read groups 2206 and the write groups 2205 each contain the read and write addresses of one or more memory reads and writes in a transaction.

[0186] According to the invention, inputs to the transactional memory system A 1104 include non-prefetch operations (i.e., non-prefetch memory accesses) on a non-prefetch operations bus 2210 and prefetch operations (i.e., prefetch request memory accesses) on a prefetch operations bus 2209. An address decryptor 2211 decrypts the address of a non-prefetch operation and asserts the target address of a non-prefetch read operation on a read address bus 2201 and the target address of a non-prefetch write operation on a write address bus 2202. An address decryptor 2212 decrypts the address of a prefetch operation and asserts the target address of a prefetch read operation on a prefetch read address bus 2203 and the target address of a prefetch write operation on a prefetch write address bus 2204.The prefetch manager 2208 is connected to the abort manager 2207 and receives all prefetch operations input to the transactional memory system A 1104 from the prefetch bus 2209.

[0187] Read address bus 2201 and write address bus 2202 are inputs to abort manager 2207. In one embodiment, if abort manager 2207 detects that an address on read address bus 2201 matches an address in write group 2205, abort manager 2207 aborts the transaction-bound memory operation that includes write operations for the address. In one embodiment, if abort manager 2207 detects that an address on write address bus 2202 matches an address in write group 2205 or matches an address in read group 2206, abort manager 2207 aborts the transaction.

[0188] Prefetch read address bus 2203 and prefetch write address bus 2204 are inputs to abort manager 2207. In one embodiment, when abort manager 2207 detects that an address on prefetch read address bus 2203 matches an address in write group 2205, abort manager 2207 notifies prefetch manager 2208 that a conflict has occurred with an address of a current prefetch operation. In one embodiment, abort manager 2207 does not abort the transaction containing a write operation to the address. In one embodiment, abort manager 2207 aborts the transaction containing a write operation to the address when instructed to abort by prefetch manager 2208.In one embodiment, if the abort manager 2207 detects that an address on the prefetch write address bus 2204 matches an address in the write group 2205 or matches an address in the read group 2206, the abort manager 2207 notifies the prefetch manager 2208 that a conflict has occurred with an address of a current prefetch operation. In one embodiment, the abort manager 2207 does not abort the transaction. In one embodiment, the abort manager 2207 aborts the transaction when instructed to do so by the prefetch manager 2208.

[0189] In one embodiment, prefetch manager 2208 includes a prefetch queue 2213, a queue manager 2214, and a local / remote prefetch 2216. In one embodiment, local / remote prefetch 2216 classifies a prefetch operation as either a local prefetch operation (e.g., from a local processor) or a remote prefetch operation (e.g., from a remote processor). A prefetch operation is either a local prefetch operation or a remote prefetch operation with respect to a memory transaction. A prefetch operation is local to a memory transaction if the prefetch operation is generated on the same processor (e.g., the local processor) that generates the memory transaction and accesses a cache accessed by the memory transaction; otherwise, the prefetch operation is a remote prefetch operation (e.g.,from a remote processor) with respect to the memory transaction. Therefore, a prefetch operation is local to all transactional memory operations created by the execution thread that created the prefetch operation. In one embodiment, a prefetch operation that local / remote prefetch 2216 classifies as local relative to a transaction may not conflict with read or write operations in write groups 2205 or read groups 2206 associated with the transaction. In particular, a prefetch is local if it is associated with the present transactional memory system (e.g., if multiple threads in a single processor, e.g., processor 1102, or multiple processors are associated with the present transactional memory system).

[0190] In one embodiment, if prefetch manager 2208 is notified that a current prefetch operation conflicts with a transaction, prefetch manager 2208 discards the prefetch operation and the prefetch request associated with the prefetch operation. In one embodiment, if prefetch manager 2208 discards a prefetch operation, prefetch manager 2208 sends a response to the prefetch source notifying the source that the prefetch has been discarded. In one embodiment, if prefetch manager 2208 is notified by abort manager 2207 that a current prefetch operation conflicts with an existing transaction, queue manager 2214 places the prefetch operation in prefetch queue 2213.In one embodiment, prefetch manager 2208 processes the queued prefetch operation after the conflicting transactional memory operation has either completed or aborted. In one embodiment, when queue manager 2214 places a prefetch operation in prefetch queue 2213, prefetch manager 2208 sends a response to the prefetch source notifying the source that the prefetch operation has been queued.

[0191] In one embodiment, prefetch operations on the prefetch bus 2209 participate in a prefetch protocol that includes a prefetch request. A prefetch request describes a memory access to data that is a prefetch memory access, rather than a regular read or write memory access for which there is a secure need for data. In one embodiment, a prefetch request is identified with a unique tag. In one embodiment, the unique tag cannot be reused in a system until a duplicate of the unique tag cannot possibly continue to exist in the system.In one embodiment, the prefetch request includes prefetch information (included in the prefetch request at the source of the prefetch request) that enables the prefetch operation (which is the result of the prefetch request) to be efficiently managed by the prefetch manager 2208 when processing the prefetch operation. In one embodiment, the prefetch information includes priority information about the source of the prefetch operation. In one embodiment, if the prefetch manager finds that the priority of the source of the prefetch operation exceeds the priority of the process that created a transactional memory operation (which is in address conflict with the prefetch operation), the prefetch manager 2208 allows the prefetch operation to proceed, and the transaction is aborted by the abort manager 2207.In one embodiment, the prefetch manager finds that the ratio of the priority of the source of the prefetch operation to the priority of the process that created a transaction (which is in address conflict with the prefetch operation) exceeds a threshold, the prefetch operation continues, and the transaction is aborted. In one embodiment, if a prefetch operation causes a memory transaction to be aborted, the prefetch manager 2208 sends a notification to the source of the prefetch operation indicating that the prefetch operation has continued with a subsequent transaction abort. Furthermore, a transaction abort is communicated to the processor corresponding to the transaction to roll back the transaction state and restart execution of the transaction or an alternate path.

[0192] In one embodiment, prefetch information in a prefetch request includes a prefetch horizon. In one embodiment, the prefetch horizon is specified as a number of cycles for which queue manager 2214 may place the prefetch operation in prefetch queue 2213 if it is in address conflict with a transactional memory operation. In one embodiment, the prefetch horizon is specified at the source of the prefetch as a number of instruction executions, which is converted to a number of cycles at the receiver of the prefetch. In one embodiment, a conversion of the number of instruction executions at the source to a number of cycles is performed at the receiver using a static conversion factor that does not change during execution.In one embodiment, a conversion of the number of instruction executions at the source to a number of cycles at the receiver is performed during execution using a dynamically calculated conversion factor (e.g., by multiplying the number of instructions by a current cycles-per-instruction conversion factor) calculated by the prefetch receiver.

[0193] In one embodiment, a local processor receiving prefetch requests keeps track of the number of cycles expected to remain in the one or more current transactions with which the prefetch conflicts. In one embodiment, a count of the expected number of cycles remaining in one or more transactions is maintained. In one embodiment, if the end time of a time horizon for a prefetch is before the end time after an expected number of cycles remaining in a conflicting transaction, the prefetch request is discarded; otherwise, the prefetch request is queued.

[0194] In one embodiment, if the prefetch operation has not been processed within the number of cycles of the prefetch horizon, queue manager 2214 removes the prefetch operation from prefetch queue 2213 and discards it. In one embodiment, if the prefetch operation is discarded after the number of cycles of the prefetch horizon, prefetch manager 2208 notifies the prefetch source that the prefetch was discarded because the prefetch horizon was exceeded.

[0195] In one embodiment, the prefetch information in a prefetch request includes the importance of the source of the prefetch request. In one embodiment, if the importance of the source exceeds a threshold, and if the prefetch request has an address conflict with a transactional memory operation, the prefetch manager 2208 queues the prefetch operation; otherwise, the prefetch request is discarded.

[0196] In one embodiment, if the prefetch queue 2213 is full, an incoming prefetch request is discarded. In one embodiment, if an incoming prefetch request is discarded because the prefetch queue 2213 is full, a response to the prefetch request is sent to the source of the prefetch request informing it that the prefetch request was discarded because the prefetch queue 2213 was full. In one embodiment, if the prefetch queue 2213 is full, an entry is selected from the prefetch queue 2213 and discarded. In one embodiment, the queue manager 2214 selects a prefetch in the prefetch queue 2213 to discard.In one embodiment, when queue manager 2214 discards a prefetch from prefetch queue 2213 to make room for another prefetch request, the source of the discarded prefetch is notified that the prefetch request has been discarded to make room for another prefetch request. In one embodiment, when a prefetch request causes an eviction from prefetch queue 2213, the source of the prefetch request is notified that the prefetch request has caused an eviction from prefetch queue 2213. In one embodiment, when an eviction from the prefetch queue is performed, the oldest entry in the prefetch queue is discarded.In another embodiment, the pre-access request discarded from the pre-access queue 2213 is the pre-access request at the head of the pre-access queue 2213. In another embodiment, the pre-access request in the pre-access queue 2213 that is discarded is the pre-access request with the shortest time horizon, that is, the pre-access request with the shortest end time. In another embodiment, the pre-access request selected from the queue for discarding is the pre-access request with the longest time horizon, that is, the pre-access request with an end time farthest in the future.

[0197] In other embodiments, different buses are Fig.5. Thus, in one or more of such embodiments, as non-limiting examples, the function of read address bus 2201 and prefetch read address bus 2203 is performed by a single read bus with a request type indicator; the function of write address bus 2202 and prefetch write address bus 2204 is performed by a single write bus with a request type indicator; the function of prefetch buses 2203 and 2204 is performed by a single prefetch bus with a request type indicator; the functions of non-prefetch buses 2201 and 2203 are performed by a single non-prefetch bus with a request type indicator; the functions of all buses are performed by a single address bus with a request type indicator. In other embodiments, request and address buses are combined.

[0198] In some scenarios and embodiments, a prefetch protocol includes a read request that is acknowledged on a bus-based or switch-based connection that includes information regarding the purpose of a prefetch request. In one embodiment, in a system that includes a MESI cache coherence protocol, the cache coherence protocol is extended to include a read request that includes a specification regarding whether the read request is a write / read request (i.e., a non-prefetch request of data) or a prefetch request. In one embodiment, a prefetch request is a prefetch for an exclusive responsibility for data (e.g., to read or write data) or a prefetch for a shared responsibility for data (e.g., to read the data).

[0199] Fig.Figure 6 illustrates one embodiment of a format of a read request indicating a read / write request or a pre-access request, and further indicating whether a pre-access request is for exclusive data ownership or a pre-access request is for shared data ownership. In one embodiment, a read request 3311 includes at least six information fields: read information field 3301, tag information field 3302, exclusive or shared information field 3303, p (pre-access) information field 3304, address information field 3305, and CRC (cyclic redundancy check) information field 3306.In one embodiment, the read information field 3301 identifies the read request 3311 as a read request; the tag information field 3302 identifies a particular instance of the read request 3311; the exclusive or shared information field 3303 indicates whether the read request 3311 is for exclusive or shared access to data; the p information field 3304 indicates whether the read request 3311 is a pre-access request; the address information field 3305 indicates the address of the data requested by the read request 3311; the CRC information field 3306 contains a value calculated from the bit patterns of the values in the other information fields in the read request 3311 and is used to detect an error that occurred during a transfer of the read request 3311 from a source of the read request 3311 to a destination of the read request 3311.

[0200] Fig.6 also illustrates one embodiment of a format of a response to read request 3311, a read response 3312. In one embodiment, read response 3312 includes at least four information fields: read response information field 3307, tag information field 3308, data information field 3309, and CRC information field 3310. In one embodiment, read response information field 3307 identifies read response 3312 as a read response; tag information field 3308 identifies read response 3312 as a response to read request 3311, where tag information field 3302 is the same as tag information field 3308; data information field 3309 contains data corresponding to the data requested by read request 3311, where tag information field 3302 is the same as tag information field 3308.The CRC information field 3310 contains a value calculated from the bit patterns of the values in the other information fields in the read response 3312 and is used to detect an error that occurred during a transfer of the read response 3312 from a source of the read response 3312 to a destination of the read response 3312.

[0201] In some scenarios and embodiments, a prefetch protocol includes a prefetch request acknowledged on a bus-based or switch-based connection that includes information regarding the purpose of a prefetch request. In one embodiment, in a system including a MESI cache coherence protocol, the cache coherence protocol is extended to include a prefetch request that includes a specification regarding whether the prefetch request is a prefetch for an exclusive data jurisdiction or a prefetch for a shared data jurisdiction.

[0202] Fig. 7 illustrates one embodiment of a format of a pre-access request 4416 and one embodiment of a format for each of the three types of response to the pre-access request 4416: Pre-access response 4417 (indicating that pre-access request 4416 was executed), pre-access rejection response 4418 (indicating that pre-access request 4416 was rejected), and Pre-accession delay response 4419 (indicating that the pre-accession request 4416 has been delayed). In one embodiment, the pre-accession request 4416 indicates whether a pre-accession request 4416 is for an exclusive data jurisdiction or a pre-accession for a shared data jurisdiction. In one embodiment, the pre-accession request 4416 has at least five information fields: pre-accession information field 4401, tag information field 4402, exclusive or shared information field 4403, address information field 4404, and CRC (cyclic redundancy check) information field 4405. In one embodiment, the pre-accession information field 4401 identifies the read request 4416 as a pre-accession request; the tag information field 4402 identifies a particular instance of the Pre-access request 4416; the Exclusive or Shared information field 4403 indicates that the Pre-access request 4416 is either a Pre-access request for Exclusive access or a Pre-access request for Shared access; the Address information field 4404 indicates the address of the data requested by the Pre-access request 4416; the CRC information field 4405 contains a value calculated from the bit patterns of the values in the other information fields in the Pre-access request 4416 and is used to detect an error that occurred during a transfer of the Pre-access request 4416 from a source of the Pre-access request 4416 to a destination of the Pre-access request 4416.

[0203] Fig.7 also illustrates one embodiment of a format for a pre-access response 4417. In one embodiment, the pre-access response 4417 is sent in response to a pre-access request 4416 from a target that received the pre-access request 4416 when the action in the pre-access request 4416 is performed. In one embodiment, the pre-access response 4417 includes at least four information fields: pre-access response information field 4406, tag information field 4407, data information field 4408, and CRC information field 4409.In one embodiment, pre-access information field 4406 identifies pre-access response 4417 as a pre-access response; tag information field 4407 identifies a pre-access response 4417 as a response to a pre-access request 4416 having a tag information field 4407 corresponding to tag information field 4402; data information field 4408 contains data that is the data requested by a pre-access request 4416 having a tag information field 4402 corresponding to tag information field 4407; the CRC information field 4409 contains a value calculated from the bit patterns of the values in the other information fields in the pre-access response 4417 and used to detect an error that occurred during a transfer of the pre-access response 4417 from a source of the pre-access response 4417 to a destination of the pre-access response 4417.

[0204] Fig.7 also illustrates an embodiment of a format of a second type of response to pre-access request 4416, a pre-access rejection response 4418. In one embodiment, pre-access rejection response 4418 is sent in response to pre-access request 4416 from a target that received pre-access request 4416 when the action requested by pre-access request 4416 is rejected (i.e., aborted). In one embodiment, pre-access rejection response 4418 includes at least three information fields: pre-access rejection information field 4410, tag information field 4411, and CRC information field 4412.In one embodiment, the pre-access denial information field 4410 identifies the pre-access denial response 4418 as a pre-access denial response; the tag information field 4411 identifies a pre-access denial response 4418 in response to a pre-access request 4416 having a tag information field 4402 corresponding to the tag information field 4411; the CRC information field 4409 contains a value calculated from the bit patterns of the values in the other information fields in the pre-access denial response 4418 and used to detect an error that occurred during a transfer of the pre-access denial response 4418 from a source of the pre-access denial response 4418 to a destination of the pre-access denial response 4418.

[0205] Fig.7 also illustrates an embodiment of a format of a third type of response to pre-access request 4416, a pre-access delay response 4419. In one embodiment, pre-access delay response 4419 is sent in response to a pre-access request 4416 from a target that received pre-access request 4416 when the action requested by pre-access request 4416 is deferred or delayed (i.e., queued at the target). In one embodiment, pre-access delay response 4419 includes at least three information fields: pre-access delay information field 4413, tag information field 4414, and CRC information field 4415.In one embodiment, the pre-access delay information field 4413 identifies the pre-access delay response 4419 as a pre-access delay response; the tag information field 4414 identifies a pre-access delay response 4419 in response to a pre-access request 4416 having a tag information field 4402 corresponding to the tag information field 4414; the CRC information field 4415 contains a value calculated from the bit patterns of the values in the other information fields in the pre-access delay response 4419 and used to detect an error that occurred during a transfer of the pre-access delay response 4419 from a source of the pre-access delay response 4419 to a destination of the pre-access delay response 4419.

[0206] In one embodiment, a prefetch protocol includes a packed prefetch request consisting of a plurality of prefetch operations, each prefetch operation associated with an address specified in the packed prefetch request. In one embodiment, all prefetch operations in the packed prefetch request are of the same type of packed prefetch request. In one embodiment, a packed prefetch request may include either a plurality of prefetch requests for an exclusive data jurisdiction or a plurality of prefetch requests for a shared data jurisdiction, but not both. In one embodiment, each prefetch operation in the packed prefetch request may be a type of prefetch operation (e.g.,an exclusive jurisdiction prefetch or a shared jurisdiction prefetch) that is independent of any type of other prefetch operation contained in the packaged prefetch request.

[0207] In one embodiment, the one or more addresses in a packed prefetch request are in a sequence of addresses, and each address has a unique index associated with the address. In one embodiment, the value of an index of an address (i.e., the address index) takes the value of the position of the address in the sequence of addresses in the packed prefetch request. In one embodiment, a packed prefetch request includes a packed prefetch request tag that identifies the packed prefetch request. In one embodiment, the packed prefetch request tag consists of a unique tag for the packed prefetch request and a unique tag for each address in the sequence of addresses.

[0208] In one embodiment, a response to a packed prefetch request (e.g., a response containing data, a response indicating a rejection of a prefetch operation, or a response indicating that an operation has been deferred) is associated in the prefetch log with a prefetch operation included in the packed prefetch request. The response includes the tag of the packed prefetch request and an identifier of the address (i.e., an address identifier) of the prefetch operation in the packed prefetch request to which the response is associated. Because a packed prefetch request may contain one or more prefetch operations, a response to an operation in the packed prefetch request must identify the individual prefetch operation to which the response is associated.In one embodiment, the address identifier is the index of an address in the sequence of one or more addresses in the packed prefetch request. In one embodiment, the address identifier is a unique tag for the address included in the packed prefetch tag of the packed prefetch request.

[0209] Fig.8 illustrates one embodiment of a format of a packed prefetch request 5501 and one embodiment of a format for each of three types of response to the packed prefetch request 5501: packed prefetch response 5502 (indicating that a prefetch operation in the packed prefetch request 5501 was performed), packed prefetch reject 5503 (indicating that a prefetch operation in the packed prefetch request 5501 was rejected), and packed prefetch delay 5504 (indicating that a prefetch operation in the packed prefetch request 5501 was delayed). In one embodiment, the packed pre-access request 5501 has at least five information fields: Packed Pre-Access Information Field 5505, Packed Pre-Access Tag Information Field 5506, Exclusive or Shared Information Field 5507,Address information field 5508 and CRC (Cyclic Redundancy Check) information field 5501. In one embodiment, the packed pre-access information field 5505 identifies the packed pre-access request 5501 as a packed pre-access request; the packed tag information field 5506 identifies a particular instance of the packed pre-access request 5501; the exclusive or shared information field 5507 indicates that the pre-access request 5501 is either an exclusive pre-access request or a shared pre-access request; the address information field 5508 indicates one or more addresses of the data requested by the packed pre-access request 5501; the CRC information field 5509 contains a value calculated from the bit patterns of the values in the other information fields in the packed prefetch request 5509 and is used to detect an error,that occurred during a transfer of the packed pre-access request 5509 from a source of the pre-access request 5509 to a destination of the packed pre-access request 5509. In one embodiment, the packed tag information field 5506 contains a unique tag for the packed pre-access request and a unique tag for each address in the address information field 5508.

[0210] Fig.8 also illustrates one embodiment of a format of the packed prefetch response 5502, which is a response to a prefetch operation in the packed prefetch 5501. In one embodiment, the packed prefetch response 5502 is sent by a target that has received a packed prefetch request 5501 in response to the packed prefetch request 5501 when a prefetch operation is performed in the packed prefetch request 5501. In one embodiment, the packed pre-access response 5502 consists of at least four information fields: packed pre-access response information field 5513, operation tag information field 5510, data information field 5514, and CRC information field 5515. In one embodiment, the packed pre-access response information field 5513 identifies the packed pre-access response 5502 as a packed pre-access response;the operation tag information field 5510 identifies the packed prefetch response 5502 in response to a prefetch operation in the packed prefetch request 5501 with the packed prefetch tag 5506; the data information field 5514 contains data that is the data requested by a prefetch operation in the prefetch request 5501 with an operation tag information field 5510 that contains the unique tag that identifies a unique instance of the packed prefetch 5501 and an address identifier that identifies a unique prefetch operation in the unique instance of the packed prefetch 5501;The CRC information field 5515 contains a value calculated from the bit patterns of the values in the other information fields in the packed pre-access response 5502 and used to detect an error that occurred during a transfer of the packed pre-access response 5502 from a source of the packed pre-access response 5502 to a destination of the packed pre-access response 5502. In one embodiment, the address identifier in the operation tag 5510 is a unique address tag contained in the packed pre-access tag information field 5506. In one embodiment, the address identifier in the operation tag 5510 is an index of an address contained in the address information field 5508.

[0211] Fig.8 also illustrates an embodiment of a format of a second type of response to the packed pre-access request 5501, a packed pre-access rejection response 5503. In one embodiment, the packed pre-access rejection response 5503 is sent in response to a pre-access operation in the packed pre-access request 5501 from a target that received the packed pre-access request 5501 when a pre-access operation included in the pre-access request 5501 is rejected (i.e., aborted). In one embodiment, the packed pre-access rejection response 5503 consists of at least three information fields: Packed pre-access rejection response information field 5516,Operation tag information field 5511 and CRC information field 5517. In one embodiment, the packed pre-access rejection response information field 5516 identifies the pre-access rejection response 5503 as a pre-access rejection response; the operation tag information field 5511 identifies the packed pre-access rejection response 5503 as a response to a pre-access operation in a packed pre-access request 5501 with the operation tag information field 5511 containing the unique tag identifying a unique instance of the packed pre-access request 5501 and an address identifier identifying a unique pre-access operation; the CRC information field 5517 contains a value calculated from the bit patterns of the values in the other information fields in the packed pre-accession rejection response 5503 and is used to detect an error,that occurred during a transfer of the packed pre-access rejection response 5503 from a source of the packed pre-access rejection response 5503 to a destination of the packed pre-access rejection response 5503. In one embodiment, the address identifier in the operation tag 5511 is a unique address tag contained in the packed pre-access tag information field 5506. In one embodiment, the address identifier in the operation tag 5511 is an index of an address contained in the address information field 5508.

[0212] Fig.8 also illustrates an embodiment of a format of a third type of response to the packed prefetch request 5501, a packed prefetch delay response 5504. In one embodiment, the packed prefetch delay response 5504 is sent in response to a prefetch operation in the packed prefetch request 5501 from a target that received the packed prefetch request 5501 when a prefetch operation included in the packed prefetch request 5501 is delayed (i.e., queued for delayed execution or delayed until certain conditions are met in a local processor or in a remote processor). In one embodiment, the packed prefetch delay response 5504 consists of at least three information fields: Packed Prefetch Delay Response Information Field 5518,Operation tag information field 5512 and CRC information field 5519. In one embodiment, the packed prefetch delay response information field 5518 identifies the prefetch delay response 5504 as a prefetch delay response; the operation tag information field 5512 identifies the packed prefetch delay response 5504 as a response to a prefetch operation in the packed prefetch request 5501 with the operation tag information field 5512 containing the unique tag identifying a unique instance of the packed prefetch 5501 and an address identifier identifying a unique prefetch operation; the CRC information field 5519 contains a value calculated from the bit patterns of the values in the other information fields in the packed prefetch delay response 5504 and used to detect an error.that occurred during a transfer of the prefetch delay response 5503 from a source of the packed prefetch delay response 5504 to a destination of the packed prefetch delay response 5504. In one embodiment, the address identifier in the operation tag 5511 is a unique address tag contained in the packed prefetch tag information field 5506. In one embodiment, the address identifier in the operation tag 5511 is an index of an address contained in the address information field 5508.

[0213] In other embodiments, a prefetch log includes further responses to the prefetch request 4416 and the packed prefetch request 5501. In one embodiment, a cache associated with a source of a prefetch operation marks data that the cache receives from a prefetch operation as prefetch data. In one embodiment, the cache uses the information contained in the prefetched data to manage cache space (e.g., to determine which data needs to be evicted to accommodate new data).

[0214] Fig.9 is a flowchart illustrating the operation of transactional memory system A 1104 in one embodiment. In one embodiment, transactional memory system A 1104 receives a protocol request for a memory operation on non-prefetch bus 2210 or prefetch bus 2209 (step 301). Abort manager 2207 determines whether the memory operation is a read operation, i.e., a read-type memory operation (decision step 302). If the memory operation is a read operation (decision step 302, branch YES), abort manager 2207 determines whether the address of the read operation matches an address in write group 2205 of a transaction (decision step 303).If the address of the read operation does not match an address in the write group 2205 of a transaction (decision step 303, branch NO), the read operation is executed and, in the case of remote requests, a response is returned by the abort manager 2207 according to an exemplary protocol of . Fig. 7 is sent (step 304) and the process is terminated.

[0215] If the address of the read operation matches an address in the write group 2205 of a transaction (decision step 303, branch YES), the abort manager 2207 determines whether the read operation is a prefetch operation (decision step 305). If the read operation is a prefetch operation (decision step 305, branch YES), the read prefetch is managed by the prefetch manager 2208 (step 312), and the process ends (step 313). If the read operation is not a prefetch operation (decision step 305, branch NO), the transaction associated with the address of the read operation is aborted (step 307), and the process ends (step 313).

[0216] If the memory operation is not a read operation (decision step 302, branch NO), the abort manager 2207 determines that the memory operation is a write operation, and the abort manager 2207 determines whether the address of the write operation matches an address in the read group 2206 of a transaction (decision step 308). If the address of the write operation does not match an address in the read group 2206 of a transaction (decision step 308, branch NO), the abort manager 2207 determines whether the address of the write operation matches an address in the write group 2205 of a transaction (decision step 309). If the address of the write operation does not match an address in the write group 2205 of a transaction (decision step 309, branch NO), the write operation is executed (i.e.the write operation is performed for a local write request, and if the write request came from a remote processor or node, the corresponding cache line is transferred to the remote processor with exclusive jurisdiction) (step 314), and the process ends (step 313). If the address of the write operation matches an address in the write group 2205 of a transactional memory operation (decision step 309, branch YES), the abort manager 2207 determines whether the write operation is a prefetch operation (decision step 310). If the address of the write operation matches an address in the read group 2206 of a transaction (decision step 308, branch YES), the abort manager 2207 determines whether the write operation is a prefetch operation (decision step 310).If the abort manager 2207 determines that the write operation is not a prefetch operation (decision step 310, branch NO), the transaction associated with the address of the write operation is aborted by the abort manager 2207 (step 311), and the process ends (step 313). If it is determined that the write operation is a prefetch for a write operation (decision step 310, branch YES), the write prefetch is managed by the prefetch manager 2208 (step 312), and the process ends (step 313).

[0217] Fig. 10 is a flowchart illustrating the main operations of the prefetch manager 2208 for managing a prefetch operation performed in step 312 of Fig.9 for managing prefetching. Local / Remote Prefetch 2216 in Prefetch Manager 2208 determines whether the prefetch operation is local or remote (decision step 401). If the prefetch operation is local (decision step 401, branch YES), the prefetch operation is executed (step 402) and the process ends (step 419). (In at least one embodiment, performing a prefetch includes indicating a successful prefetch without further action, such as when the requestor and the request receiver access the same cache.) If the local / remote prefetch 2216 determines that this is a remote prefetch operation (decision step 401, branch NO), the queue manager 2214 determines whether to insert the prefetch operation into the prefetch queue 2213 (decision step 420).If queue manager 2214 determines not to insert the prefetch operation into prefetch queue 2213 (step 420, branch NO), prefetch manager 2208 discards the prefetch operation (step 413). Prefetch manager 2208 determines whether to notify the source of the prefetch operation that the prefetch operation has been discarded (decision step 414). If prefetch manager 2208 determines to notify the source of the prefetch operation that the prefetch operation has been discarded (decision step 414, branch YES), the source of the prefetch operation is notified by prefetch manager 2208 (step 415), and operations in prefetch manager 2208 are completed (step 419).If the prefetch manager 2208 determines that the source of the prefetch operation should not be notified that the prefetch operation has been discarded (decision step 414, branch NO), operations in the prefetch manager 2208 are completed (step 419).

[0218] If queue manager 2214 determines that the prefetch operation should be inserted into prefetch queue 2213 (decision step 420, branch YES), queue manager 2214 determines whether prefetch queue 2213 is full (decision step 421). If prefetch queue 2213 is not full (decision step 421, branch NO), queue manager 2214 inserts the prefetch operation into prefetch queue 2213 (step 416). Prefetch manager 2208 determines whether the source of the prefetch operation should be notified that the prefetch operation has been queued (decision step 417).If the prefetch manager 2208 determines that the source of the prefetch operation should be notified (decision step 417, branch YES), the prefetch manager 2208 notifies the source of the prefetch operation (step 418), and then the process completes (step 419). If the prefetch manager 2208 determines that the source of the prefetch operation should not be notified (decision step 417, branch NO), the process completes (step 419).

[0219] If queue manager 2214 determines that prefetch queue 2213 is full (decision step 421, branch YES), queue manager 2214 determines whether to evict an entry from prefetch queue 2213 (decision step 403). If queue manager 2214 determines that an entry needs to be evicted (decision step 403, branch YES), queue manager 2214 evicts an entry from prefetch queue 2215 (step 407). Prefetch manager 2208 determines whether to notify the source of the evicted prefetch operation of the evicted prefetch operation (decision step 408).If the prefetch manager 2208 determines to notify the source of the cleaned prefetch operation (decision step 408, branch YES), the source of the cleaned prefetch operation is notified of the cleanup by the prefetch manager 2208 (step 409), and the new prefetch operation is placed in the prefetch queue 2213 by the prefetch manager 2208 (step 410). If the prefetch manager 2208 determines that the source of the cleaned prefetch operation should not be notified (decision step 408, branch NO), the new prefetch operation is placed in the prefetch queue 2213 by the prefetch manager 2208 (step 410). The prefetch manager 2208 determines whether to notify the source of the new prefetch operation that the new prefetch operation has been queued (decision step 411).If the prefetch manager 2208 determines that the source of the new prefetch operation should be notified that the new prefetch operation has been queued (decision step 411, branch YES), the source of the new prefetch operation is notified by the prefetch manager 2208, indicating that the new prefetch operation has been queued (step 412), and the process completes (step 419). If the prefetch manager 2208 determines that the source of the new prefetch operation should not be notified (decision step 411, branch NO), the process completes (step 419).

[0220] If queue manager 2214 determines that a prefetch operation should not be cleaned up (decision step 403, branch NO), the new prefetch operation is discarded by prefetch manager 2208 (step 404). Prefetch manager 2208 determines whether to notify the source of the discarded prefetch operation of the discarding of the prefetch operation (decision step 405). If prefetch manager 2208 determines that the source of the discarded prefetch operation should be notified that the prefetch operation has been discarded (decision step 405, branch YES), the source of the discarded prefetch operation is notified of the discarding by prefetch manager 2208 (step 406), and the process completes (step 419).If the prefetch manager 2208 determines that the source of the cleaned prefetch operation should not be notified of the discarding of the prefetch operation (decision step 405, branch NO), the process completes (step 419).

[0221] If queue manager 2214 determines that prefetch queue 2213 is not full (decision step 402, branch NO), the prefetch operation is queued by queue manager 2214 (step 416), and prefetch manager 2208 determines whether to notify the source of the prefetch operation that the prefetch operation has been queued (decision step 417). If the prefetch manager 2208 determines that the source of the prefetch operation should be notified that the prefetch operation has been queued (decision step 417, branch YES), the prefetch manager 2208 notifies the source of the prefetch operation (step 418), indicating that the new prefetch operation has been queued (step 418), and the process completes (step 419).If the prefetch manager 2208 determines that the source of the prefetch operation should not be notified that the prefetch operation has been queued (decision step 417, branch NO), the process completes (step 419).

[0222] According to one aspect of an embodiment, queued prefetches are executed when the conflicting transaction terminates either through successful completion or through transaction abort.

[0223] In at least one embodiment of the prefetch manager 2208, the determinations as to whether notifications should be executed in response to discarding, queuing, removing, etc., prefetch requests are set at a design time of the prefetch manager 2208 and incorporated into the logic of the prefetch manager 2208. In one embodiment, such determinations correspond to one or more control values that can be configured using various means, such as software-writable control registers, firmware-writable control registers, control values configured at boot time, e.g., from read-only memory (such as a boot EPROM), or by other reversible or irreversible means (e.g.,by disconnecting a connection at chip production time using a fuse or other means for controllably breaking an electrical connection, the presence or absence of the connection being indicative of a selection.

[0224] Fig.8 illustrates a computer system 800, which is an example of a system including a computer system 1101. Processors 804 and cache 816 generally correspond to processor A 1102, cache system A 1103, processor B 1106, and cache system B 1107. The computer system 800 includes a communications structure 802 that provides communications between the computer processors 804, a memory 806, persistent storage 808, a communications unit 810, and input / output (I / O) interfaces 812. The communications structure 802 may be implemented using any architecture configured for passing data and / or control information between processors (such as microprocessors, communications and network processors, etc.), system memory, peripheral devices, and other hardware components in a system.For example, the data transmission structure 402 may be implemented with one or more buses.

[0225] The memory 806 and persistent storage 808 are computer-readable storage media. In this embodiment, the memory 806 includes random access memory (RAM). In general, the memory 806 may include any suitable volatile or non-volatile computer-readable storage media. A cache 816 is fast memory that increases the performance of the processors 804 by retaining recently accessed data and data near data accessed by the memory 806.

[0226] Program instructions and data used to practice embodiments of the present invention may be stored in persistent storage 808 for execution by one or more of the respective processors 804 via cache 816 and one or more memories of memory 806. In one embodiment, persistent storage 808 includes a magnetic hard disk drive. Alternatively, or in addition to a magnetic hard disk drive, persistent storage 808 may include a solid-state hard disk drive, a semiconductor memory device, read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, or any other computer-readable storage media capable of storing program instructions or digital information.

[0227] The media used by the persistent storage 808 may also be removable. For example, a removable hard drive may be used for the persistent storage 808. Other examples include optical or magnetic disks, thumb drives (USB sticks), and smart cards that are inserted into a drive for transfer to another computer-readable storage medium that is also part of the persistent storage 808.

[0228] In these examples, the communications unit 810 provides for data exchange with other data processing systems or devices. In these examples, the communications unit 810 includes one or more network interface cards. The communications unit 810 may provide data transfers through the use of physical or wireless communications links, or both. Program instructions and data used to practice embodiments of the present invention may be downloaded to persistent storage 808 via the communications unit 810.

[0229] I / O interfaces 812 allow the input and output of data with other devices that may be connected to any computer system. For example, I / O interface 812 may provide a connection to external devices 818, such as a keyboard, keypad, touchscreen, and / or any other suitable input device. External devices 818 may also include portable computer-readable storage media, such as thumb drives (USB flash drives), portable optical or magnetic disks, and memory cards. Program instructions and data used to practice embodiments of the present invention may be stored on such portable computer-readable storage media and may be loaded into persistent memory 808 via I / O interfaces 812. I / O interfaces 812 also connect to a display 820.

[0230] The display 820 provides a mechanism for displaying data to a user and may be, for example, a computer monitor.

[0231] The programs described herein are identified based on the application for which they are implemented in a particular embodiment of the invention. It should be understood, however, that any specific program nomenclature used herein is for convenience only, and thus, the invention is not intended to be limited to use in any particular application identified and / or implied by such nomenclature.

[0232] The present invention may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium (or storage media) having computer-readable program code thereon for causing a processor to perform aspects of the present invention.

[0233] The computer-readable storage medium may be a tangible device capable of retaining and storing instructions for use by an instruction execution unit. A computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing.A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer diskette, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), a portable CD-ROM, a DVD drive (DVD), a memory stick, a floppy disk, a mechanically encrypted device such as punched cards or raised structures in a groove with instructions recorded thereon, and any suitable combination of the foregoing. A computer-readable storage medium, as used herein, need not be designed to transmit transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g.,Light pulses traveling through a fiber optic cable or electrical signals transmitted through a wire.

[0234] Computer-readable program instructions described herein may be downloaded to respective computing / processing units from a computer-readable storage medium or to an external computer or storage device via a network, for example, the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission lines, fiber optic cables, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing unit receives computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in the respective computing / processing unit.

[0235] Computer-readable program instructions for performing operations of the present invention may be assembly language instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or other source code or object code written in any combination of one or more programming languages, including Smalltalk, C++, or the like, and conventional procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server.In the latter scenario, the remote computer may be connected to the user's computer over any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., over the Internet using an Internet service provider). In some embodiments, electronic circuits, including, for example, programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), may execute the computer-readable program instructions using state information of the computer-readable program instructions to personalize the electronic circuit to perform aspects of the present invention.

[0236] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, may be implemented by computer-readable program instructions.

[0237] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other devices that process programmable data to produce a machine such that the instructions, executing via the processor of the computer or other devices that process programmable data, create means for implementing the functions / acts specified in the flowchart and / or the block or blocks of the block diagram.These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium having the instructions stored therein comprises an article of manufacture, including instructions that implement the function / act specified in the flowchart and / or the block or blocks of the block diagram.

[0238] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operations to be performed on the computer, other programmable device, or other device to produce a computer-implemented process, such that the instructions executing on the computer, other programmable device, or other device implement the functions / acts specified in the flowchart and / or the block or blocks of the block diagram.

[0239] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or section of instructions comprising one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions specified in the block may occur out of the order indicated in the figures. For example, two blocks shown in succession may actually execute substantially in parallel, or the blocks may sometimes execute in the reverse order, depending on the functionality involved.It is also noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, may be implemented by special purpose hardware-based systems that perform the specified functions or acts, or by combinations of special purpose hardware and computer instructions.

[0240] The descriptions of the various embodiments of the present invention have been prepared for the purpose of illustration, but are by no means intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the invention. The terminology used herein has been chosen to best explain the principles of the embodiment, practical application, or technical improvement over technology found in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.

[0241] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present invention. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.

[0242] Each of the figures, in addition to illustrating methods and functionality of the present invention at various levels, also illustrates the logic of the method as implemented in whole or in part by one or more units and structures. Such units and structures are configured (i.e., include one or more components such as resistors, capacitors, transistors, and the like connected to enable execution of a process) to implement the method of merging one or more non-transactional memory operations and one or more thread-specific transactional memory operations into one or more cache line templates and a memory buffer in a memory cache.In other words, one or more computer hardware units may be created that are configured to implement the method and processes described herein with reference to the figures and their corresponding descriptions.

[0243] The descriptions of the various embodiments of the present invention have been prepared for the purpose of illustration, but are by no means intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been chosen to best explain the principles of the embodiment, practical application, or technical improvement over technology found in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.

[0244] Embodiments of the present invention may be used in a variety of electronic applications, including advanced sensors, memory / storage, semiconductors, microprocessors, and other applications.

[0245] A resulting unit and structure, such as an integrated circuit (IC) chip, could be distributed by the manufacturer in raw wafer form (i.e., a single wafer with multiple unpackaged chips), as an unpopulated chip, or in a packaged form. In the latter case, the chip is provided in a single-chip package (such as a plastic carrier with leads attached to a control board or other higher-level carrier) or in a multi-chip package (such as a ceramic carrier having either surface interconnects or embedded interconnects, or both). In either case, the chip is then integrated with other chips, discrete circuit elements, and / or other signal processing units as part of (a) an intermediate product, such as a control board, or (b) a final product.The end product can be any product containing IC chips, from toys and other simple applications to high-end computer products with a display, keyboard or other input device, and a central processor.

[0246] The corresponding structures, materials, acts, and equivalents of all means or step-plus-function elements in the following claims are intended to include any structures, materials, or acts for performing the function in combination with other claimed elements as specifically claimed. The description of the present invention has been presented for purposes of illustration and description, but is by no means intended to be exhaustive or limited to the invention in the form disclosed. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the invention.The embodiment was chosen and described in order to best explain the principles of the invention and the practical application and to enable others skilled in the art to understand the invention for various embodiments with various modifications as are suited to the particular use contemplated.

[0247] While the invention has been described in detail in connection with only a limited number of embodiments, it should be absolutely clear that the invention is not limited to such disclosed embodiments. Rather, the invention may be modified to incorporate any number of variations, changes, substitutions, or equivalent arrangements not previously described that are within the spirit and scope of the invention. Furthermore, although various embodiments of the invention have been described, it should be understood that aspects of the invention may be embodied in only some of the described embodiments. Accordingly, the invention is not to be considered limited by the foregoing description. Reference to an element in the singular is not intended to mean "one and only one" unless specifically stated, but instead "one or more."All structural or functional equivalents to the elements of the various embodiments described throughout the disclosure, known to those skilled in the art, or later discovered, are expressly incorporated herein by reference and are intended to be encompassed by the invention. It is therefore to be understood that changes made to certain disclosed embodiments are within the scope of the present invention as set forth by the appended claims.

Claims

[1] A computer system (1101) for preventing a prefetch memory operation from causing a transaction to abort, the computer system (1101) comprising: one or more computer processors (804), one or more computer-readable storage media, and program instructions stored on the one or more computer-readable storage media for execution by at least one of the one or more computer processors (804), the program instructions comprising: Program instructions for receiving a pre-access request (4416) from a remote processor by a local processor; Program instructions for determining whether the prefetch request (4416) conflicts with a transaction of the local processor; Program instructions for responding to at least one of i) a determination that the local processor does not have a transaction and ii) a determination that the prefetch request (4416) does not conflict with a transaction by providing the requested prefetch data; and Program instructions for responding to a determination that the pre-access request (4416) conflicts with a transaction by suppressing processing of the pre-access request (4416), wherein the computer system (1101) comprises a transaction-bound memory system (A1104), and further wherein into the transactional memory system (A1104) non-prefetch operations (ie non-prefetch memory accesses) on a non-prefetch bus (2210) and prefetch operations (iePrefetch request memory accesses) on a bus for prefetch operations (2209), and an address decryption device (2211) decrypts the address of a non-prefetch operation and confirms the target address of a non-prefetch read operation on a read address bus (2201) and the target address of a non-prefetch write operation on a write address bus (2202), and further an address decryption device (2212) decrypts the address of a prefetch operation and confirms the target address of a prefetch read operation on a bus for a prefetch read address (2203) and the target address of a prefetch write operation on a bus for a prefetch write address (2204), and the prefetch manager (2208) is connected to an abort manager (2207) and receives all prefetch operations (2209) that are input to the transactional memory system (A1104) from the prefetch bus (2209). [2] A method for preventing a prefetch memory operation from causing a transaction to abort, the method comprising: Use of a device according to claim 1, Receiving a pre-access request from a remote processor by a local processor from a group of one or more processors, Determining, by the group of one or more processors, whether the prefetch request conflicts with a transaction of the local processor; In response to at least one of i) a determination that the local processor does not have a transaction and ii) a determination that the prefetch request (4416) does not conflict with a transaction, providing, by the group of one or more processors, the requested prefetch data; and in response to a determination that the prefetch request (4416) conflicts with a transaction, suppressing processing of the prefetch request (4416) by the group of one or more processors. [3] The method of claim 2, wherein the pre-access request (4416) is at least one of i) a read pre-access request and ii) a write pre-access request. [4] The method of claim 2, wherein suppressing processing of the pre-access request (4416) includes: Processing a prefetch request cancellation (4416) by the group of one or more processors. [5] The method of claim 4, wherein the method further comprises: Sending a notification by the group of one or more processors to the remote processor indicating that the prefetch request (4416) has been aborted. [6] The method of claim 2, wherein suppressing processing of the pre-access request (4416) includes: queue the prefetch request (4416) by the group of one or more processors. [7] The method of claim 6, wherein the method further comprises: Notifying, by the group of one or more processors, the remote processor that the prefetch request (4416) has been queued. [8] The method of claim 6, wherein the method further comprises: removing the prefetch request (4416) from the queue by the group of one or more processors; and Executing the prefetch request (4416) by the group of one or more processors if the memory address associated with the prefetch request (4416) does not conflict with a transaction. [9] The method of claim 8, wherein the method further comprises: Notifying, by the group of one or more processors, the remote processor that the prefetch request (4416) has been completed. [10] The method of claim 2, wherein suppressing processing of the pre-access request (4416) includes: Executing a delay by the group of one or more processors before executing the prefetch request (4416), wherein a duration of the delay is determined by one or both of i) one or more conditions in the local processor and ii) one or more conditions in the remote processor.

Citation Information

Patent Citations

  • Shared bypass bus structure

    US20030163649A1

  • Indicating a low priority transaction

    US20150212851A1