Method and system for thread transition management

DE112012000965B4Active Publication Date: 2025-07-17INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
DE112012000965
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2011-02-23
Filing Date
2012-02-20
Publication Date
2025-07-17
Estimated Expiration
2032-02-20

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for managing thread transitions in a computer system, the method comprising: Determining that a transition is to be made with respect to the relative usage of two data register sets, the transition being from a mode in which the data register sets contain mirrored data to one in which they contain different data; Determining, based on the transition determination, whether thread data in at least one of the data register sets is to be moved to second-level registers, i.e., to non-register memory that simulates registers; and Move the thread data from at least one set of data registers to second-level registers based on the move detection.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] The present invention relates to processors and, more particularly, to processor threads.

[0002] Advanced processors can typically create a number of threads (e.g., four), which can be sub-elements of a process. These threads are typically assigned identifiers (e.g., 0, 1, 2, and 3) and are executed in a time-multiplexed manner by the processor. Furthermore, the threads may share the same memory (e.g., registers) in the processor or have dedicated memory (e.g., specific registers). When a thread completes, its data is typically removed from processor memory by halting operations, moving the data to be saved to another memory (e.g., main memory), invalidating the processor memory, and then storing the data to be saved back into processor memory.

[0003] The paper “How to Fake 1000 Registers” by Oehmke et al., published in the “Proceedings of the 28th Annual IEEE / ACM International Symposium on Microarchitecture (MICRO'05)” describes a virtual context architecture (VCA), a register file architecture that virtualizes logical register contexts.

[0004] Document US 2003 / 0 033 509 A1 describes a microprocessor with multiple register files. In single-threaded mode, the microprocessor allows a single thread to access multiple register files. In multi-threaded mode, each thread has access to the respective register files. In multi-threaded mode, multiple threads execute concurrently. Circuitry and hardware are provided to facilitate the respective modes and transitions between modes.

[0005] The paper "Register Multimapping: Reducing Register Bank Conflicts Through One-to-Many Logical-to-Physical Register Mapping" by Nam Le Duong and Rakesh Kumar, published by the Coordinated Science Laboratory of the University of Illinois at Urbana-Champaign in August 2008, describes register multimapping, a technique for mapping one architectural register to multiple physical registers.

[0006] The paper "Motivation for Multithreaded Architectures" published as part of the course "CSE 548" at the University of Washington in Winter 2006 describes multithreaded architectures and multithreaded processors.

[0007] The paper “Exploiting Choice: Instruction Fetch and Issue on an Implementable Simultaneous Multithreading Processor” by Tullsen et al., published in the “Proceedings of the 23rd annual international symposium on Computer architecture (ISCA'96)” describes an architecture for simultaneous multithreading. SUMMARY

[0008] The invention is described by the features of the independent claims. Embodiments are specified in the dependent claims.

[0009] The invention is based on the object of providing an improved method and system for thread transition management. This object is achieved by the subject matter of the independent claims.

[0010] In one implementation, a process for managing thread transitions may include determining that a transition is to be made with respect to the relative usage of two data register sets, and determining, based on the transition determination, whether to move thread data in at least one of the data register sets to second-level registers. The process may also include moving the thread data from at least one data register set to second-level registers based on the move determination. The process may be implemented, for example, by a processor.

[0011] The details and features of various implementations are conveyed by the following description together with the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Embodiments of the invention will now be described, by way of example only, with reference to the accompanying drawings, in which: Fig. 1 shows a block diagram illustrating an example system for managing thread transitions. Fig. 2 shows a flowchart illustrating an example process for managing thread transitions. Fig. 3 shows a flowchart illustrating another example process for managing thread transitions. Fig. 4 shows a flowchart illustrating an additional example process for managing thread transitions. Fig. Figure 5 shows a flowchart illustrating another example process for managing thread transitions. Fig. 6 shows a flowchart illustrating another example process for managing thread transitions. Fig. Figure 7 shows a block diagram illustrating an example computer system for which thread transitions can be managed. DETAILED DESCRIPTION

[0013] Processor thread transitions can be managed using a variety of techniques. In specific implementations, managing thread transitions may include the ability to move thread data between data register sets and second-level registers. Being able to move thread data between data register sets and second-level registers can provide a variety of benefits, such as enabling threads to start without halting ongoing operations, enabling the operation of data register sets for multiple threads in a mirrored manner, and / or enabling the execution of a larger number of threads.

[0014] As will be appreciated by one skilled in the art, aspects of the present disclosure may be implemented as a system, method, or computer program product. Accordingly, aspects of the present disclosure may be embodied in the form of a complete hardware environment, a complete software embodiment (including firmware, resident software, microcode, etc.), or in an implementation combining software and hardware aspects, all of which may be generally referred to herein as a "circuit," "module," or "system." Furthermore, aspects of the present disclosure may be embodied in the form of a computer program product, which may be embodied in one or more computer-readable media having computer-readable program code embodied thereon.

[0015] Any combination of one or more computer-readable media may be used. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may include, but is not limited to, a system, apparatus, or device of an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor nature, as well as any suitable combination of the foregoing.More specific examples of a computer-readable storage medium may include (but are not limited to): an electrical connection having one or more wires, a portable computer diskette, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this disclosure, a computer-readable storage medium may be any tangible medium that can include or store a program for use by or in connection with an instruction-executing system, apparatus, or device.

[0016] A computer-readable signal medium may include a propagated data signal having computer-readable program code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal may be embodied in any of a variety of forms, including, but not limited to, electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium, other than a computer-readable storage medium, that can transmit, disseminate, or transport a program for use by or in connection with an instruction-executing system, apparatus, or device.

[0017] The program code embodied in a computer-readable medium may be transmitted by any medium, including, but not limited to, wireless, wired, fiber optic cable, radio frequency (RF), etc., or any suitable combination of the foregoing.

[0018] Computer program code for performing operations for aspects of the disclosure may be written in any combination of one or more programming languages, such as Java, Smalltalk, C++, or the like, and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server.In the latter scenario, the remote computer may be connected to the user's computer over any type of network, such as a local area network (LAN) or a wide area network (WAN), or it may be connected to an external computer (for example, via an Internet service provider over the Internet).

[0019] Aspects of the disclosure are described below with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to implementations. It should be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, may be implemented by computer program instructions.These computer program instructions may be provided to a processor of a general-purpose computer, a dedicated computer, or other programmable data processing apparatus to produce a machine such that the instructions, executing via the processor of the computer or other programmable data processing apparatus, produce a means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0020] These computer program instructions may also be stored in a computer-readable medium that can direct a computer, other programmable data processing apparatus, or other device to function in a particular manner such that the instructions stored in the computer-readable medium produce an article of manufacture including instructions that perform the functions / acts specified in the block or blocks of the flowchart and / or block diagram.

[0021] The computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable device, or other devices to produce a computer-implemented process such that the instructions executing on the computer or other programmable device provide processes for implementing the functions / acts specified in the block or blocks of the flowchart and / or block diagram.

[0022] Fig. Figure 1 illustrates an example system 100 for managing thread transitions. System 100 includes a processor 110, data registers 120, instruction registers 130, main memory 140, and cache memory 150.

[0023] The processor 110 includes an execution core 112, a memory allocator 114, an instruction fetch unit 116, and an instruction queue 118. The execution core 112 is responsible for processing data under the direction of program instructions (e.g., from software) and includes an arithmetic logic unit (ALU) 113 that assists in executing the instructions. The execution core 112 is also capable of concurrently executing multiple independent hardware execution threads. The memory allocator 114, which is explained in more detail below, is responsible for managing the mapping of register data between the data registers 120 and other system memory (e.g., main memory 140 and / or cache memory 150). The instruction fetch unit 116 is responsible for fetching and ordering instructions for execution in multiple parallel hardware threads.The instruction fetch unit 116 is connected to the instruction queue 118, to which the instruction fetch unit dispatches instructions and from which instructions may be issued out of order to the execution core 112 for execution.

[0024] Instructions can take a variety of forms for different instruction types or in different instruction set architectures. For example, instructions can take the form: "opcode RT, RA, RB", where "opcode" is the operation code of the instruction, RT, if present, specifies a logical destination register that receives the execution result (destination operand) of the instruction, and RA and RB, if present, specify the logical source register(s) that provide source operands of the instruction.

[0025] Data registers 120 locally store data that execution core 112 manipulates (e.g., source operands). The data stored by the registers may include, for example, fixed-point values, floating-point values, vector values, decimal values, condition code values, count values, and / or any other suitable data type. Data registers 120 may also store data resulting from operations (e.g., destination operands). The data may also be stored in different locations (e.g., main memory 140 and / or cache memory 150). The contents of data registers 120 may include both architected register values that reflect the current non-speculative state of threads, as well as non-architected register values that reflect working or in-flight values that have not yet been incorporated into the architected state of the threads.The data registers 120 can be universal registers. This means they can store data or pointers (e.g., addresses) to data.

[0026] In the illustrated implementation, the data registers 120 are divided into two sets 122. This division may, for example, enable better access to the execution core 112. For example, by using two register sets 122, the number of read ports providing access to, for example, load / store units, load units, and / or floating-point units may be increased. The depth of (i.e., the number of registers within) each register set 122 is typically limited to a size intended to provide adequate data storage capacity for the number of supported threads while meeting desired targets for access latency and / or power consumption.

[0027] The instruction registers 130 locally store the instructions executed by the execution core 112. The instructions may also be stored in different locations (e.g., main memory 140 and / or an instruction cache).

[0028] Main memory 140 is responsible for storing an extended list of executed instructions along with the associated data. Data registers 120 can access the specific data from main memory 140 as needed, and instruction registers 130 can access the specific instructions from main memory 140 as needed. In certain implementations, main memory 140 may include random access memory (RAM).

[0029] Cache memory 150 is responsible for storing some of the data in main memory 140. Processor 110 can typically access cache memory 150 more quickly than main memory 140. Thus, processor 110 can attempt to store more frequently accessed data in cache memory 150 to speed up operations. Cache memory 150 may be implemented on-chip or off-chip.

[0030] Processor 110 may use data register sets 122 in a variety of ways depending on the operations accessed by a process. For example, if processor 110 is executing only a single thread of a process, which may actually be the process itself, data register sets 122 may contain mirrored data. This means the contents of data register sets 122 may be the same. Thus, the thread may share the data register entries and access execution core 112 through all ports available to data registers 120.

[0031] As another example, if processor 110 is executing two threads of a process, data register sets 122 may contain mirrored data or split data (i.e., one thread's data on each set). For example, if data register sets 122 each contain 64 registers and each thread requires 32 registers, each data register set 122 could contain all the data for each thread, which would allow data mirroring. Thus, each thread can share the data register entries and access execution core 112 through all ports available to data registers 120. However, the data for the two threads could also be split between data register sets 122 (e.g., the data for the first thread could be located on data register set 122a and the data for the second thread could be located on data register set 122b).Thus, each data register set 122 could contain data for only a single set of threads, and each thread could access only a portion of the execution core 112 via half of the ports available to the data registers 120. This mode could be advantageous because it typically leaves more registers in the data register set for "momentary" use, which could be a bottleneck for some workloads.

[0032] Continuing with the example register sizes and thread requirements, once the thread count increases above two, the data register sets 122 may be used in a split mode, as there is insufficient space to store all of the threads' data in a single set. For example, in a four-thread situation, the data for the first and third threads could be located on data register set 122a, and the data for the second and fourth threads could be located on data register set 122b. This would allow each thread to access the execution core 112 through half of the ports available to the data registers 120.

[0033] However, in some situations, there may be a need for more threads than the data registers 120 can support. For example, if four threads are executing and each requires 32 registers, a total of 128 registers are required, which are typically available between set one 122a and set two 122b. However, if eight threads need to be executed, 256 registers are required, which are typically unavailable.

[0034] The memory allocator 114 is responsible for managing register memory when the required memory exceeds that available in the data registers 120. In situations where the required register memory is larger than the available data registers 120, the memory allocator 114 may allocate regions of non-register memory in the system 100 (e.g., the main memory 140 and / or the cache memory 150) to serve as registers. Thus, the data registers 120, in combination with other memories of the system 100, may serve as an increased number of registers, creating a first register level (i.e., the data registers 120) and a second register level (i.e., the regions of non-register memory). This increased number of registers may be operated in similar ways (mirrored or split) as the registers 120, except for the latency introduced by the second form of memory, which typically takes longer to access.

[0035] In more detail, memory allocator 114 may include data structures and logic to track architected and unarchitected register values in processor 110 by mapping logical registers referenced by instructions executed in execution core 112 to specific physical registers in data registers 120 or in second-level registers. For example, memory allocator 114 may assign each thread an identifier (e.g., a number) and have physical pointers to the architected data. Thus, logical memory locations for a thread may be converted to physical pointers to map a thread to a register.

[0036] For example, memory allocator 220 may include a current allocator containing a plurality of entries that track, for all concurrent threads, the physical registers in data registers 120 assigned as target registers for current instructions that have not yet committed execution results to the architected state of processor 110. The data register 120 assigned as the target register of a particular instruction may be specified by placing a register tag (RTAG) of the physical register in the current allocator entry assigned to that particular current instruction.

[0037] Additionally, the memory allocator 114 may include a mapping data structure to track the assignment of the data registers 120 to architected logical registers referenced by instructions across all concurrent threads. In the illustrated implementation, this data structure is implemented as architected allocator caches 115, each of which maps to one of the data register sets 122. (In mirrored memory situations, only one of the architected allocator caches 115 may be used.) For example, the architected allocator caches 115 may include a plurality of lines, each of which contains multiple entries. The lines may be indexed by an architected logical register (LREG) or a subset of the bits that comprise an LREG.The data registers 120 may contain multiple physical registers for the different threads that correspond to the same LREG or LREG group specified by the row index.

[0038] Consequently, each line of an architected allocator cache 115 may contain multiple allocations for a given LREG or LREG group across the multiple concurrent hardware threads.

[0039] Each row of the architected allocator cache 115 may also have a respective associated replacement order vector specifying a replacement order of its entries according to a selected replacement policy (e.g., least recently used (LRU)). The replacement order vector may be updated when a row is accessed at instruction dispatch (when the logical source register of the dispatched instruction hits an architected allocator cache 115), upon completion of an instruction with a logical destination register mapped by the architected allocator cache 115, and when a swap request accessing an architected allocator cache is issued.For example, a swap request may be triggered when a dispatched instruction needs to acquire data that is not located in data registers 120 designated by the architected allocator cache. Swap operations are queued along with the waiting instruction servicing the swap, and a swap then "posts" swap data from the second-level registers in system 100 to a victim in data registers 120. The victim is determined by the architected allocator cache using an LRU algorithm. Once the swap is complete, the waiting instruction may be notified.

[0040] An entry in an architected allocator cache 115 may include a number of fields, including an RTAG field that specifies a physical register in the data registers 120 mapped by that architected entry, and a thread ID field that specifies the hardware thread currently using the specified physical register. The architected logical register currently mapped to the data register 120 specified by the RTAG field may be explicitly specified by an additional field in the entry, or it may be implicitly specified by the index into the architected allocator cache 115 associated with the entry.

[0041] In certain operating modes, the memory allocator 114 may also monitor accesses to the data registers (actual and simulated) by the processor 110 and move frequently accessed data into the data registers 120 and move unaccessed data into non-register memory that simulates registers (i.e., second-level registers). For example, the memory allocator 114 may include swap control logic that manages the transfer of operands between the data registers 120 and second-level registers (e.g., areas of the cache 150 and / or main memory 140 designated as registers). In a preferred embodiment, the swap control logic may be implemented using a first-in, first-out (FIFO) queue that stores operand transfer requests from the memory allocator 114 until serviced.

[0042] Continuing with the example register sizes and thread requirements, if eight threads are required for a process, the data register sets 122 may be used in a split mode because there is not enough space to store all of the threads' data in one set. However, there is also not enough space to store all of the threads' data in the data registers 120. Thus, the memory allocator 114 may allocate additional memory in the system 100 to serve as registers. For example, the memory allocator 114 could allocate 16 registers of main memory 140 to each thread, resulting in each thread having 16 registers in the data registers 120 and 16 second-level registers (in main memory 140).

[0043] The memory allocator 114 may even enable mirroring of the register sets in high-thread-count situations. For example, if there are more than two threads for the example register and thread mapping, there will be an insufficient number of registers in each of the data register sets 122 to enable mirroring of the data register sets. By allocating sufficient space in additional memories of the system 100, the register sets could behave as if they were mirrored. For example, if there are four threads requiring processing, the memory allocator 115 could allocate space in main memory 140 equivalent to each of the data register sets 122 to serve as registers.Thus, data register set 122a plus its equivalent mapping in main memory 140 could hold all the data required by the threads, and data register set 122b plus its equivalent mapping in main memory 140 could hold all the data required by the threads. Memory allocator 114 could then monitor processor accesses to the data and move frequently accessed data into data registers 120 and move unaccessed data into the register-simulating area of main memory 140. One of the architected allocator caches 115 can be used to manage the mapping of registers between register sets 120 and the second-level registers.

[0044] Moving data into the data registers 120 from second-level registers can be accomplished in a variety of ways. For example, the move can be achieved using a swap, as previously explained. While this is occurring, other threads can execute normally without being placed into a wait state. In addition, the move can be achieved using castouts, which occur when a target exits and does not have older architected data for the same logical register in the data registers 120 (i.e., it is located in the second-level registers). The architected memory cache is used (using LRU) to find a victim to move from the first-level registers to the second-level registers to make room for the exiting target.

[0045] When operating in a split mode, threads may be assigned to specific data register sets 122 when they start. For example, a first thread may be assigned to data register set 122a, a second thread to data register set 122b, a third thread to data register set 122a, and a fourth thread to data register set 122b. When starting with a given number of threads, the assignment may be predetermined or determined by an algorithm (e.g., assigning one of the threads to data register set 122 that has the fewest number of active threads). The processor 110 may operate using these threads for any suitable period of time.

[0046] Furthermore, when a thread pauses, its data is removed from data registers 120. For example, when the third thread pauses, its data may be removed from data register set 122a. To accomplish this, memory allocator 114 may be accessed with a thread detector to determine which data register set 122 contains registers for the paused thread and to determine the specific registers for the thread in that set. These registers may then be invalidated through an architected memory cache. For example, the registers may be invalidated immediately, or they may just be marked and cast out at a later time via dummy terminations.The castout operation in this situation is similar to that performed for termination (mentioned above), except that the architected memory cache entry is invalidated and the victim entry is moved from the first-level register (i.e., data registers 120) to the second-level registers. These operations can free up registers for other threads residing on that particular data register set 122, and the entries are typically used relatively quickly by other threads' current data. If the data is not removed, the corresponding entries in the architected allocator cache 115 for the corresponding thread will eventually be evicted by other threads.

[0047] As additional threads are added, memory allocator 114 may allocate registers to the thread without halting operations for the already active threads. For example, memory allocator 114 may allocate second-level registers to serve as registers for a thread being added. Then, as the thread becomes more active, the data may be migrated to data registers 120. The data register set 122 for servicing the thread may be allocated based on a balancing algorithm.

[0048] At some points during operation, processor 110 may decide to transition data register sets 122 from a split mode to a mirrored mode. For example, if the number of threads drops from eight to two, the processor may determine that mirrored mode is more efficient. However, in mirrored mode, the contents of data register sets 122 are synchronized. To achieve this transition, current instructions on all threads may be removed, architected allocator caches 115 may be flushed, the data registers for the remaining threads may be moved to second-level registers (e.g., via dummy exit operations), and data registers 120 may be invalidated. Flushing an architected memory cache may be achieved by performing a castout via dummy exits.For example, the entries in the architected memory cache may be invalidated and their victims moved from the data registers 120 to non-register memory. Therefore, when this is complete, the architected memory caches are empty, meaning that no architected data is present in the data registers 120. Thus, when execution resumes in mirrored mode, the data from the second-level registers may be returned to both sets of data registers 122. As instructions complete, they allocate entries for their target registers in an architected allocator cache 115, evicting victims from other threads if necessary (although initially, no other threads are in the architected memory cache).

[0049] At other points in operation, the processor 110 may decide to transition data register sets 122 from a mirrored mode to a split mode. For example, the processor may decide to transition from a two-mirrored-thread mode to a two-split-thread mode. In these cases, all threads present on the data register set 122a that are to be transitioned to the data register set 122b are flushed from the architected allocator cache 115a so that they can be reloaded into the data register set 122b and its associated architected allocator cache 115b when execution resumes. In addition, when transitioning from a mirrored mode (e.g., a two-threaded mirrored mode) to a more split mode (e.g.,If the thread transitions to a split mode (e.g., a four-thread split mode), instead of moving register data for an existing thread to the second-level registers, the thread may be assigned to data register set 122a, while some / all of the newly added threads are assigned to data register set 122b. However, in situations where a relatively large imbalance will occur (e.g., four threads on one side and one thread on another), some of the data for the existing threads may be moved to second-level registers and reloaded onto data register set 122b to rebalance the sets. It should be noted that if data is moved from one of the data register sets 122 to the second-level registers, the threads may need to wait or pause on that data register set to avoid potential data usage contention.

[0050] In certain operating modes, an imbalance may be created between the data register sets 122. For example, due to persistent threads, the data register set 122a may contain data for four threads and the data register set 122b may contain data for a single thread. In this case, the processor 110 may move the data in the data registers 120 to the second-level registers and reassign the threads to different data register sets 122 (e.g., two and two). The data may then be brought back in from the second-level registers.

[0051] System 100 has a variety of features. For example, when threads pause, the remaining active threads do not need to have their identifiers reassigned, which involves writing the data in the registers to another memory and loading the data back into the registers. Furthermore, a given thread does not need to be on a given set of registers. As another example, active threads do not need to be put into a wait state to allow registers to be allocated for threads that are added. As an additional example, higher numbers of concurrent threads can be realized, and higher numbers of threads can be used in a mirrored mode.

[0052] In further implementations, system 100 may include fewer or additional elements. For example, system 100 may not include cache memory. As another example, system 100 may include one or more memory controllers, network interface cards, I / O adapters, non-volatile data storage, and / or bus bridges, as well as other known elements.

[0053] Fig. Figure 2 illustrates an example process 200 for managing thread transitions. Process 200 may be implemented, for example, by a processor such as processor 110.

[0054] In process 200, a determination is made to determine whether a thread is starting (operation 204). Determining whether a thread is starting can be achieved, for example, by determining whether an interrupt has occurred. If no thread is starting, process 200 continues to wait for a thread to start.

[0055] Once a thread starts, process 200 calls to determine whether registers in a data register set are allocable to the thread (operation 208). Registers in a data register set may be allocable, for example, if the data register set is not currently in use.

[0056] If no registers are allocable in a data register set (e.g., because the data register set already has an allocated thread), process 200 calls an allocation of registers for the thread in second-level registers (e.g., main memory) (operation 212). The registers can be allocated, for example, by marking them for the thread in a memory allocator. Data for executing the thread instructions can then be loaded into the second-level registers.

[0057] Process 200 also calls for moving data for the thread from the second-level registers to the data register set based on usage (operation 216). For example, if data is required to execute thread instructions, the data may be moved from the second-level registers to a data register set.

[0058] However, if registers in a data register set are allocable, process 200 calls an allocate registers for the thread in a data register set (operation 220). Data for executing the thread instructions can then be loaded into the registers in the data register set.

[0059] Fig. 3 illustrates another example process 300 for managing thread transitions. Process 300 may be implemented, for example, by a processor such as processor 110.

[0060] In process 300, a determination is made to see if a thread is halting (operation 304). If no thread is halting, process 300 continues to wait for a thread to halt.

[0061] Once a thread stalls, process 300 calls for locating registers for the stalled thread (operation 308). Locating registers for the stalled thread may be accomplished, for example, by providing a thread identifier to a memory allocator. Process 300 also calls for removing thread entries in the associated registers (operation 312). For example, these registers may be immediately invalidated or marked for castout at a later time. The remaining threads may continue operating as before. This means they do not need to be re-identified and / or re-allocated just because a thread stalls.

[0062] Fig. 4 illustrates an additional example process 400 for managing thread transitions. Process 400 may be implemented, for example, by a processor such as processor 110.

[0063] In process 400, a determination is called to determine whether a transition from split mode to mirrored mode occurs for two data register sets (operation 404). If a transition from split mode to mirrored mode does not occur for two data register sets, process 400 continues to wait for a transition from split mode to mirrored mode.

[0064] Once a transition from a split mode to a mirrored mode occurs, process 400 calls for moving data for the threads from a data register set to second-level registers (operation 408). Moving the data may, for example, involve flushing the architected memory caches using dummy exit operations.

[0065] Process 400 also calls for moving thread data from the second-level registers to data register sets based on usage (operation 412).

[0066] Fig. 5 illustrates another example process 500 for managing thread transitions. Process 500 may be implemented, for example, by a processor such as processor 110.

[0067] In process 500, a determination is called to determine whether an inappropriate imbalance exists between threads assigned to data register sets. For example, if one data register set has four assigned threads and another data register set has zero assigned threads, an inappropriate imbalance may exist because the execution elements are not fully utilized. Imbalances can occur, for example, when threads stall. If no inappropriate imbalance exists, process 500 calls to continue checking for an inappropriate imbalance.

[0068] Once an inappropriate imbalance occurs, process 500 calls for determining the threads to be moved to the other data register set (operation 508). For example, the threads could be evenly distributed between the data register sets. Process 500 also calls for moving data for the threads to be moved from the data register set(s) to second-level registers (operation 512) and reallocating the data register sets for the threads (operation 516). For example, the threads could even be distributed between the data register sets.

[0069] Additionally, process 500 calls for moving data from the second-level registers to the data register sets (operation 520). For example, the data could be moved based on usage or loaded before resuming execution.

[0070] Fig. 6 illustrates an additional example process 600 for managing thread transitions. Process 600 may be implemented, for example, by a processor such as processor 110.

[0071] In process 600, a determination is called to determine whether a transition from a mirrored mode to a split mode for two data register sets occurs (operation 604). If a transition from a mirrored mode to a split mode for two data register sets does not occur, process 600 continues to wait for a transition from a mirrored mode to a split mode.

[0072] Once a transition from a mirrored mode to a split mode occurs, process 600 invokes determining whether new threads are being added (operation 608). If no new threads are being added, process 600 invokes determining which threads are to be transitioned to the second register set (operation 612) and moving data for the transitioning threads from a data register set to second-level registers (operation 616). To accomplish this, for example, threads present on the first data register set and to be transitioned to the second data register set may be taken from an architected allocator cache. Process 600 also invokes moving data for the transitioning thread to the second data register set based on usage when execution resumes (operation 620).

[0073] When new threads are added, process 600 calls for leaving data for existing threads in one data register set (operation 624) and adding the new threads to the second data register set (operation 628). If an inappropriate imbalance occurs thereafter, the data register sets may be rebalanced (e.g., by a process similar to process 500).

[0074] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of systems, methods, and computer program products according to various implementations of the disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of code that may include one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some varying implementations, the functions specified in the blocks may occur in a different order than that specified in the figures. For example, two consecutively depicted blocks may actually execute substantially concurrently, or the blocks may sometimes execute in reverse order depending on the functionality involved.It is also noted that each block of the block diagrams and / or flowchart illustrations, and each combination of blocks in the block diagrams and / or flowchart illustrations, may be implemented by dedicated hardware-based systems or combinations of dedicated hardware and computer instructions that perform the specified functions or acts.

[0075] Fig. Figure 7 illustrates an example computer system 700 in which thread transition management may be performed. System 700 includes a central processing unit (CPU) 710, an input / output system 720, and memory 730, interconnected by a network 740.

[0076] The central processing unit 710 may be, for example, a microprocessor, a microcontroller, or an application-specific integrated circuit, and may include a processor and memory (e.g., registers and / or cache memory). Furthermore, the processor of the central processing unit may operate according to Reduced Instruction Set Computer (RISC) or Complex Instruction Set Computer (CISC) principles. In general, the central processing unit may be any unit that manipulates data in a logical manner.

[0077] The input / output system 720 may include, for example, one or more data transmission interfaces and / or one or more user interfaces. A data transmission interface may be, for example, a network interface card (wireless or wireless) or a modem. A user interface may be, for example, a user input device (e.g., a keyboard, touchpad, stylus, or microphone) or a user output device (e.g., a monitor, display, or speaker). In general, the system 720 may be any combination of devices through which a computer system can receive and output data.

[0078] Memory 730 may include, for example, random access memory (RAM), read-only memory (ROM), and / or disk storage. Various elements may be stored in different locations of memory at different times. Memory 730 may generally be any combination of units for storing data.

[0079] Memory 730 contains instructions 732 and data 736. The instructions include an operating system 733 (e.g., Windows, Linux, or Unix) and applications 734 (e.g., word processing, spreadsheets, drawing, scientific applications, etc.). Data 736 includes the data required by and / or generated for the applications 734.

[0080] The network 740 is responsible for transferring data between the processor 710, the input / output system 720, and the memory 730. The network 740 may, for example, include a number of different types of buses (e.g., serial and parallel).

Claims

[1] A method for managing thread transitions in a computer system, the method comprising: Determining that a transition is to be made with respect to the relative usage of two data register sets, the transition being from a mode in which the data register sets contain mirrored data to one in which they contain different data; Determining, based on the transition determination, whether thread data in at least one of the data register sets is to be moved to second-level registers, i.e., to non-register memory that simulates registers; and Move the thread data from at least one set of data registers to second-level registers based on the move detection. [2] The method of claim 1, further comprising: Determine whether starting at least one additional thread is associated with the transition; Determining the identity of active threads to be transferred to the second data register set based on a non-starting additional thread; Moving the data for the identified threads from a data register set to second-level registers based on the identification determination; and Move the identified thread data to the second data register set. [3] The method of claim 2, wherein moving the identified thread data to the second data register set is performed based on the usage of the thread data. [4] The method of claim 1, further comprising: Determine whether starting at least one additional thread is associated with the transition; Maintaining data for existing threads on the first data register set based on an additional starting thread; and Assign registers on the second data register set to threads being started. [5] The method of claim 1, wherein the transition is from a mode in which the data register sets contain different data to one in which they contain mirrored data. [6] The method of claim 5, further comprising moving the data for the active threads from the data register sets to the second level registers. [7] The method of claim 6, further comprising moving the thread data from the second level registers to the data register sets based on usage. [8] The method of claim 1, further comprising: Determine whether a thread starts; Determining, based on a starting thread, whether registers in a data register set are assignable to the thread; and Allocating second-level thread registers based on thread-unallocable data registers. [9] The method of claim 8, further comprising moving data for the starting thread from the allocated second level registers to a data register set based on usage. [10] The method of claim 1, further comprising: Determine whether a thread is halting; Determining registers for the persistent thread; and Flush the registers in a data register set for the lingering thread. [11] The method of claim 1, further comprising: Determine whether there is an inappropriate imbalance between the data register sets assigned to threads; Determining at least one thread to be moved from one data register set to another data register set; Moving data for the threads to be moved from at least one set of data registers to second-level registers; Reallocating data register sets for the threads; and Moving data from the second-level registers to the data register sets. [12] System comprising: a computer memory; two sets of data registers connected to the computer memory; and a processor connected to the two data register sets, the processor being arranged to perform all the steps of any one of the preceding claims.

Citation Information

Patent Citations

  • Architectural reuse of registers for out of order simultaneous multi-threading

    US20030033509A1