Method, computer program product, and system for sharing closely spaced data caches
By sharing data caches between processors and accessing neighboring L1 caches, the method effectively increases cache size and reduces latency, addressing the balance between cache size and latency in modern computing systems.
Patent Information
- Application Number
- DE102012224265
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2012-01-04
- Filing Date
- 2012-12-21
- Publication Date
- 2025-07-03
- Estimated Expiration
- 2032-12-21
AI Technical Summary
The challenge of balancing cache size and latency in modern computing systems, where increasing cache size for processors leads to higher hit rates but longer latency, is addressed by sharing data caches between processors without increasing physical size.
A method and system for accessing data caches involve searching directories of adjacent processors to determine data availability, allowing data retrieval from neighboring L1 caches when local caches miss, thereby effectively increasing the cache size without physical expansion.
This approach reduces cache misses and latency by allowing access to neighboring L1 caches, providing faster data retrieval with minimal additional hardware, thus optimizing cache performance without increasing chip area.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUNDField of the invention
[0001] The present invention relates generally to data caches for processors and, more particularly, to a method, computer program product, and system for sharing data caches between processors. Description of the underlying technology
[0002] The size of the various cache levels in a cache hierarchy—i.e., Level 1 (L1) cache, Level 2 (L2) cache, etc.—remains an important design feature of modern computing systems. As the size of the cache increases, the computer system can store more data in the cache, but this increases the time it takes for the processor to locate the data within the cache—i.e., latency. Thus, larger caches have higher hit rates but also longer latency. However, because caches are typically located near the processors requesting the data—e.g., on the same semiconductor chip, where area is limited—increasing the cache size to store larger amounts of data may not be feasible. These considerations must be balanced when deciding on cache size.
[0003] The document “Chip Multiprocessor Caches, Places and Management (Short Course)” by Andreas Moshovos from July 2009 (URL: https: / / web.archive.org / web / 20111113012132 / http: / / www.eecg.toronto.edu / ∼moshov os / ) describes cache designs for chip multiprocessors.
[0004] The paper "Improving Support for Locality and Fine-grain Sharing in Chip Multiprocessors" by Hemayat Hossain, Sandhya Dwarkadas, and Michael C. Huang, published in the "Proceedings of the 17th International Conference on Parallel Architectures and Compilation Techniques. New York, NY, USA: ACM, 2008 (PACT '08). pp. 155-165. - ISBN 978-1-60558.282-5, http: / / doi.acm.org / 10.1145 / 1454115.1454138" describes methods for data sharing across multiple threads or processes. The dissertation "Effective On-Chip Utilization in Chip Multiprocessors" by Hemayat Hossain, published by the University of Rochester, Department of Computer Science. pp. i, vi-xvi, 2-15, 58-60, 71. URL: https: / / www.cs.rochester.edu / ~hossain / phd thesis hemayet.pdf describes cache coherence protocols for exploiting on-chip low-latency interconnects. SUMMARY
[0005] The invention is based on the object of providing an improved method, computer program product and system for accessing data cache memories associated with multiple processors.
[0006] The solution to the problem underlying the invention is achieved by the subject matter of the independent patent claims.
[0007] Embodiments of the invention provide a method, a system, and a computer program product for accessing data caches associated with multiple processors. The method and the computer program product comprise searching a first directory to determine whether a first cache associated with a first processor contains data required to execute an instruction being executed by the first processor, the first directory comprising an index of the data stored in the first cache. The method and the computer program product comprise searching a second directory to determine whether a second cache associated with a second processor contains the required data, the second directory comprising an index of the data stored in the second cache.Upon determining that the data is located in the second cache, the method and the computer program product also comprise transmitting a request to the second processor to retrieve the data from the second cache. Upon determining that the data is not located in either the first or second cache, the method and the computer program product also comprise transmitting a request to another memory associated with the first processor to retrieve the data.
[0008] The system includes a first processor and a second processor. The system also includes a first cache in the first processor and a second cache in the second processor. The system includes a first directory in the first processor having an index of the data stored in the first cache, wherein searching the index of the first directory indicates whether the first cache contains data required to execute an instruction being executed by the first processor. The system also includes a second directory in the first processor having an index of the data stored in a second cache of the second processor, wherein searching the index of the second directory indicates whether the second cache contains the required data.After determining that the data is located in the second cache, the system sends a request to the second processor to retrieve the data from the second cache. After determining that the data is not located in either the first or second cache, the system sends a request to another memory associated with the first processor to retrieve the data. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] To realize and fully understand the above aspects, reference may be made to the detailed description of embodiments of the invention briefly summarized above in conjunction with the accompanying drawings. Fig. 1 is a chip with multiple processors according to an embodiment of the invention. Fig. 2 is a flowchart for accessing an L1 cache in an adjacent processor according to one embodiment of the invention. Fig. 3 is a system architecture view of a plurality of processors sharing an L1 cache according to an embodiment of the invention. Fig. 4 is a flowchart for determining when to insert a request to fetch data from an L1 cache of a neighboring processor according to one embodiment of the invention. Fig. 5A through 5B are block diagrams illustrating a networked system for executing client-submitted jobs on a multi-node system according to embodiments of the invention. Fig. 6 is a diagram illustrating a multi-node job system according to embodiments of the invention. DETAILED DESCRIPTION
[0010] The L1 caches for multiple computer processors can be shared to effectively create a single (i.e., virtual) L1 cache. This does not increase the physical size of the L1 cache, but it does increase the probability of a cache hit—that is, at least one of the shared L1 caches contains the requested data. The advantage is that accessing an L1 cache of a neighboring processor requires fewer clock cycles and thus lower latency than accessing the processor's L2 cache or other off-chip memory.
[0011] In many high-performance computers, where hundreds of individual processors may be located in close proximity to one another—for example, on the same semiconductor chip or semiconductor substrate—the processors may execute threads that continually load, process, and store the same data. For example, in a parallel computer system, the processors may perform different tasks within the job submitted by the user. If these tasks are related, the processors are likely to retrieve copies of the same data from main memory and store that data in their own caches for faster access. Accordingly, by providing access to adjacent processors, the size of the processor's cache is effectively increased to retrieve data from a processor's L1 cache without increasing the real estate occupied by the on-chip caches.
[0012] When a processor looks for data to execute an instruction in its pipeline, it can check whether the data is stored in its own L1 cache. The processor can also check whether the data is in the L1 cache of a neighboring processor. If the data is not in its own cache but is in its neighbor's L1 cache, it can send a request for the data to its neighboring processor. The request can then be inserted into the neighboring processor's pipeline, so that the data is passed from the neighbor's L1 cache to the requesting processor. That is, the neighboring processor treats the request as if it originally came from its own pipeline. However, after fetching the data, it is not used by the neighboring processor's pipeline but is passed back to the requesting processor.
[0013] Additionally, the processor may include selection logic that determines when a request for data should be inserted into the neighboring processor's pipeline. For example, the logic may wait until the neighboring processor is not busy or has an empty slot in its pipeline before inserting the request to ensure that the request does not interrupt the neighbor's pipeline. Or the selection logic may assign priorities to the processors such that when a request is received from a higher-priority processor, the pipeline of the lower-priority processor is interrupted to insert the request. However, if the logic determines that the request must wait, the processor may use a queue to hold requests until they can be inserted.
[0014] If the data requested by the processor is not found in the local L1 cache or the neighboring processor's L1 cache, the request may be forwarded to another cache level in a cache hierarchy or to off-chip RAM.
[0015] In the following, reference is made to embodiments of the invention. However, it should be understood that the invention is not limited to the individual embodiments described. Rather, it is intended that the invention be implemented and practiced by any combination of the following features and elements, regardless of their reference to various embodiments. Although embodiments of the invention may achieve advantages over other possible solutions and / or over the prior art, it is not a limitation of the invention whether a single advantage is achieved by a particular embodiment or not. Thus, the following aspects, features, embodiments, and advantages are merely illustrative and are not to be considered elements or limitations of the appended claims unless expressly relied upon in the claim(s).Likewise, reference to the term "the invention" is not to be construed as a generalization of any inventive subject matter disclosed herein, nor as an element or limitation of the appended claims, unless expressly stated in one or more claims.
[0016] Those skilled in the art will appreciate that aspects of the present invention may be implemented as a system, method, or computer program product. Accordingly, aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which may be generally referred to herein as a "circuit," "module," or "system." Furthermore, aspects of the present invention may take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code stored thereon.
[0017] Any combination of one or more computer-readable media may be used. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof.More specific examples (a non-exhaustive list) of the computer-readable storage medium may include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the context of this document, the computer-readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with a system, apparatus, or device for executing instructions.
[0018] A computer-readable signal medium may include a propagated data signal with program code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal may take a variety of forms, including, but not limited to, electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium, other than a computer-readable storage medium, capable of transmitting, propagating, or transporting a program for use by or in connection with an instruction-executing system, apparatus, or device.
[0019] The program code stored on a computer-readable medium may be transmitted using any suitable medium, including, but not limited to, wireless, wired, fiber optic, RF, etc., or any suitable combination thereof.
[0020] The computer program code for performing operations for aspects of the present invention may be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, Smalltalk, C++, or the like, and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code may execute entirely on a user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server.In the latter scenario, the remote computer can be connected to the user's computer over any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, over the Internet using an Internet service provider).
[0021] In the following, aspects of the present invention are described with reference to flowchart and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be appreciated that each block of the flowchart and / or block diagrams, and combinations of blocks in the flowchart and / or block diagrams, may be implemented by instructions of a computer program. These computer program instructions may be supplied to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions executed by the processor of the computer or other programmable data processing apparatus produce means for implementing the functions / acts specified in the block or blocks of the flowchart and / or block diagram.
[0022] These computer program instructions may also be stored in a computer-readable medium that can cause a computer, other programmable data processing apparatus, or other devices to operate in a particular manner such that the instructions stored in the computer-readable medium produce an article of manufacture that includes instructions that implement the function / act specified in the block or blocks of the flowchart and / or block diagram.
[0023] The instructions of the computer program may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operations to be performed on the computer, other programmable device, or other devices such that the instructions to be executed on the computer or other programmable device provide processes for implementing the functions / acts specified in the block or blocks in the flowchart and / or block diagram.
[0024] Fig. 1 is a multi-processor chip according to one embodiment of the invention. Computer processors fabricated on semiconductor wafers may include hundreds or even thousands of individual processors. As used herein, a "processor" includes one or more execution units with at least one cache. Thus, in a processor with multiple processor cores, each processor core may be considered a single processor if it includes at least one independent execution unit and at least one separate cache level. Chip 100 includes four single-processor cores 110A-110D. These may be separate processors or cores for a single multi-core processor. Each processor 110A-110D includes one or more pipelines, typically consisting of two or more execution stages. Generally, the pipelines are used to execute one or more threads 120A-120D.For example, a multithreaded processor may use a single pipeline to execute multiple threads concurrently so that if one thread is blocked, for example, due to a failed cache access, another thread in the pipeline can execute while the thread waits for data to be retrieved from memory. According to one embodiment, threads 120 access the same data set stored in main memory or other memory.
[0025] The caches 125A through 125D may represent a single level of cache or a cache hierarchy—e.g., an L1 cache, L2 cache, L3 cache, etc. According to one embodiment, at least one of the caches 125 is located in an area of the chip 100 reserved for the processor—e.g., the L1 cache is located within the geometric boundaries of the processor 110 on the chip 100—while other caches 125 may be located elsewhere on the chip 100.
[0026] According to one embodiment, at least one level of cache 125 may be written to only by one processor 110, while other levels of cache 125 may be written to by multiple processors 110. For example, in chip 100, each processor 110A-110D may have its own L1 caches that can only be written to by the processor to which it is connected, but two or more processors 110 may write to and read from the L2 caches. Furthermore, processors 110A and 110B may have equal access to the same coherent L2 cache, while all processors 110A-110D have equal access to the same L3 cache. For processors executing threads that access the same data set, sharing levels of coherent cache may be advantageous because it may save chip area without penalty to system latency.
[0027] Fig. Figure 2 is a flowchart for accessing an L1 cache in a neighboring processor according to one embodiment of the invention. Although the L1 cache (or any other cache) may be exclusive to each processor, preventing other processors from writing data to the cache, chip 100 may provide a data path (i.e., one or more dedicated traces) for a neighboring processor to read data from the L1 cache of a neighboring processor.
[0028] Fig. Figure 3 is a system architecture view of a plurality of processors having equal access to an L1 cache according to an embodiment of the invention. A method 200 for accessing the L1 cache of a neighboring processor is described in Fig. 2 shown.
[0029] According to one embodiment, processor 301 is fabricated to reside on the same semiconductor chip 100 as processor 350. The dashed line shows the separation between the hardware elements located within the two separate processors 301, 350. Each processor 301, 350 includes a pipeline consisting of a plurality of execution stages 306A through 306F and 356A through 356F. Each execution stage 306, 356 may represent an instruction fetch, decode, execute, memory access, write-back, etc., and may include any number of stages. The pipelines may be any type of pipeline, for example, a fixed-point, floating-point, or load / store pipeline. Furthermore, processors 301, 350 may include any number of pipelines, which may be of different types.The pipelines shown represent simplified versions of a single pipeline, but embodiments of the invention are not limited thereto.
[0030] In step 205 of method 200, the current instruction executing in the pipeline of processor 301 requires fetching data from memory to perform the operation associated with the instruction. As used herein, the term "memory" refers to both the main memory (e.g., RAM) located on the processor chip and the cache memory, which may be located on or off the processor chip. For example, the instruction may be a load instruction that requests that data corresponding to a particular memory address be loaded into one of registers 302. This instruction may be followed by an add instruction that then adds the data stored in two of registers 302 together. Thus, in execution stages 306C through 306F, the pipeline determines whether the data is stored in the L1 cache associated with processor 301.If so, the data is then fetched from the L1 cache, placed on the bypass bus 336, and inserted into the pipeline at the execution stage 306A. The bypass control unit 304 and the bypass and operand multiplexing unit 308 enable the processor 301 to insert the fetched data directly into the pipeline where the data is needed without first storing the data in registers 302. The operand execution unit 310 (e.g., an ALU or a multiply / divide unit) can then add the data fetched from the L1 cache with the data stored in the other register 302.
[0031] In step 210, the processor 301 determines whether the requested data is stored in the L1 cache. According to one embodiment, each processor 301, 350 includes a local cache directory 314, 364 containing an index (not shown) of the data stored in the L1 cache. In particular, the local cache directory 314, 364 contains the addresses of the various memory pages stored in the L1 cache. As the L1 cache removes and replaces memory pages according to the replacement policy, the processors 301, 350 update the local cache directories 314, 364. The embodiments disclosed herein are not limited to any particular replacement policy. When changes are made to the data stored in the L1 cache that must be written to main memory, the embodiment disclosed herein is also not limited to any particular write policy, e.g.Copy through, write back or copy back.
[0032] Using the address of the requested data, processor 301 can search local cache directory 314 to determine if the corresponding page or pages are currently in the L1 cache. If so, a cache hit has occurred. If not, a cache miss has occurred, and processor 301 must look elsewhere for the data.
[0033] According to another embodiment, the pipeline of processors 301, 350 may not include a separate hardware unit serving as the local cache directory 314, 364, but instead insert the index directly into the L1 cache to search for the data there.
[0034] Not shown is that typical pipelines also include address translation units for translating virtual memory addresses into physical memory addresses and vice versa. This translation may occur before or simultaneously with the index entry into the local cache directory 314. However, this function is not shown for clarity.
[0035] If searching the index in local cache directory 314 results in a cache hit, the processor uses cache load unit 318 to retrieve the data from the L1 cache in step 215. In execution stage 306E, the retrieved data is processed by fetch unit 322 to an expected format or alignment and then placed on bypass bus 336 for delivery to an earlier execution stage in the pipeline, as discussed above.
[0036] If the data is not found in the L1 cache at step 220, the processor 301 may search the adjacent cache directory 316 to determine if the L1 cache of a neighboring processor—i.e., processor 350—contains the data. For a typical processor, a cache miss for the L1 cache will cause the processor to work its way up through the cache (and ultimately through main memory or data stores) to find the data. However, in computing environments where neighboring processors are expected to have threads accessing the same data set, this fact may be used to increase the size of the L1 cache for the processors. In particular, an adjacent cache directory hardware unit 316 may be associated with one or both of the processors 301, 350.Since the directory 316 only needs to provide an index of what is currently stored in the L1 cache of the neighboring processor, it can be physically much smaller than would be necessary to increase the size of the L1 cache.
[0037] The term "adjacent" as used herein refers to two processors that are located at least on the same semiconductor chip 100. Furthermore, the two processors may be two cores of one and the same multi-core processor. Furthermore, the adjacent processors may be manufactured to be mirror images of each other. That is, the structure of processor 301 is a mirror image of the structure of processor 350. This places various hardware units in close proximity to the adjacent processor to facilitate access to the functional units of the adjacent processor. In particular, the select units 334 and 384, as well as the adjacent queues 332, 382 (the functions of which are discussed below), are placed near the respective processors.It should be noted, however, that although the location of the functional units of the two processors 301, 350 may be substantially mirror images of each other, the data buses / paths may be located at different locations to facilitate the transfer of data between the two processors 301, 350.
[0038] According to one embodiment, the adjacent cache directory 316 may be updated by the processor 350. That is, the processor 350 may, using a data path (not shown), update the adjacent cache directory 316 to resemble the local cache directory 364. More specifically, the processor 350 may push update data during each update of its own local cache directory 364. In this way, the adjacent cache directory 316 represents a read-only memory for the processor 301, which cooperates with the processor 350 to ensure that the data stored within the directory 316 represents what is currently stored in the L1 cache associated with the processor 350.Alternatively, for example, processor 301 may copy the index of local cache directory 364 of processor 350 to neighboring cache directory 316 from time to time.
[0039] According to one embodiment, during execution stage 306D, the local cache directory 314 and the adjacent cache directory 316 may be accessed concurrently. That is, the processor 301 may use the memory address of the requested data to search both directories 314, 316 concurrently. During execution stage 306E, the resulting identifiers (i.e., what the memory address was compared to within directories 314, 316) are sent to the identifier comparison unit 320 to determine whether a cache hit or a cache miss occurred. If the memory address is found in both directories 314, 316 (i.e., a cache hit for both directories), the data is retrieved from the local L1 cache, according to one embodiment.However, according to further embodiments, if the requested data is stored in both the local and neighboring L1 caches, the data may also be retrieved from the neighboring L1 cache if, for example, the local L1 cache is corrupted or currently unavailable.
[0040] It should be noted that Fig. 3 illustrates concurrent access to the cache using the cache load unit 318 and the local cache directory 314. Depending on whether the local cache directory 314 returns a cache hit or a cache miss, a determination is made as to whether the data retrieved from the cache using the cache load unit 318 is forwarded or flushed. In other pipeline arrangements, the pipeline may wait two additional cycles to access the cache using the cache load unit 318 after determining, via the local cache directory 314, that the data is in the L1 cache. The former technique may improve performance, while the latter may save power. Nevertheless, the embodiments disclosed herein are not limited to either technique.
[0041] According to one embodiment, the adjacent cache directory 316 may only be accessed when the tag comparison unit 320 reports a cache miss in the local cache directory 314. For example, a system administrator may configure the processors 301, 350 to enter a power-saving mode in which they do not access the directories 314, 316 concurrently. While this trade-off may save power, it increases latency. For example, the processor 301 may have to wait until execution stage 106F before detecting a cache miss in the local L1 cache. Thus, the lookup in the adjacent cache directory 316 may be delayed by approximately three clock cycles.Additionally, shared access to the L1 caches may be configured so that a user administrator can completely disable the ability of processors 301, 350 to access each other's L1 cache.
[0042] If the tag comparison unit 320 determines that neither L1 cache contains the data, in step 225, the cache miss logic unit 324 forwards the request to an L2 cache queue 330. This queue 330 manages access to an L2 cache. The L2 cache may be maintained coherently across multiple processors, or it may be accessible only by processor 301. If the memory page corresponding to the memory address of the requested data is not in the L2 cache, the request may proceed to higher levels in the cache hierarchy or to the computer system's main memory. However, if the requested data is found in the L2 cache, the data may be placed on the bypass bus 336 and forwarded to the appropriate execution stage for processing.
[0043] If the neighboring cache directory 316 returns a cache hit, but the local cache directory 314 indicates a cache miss, the processor may insert a request into a pipeline of the neighboring processor 350 to retrieve the data from its L1 cache at step 230. Fig. 3 shows that the adjacent cache hit unit 326 forwards the request to the select unit 334. The select unit 334 may determine that the request should not be immediately inserted into the pipeline of the processor 350. If so, the request may be stored in the adjacent queue 332 for later insertion. A more detailed description of the select units 334, 334 is provided in the discussion of the following Fig. 4 reserved.
[0044] Select unit 334 may use multiplexer 362 to insert the request into the pipeline of processor 350. Multiplexer 362 includes two inputs: one for receiving requests from processor 301 to retrieve data from the L1 cache, and another input for receiving requests from its own pipeline. Select unit 334 may control the select line of multiplexer 362 to determine whether to insert a request from processor 301. Not shown, select unit 334 may control additional logic within execution stages 356A through 356F to ensure that the inserted request does not corrupt the instructions and data currently in the pipeline. For example, inserting the request may require inserting a zero-operation (NOP) instruction or halting upper execution stages 356A through 356C to avoid data loss.
[0045] According to one embodiment, instead of inserting the request into the pipeline of the neighboring processor, processor 301 may include the necessary hardware and data paths to directly retrieve the data from the L1 cache of processor 350. However, by inserting the request into the pipeline of the neighboring processor instead of directly retrieving the data from the neighboring L1 cache, space on chip 100 may be saved because, according to the former solution, no additional redundant hardware units are required for processor 301, whose functions can already be performed by the hardware units located in processor 350.Thus, by adding adjacent cache directory 316 and select unit 334 to processor 301 and multiplexer 362 (and associated data paths) to processor 350, processor 301 is able to access an adjacent L1 cache with only a few additional hardware units by utilizing many of the units already associated with processor 350.
[0046] Once the select unit 334 inserts the data into the pipeline of processor 350 via multiplexer 362, in step 235, the memory address in the request is sent to the cache load unit 368, which initiates the fetching of the appropriate memory pages from the L1 cache of processor 350. The fetch unit 372 processes the data to the expected format or alignment, and the processor 350 then forwards the data to the bypass bus 336 of processor 301. The bus then inserts the data into the processor 301's own pipeline as if the data had been fetched from the L1 cache located in processor 301.
[0047] The same method 200 may be performed for the processor 350 to use the Fig. 2 to access the mirror-image functional units and data paths in the L1 cache memory of the processor 301.
[0048] Fetching data from a local L1 cache typically requires approximately four clock cycles. Fetching data from a local L2 cache typically requires 20 to 50 clock cycles. If the local cache directory 314 and the neighboring cache directory 316 are accessed concurrently, the data can be fetched from a neighboring processor's L1 cache in approximately eight clock cycles. Fetching the directories sequentially may require approximately 12 to 15 clock cycles. This demonstrates that by providing access to a neighbor's L1 cache, the size of the L1 caches is virtually doubled and the latency is shorter than accessing an L2 cache, without adding more than three or four additional functional units to the processors.
[0049] According to one embodiment, processors 301, 350 can only read from the L1 cache of a neighboring processor, so requests received from a neighboring processor do not affect which data is evicted and then stored in the local L1 cache. For example, the L1 cache of processor 350 is read-only for processor 301, so it cannot store data in the L1 cache, either directly or indirectly. That is, a replacement policy for the L1 cache of processor 350 can consider requests for data only from threads executing on the local pipeline when determining whether to invalidate and evict data in the local L1 cache.For example, many replacement policies consider least recently used (LRU) data when determining what data to replace with new data after a cache miss. Considering data most recently accessed by the neighboring processor may be irrelevant to the threads executing in the local pipeline. Thus, when determining the LRU, considering accesses by the neighboring processor may result in the removal of data that may be accessed by threads executing in the local pipeline. Thus, in this embodiment, accesses from the neighboring processor (e.g., processor 301) may be ignored, thereby preventing processor 301 from indirectly writing data to the L1 cache.
[0050] For example, if the L1 cache of processor 350 contains memory pages that are frequently accessed by processor 301 but rarely accessed by threads executing on processor 350, the cache replacement policy can evict those memory pages by ignoring the accesses by processor 301. This can actually lead to increased performance because processor 301 can now include those memory pages in its local L1 cache, which it can access even faster than the L1 cache of processor 350.
[0051] Of course, the system administrator can configure the system so that the replacement rule ignores accesses from neighboring processors when determining the LRU data. For example, the administrator may know that threads executing on both processors use the same data and thus may want to share the L1 caches to (1) prevent cache misses and (2) avoid the constant swapping of memory pages in the L1 cache for memory pages stored in the cache hierarchy.
[0052] According to one embodiment, more than two processors may be interconnected via data lines to effectively increase the size of the L1 caches. For example, processor 301 may include a second adjacent cache directory containing an index to an L1 cache located on a third processor. The third processor may be located below processor 301 and may be a mirror image with respect to a horizontal line separating the two processors. Furthermore, select unit 334 or adjacent cache hit unit 326 may be configured to determine which of the adjacent processors has the data and forward the request to the appropriate processor.
[0053] Fig. 4 is a flowchart according to one embodiment of the invention for determining when to insert a request to fetch data from an L1 cache of a neighboring processor. In particular, Fig. 4 a method 400 for inserting a request for data into the pipeline of a neighboring processor - i.e., step 230 of Fig. 2. As mentioned above, the selection unit 334 may play a role when a request is to be inserted into the pipeline of the processor 350. Preferably, this is done in such a way that the instructions and data requests associated with the processor 350 are not interrupted, although waiting for a period without interrupting the pipeline may not be desirable in some cases.
[0054] In step 405, the selection unit 334 uses predefined criteria to determine when to insert a request into the pipeline of the processor 350. These criteria may include waiting for the pipeline to be idle or interrupted, establishing a priority between processors, or maintaining a predetermined ratio.
[0055] The select unit 334 ensures that inserting the request does not affect instructions and requests already executing in the execution stages 356A through 356F. While the processor is busy executing, its pipeline may have a gap during which the execution stage is not currently moving or operating on data (i.e., during a NOP or an interrupt). By replacing a gap in the pipeline with a request to fetch data from the L1 cache, the other stages in the pipeline are not affected—e.g., the select unit 334 does not need to interrupt earlier stages to ensure that no data has been lost.Thus, according to one embodiment, the selection unit 334 may wait until a gap in the neighboring processor's pipeline reaches the execution stages 356C before inserting a request from the neighboring queue 332. .
[0056] Additionally or alternatively, processors 301, 350 may be assigned priorities based on the threads they execute. For example, if the system administrator has selected processor 301 to execute the most time-critical threads, processor 301 may be assigned a higher priority than processor 350. This priority may be communicated to selectors 334, 384. When selector 334 receives requests to retrieve data from the L1 cache of processor 350, it may immediately insert the request into the pipeline even if it requires interrupting one or more of the preceding stages 356A through 356B. This ensures that processor 301 receives the data from the adjacent L1 cache with the shortest possible latency.On the other hand, the selection unit 384 can only insert a request into the pipeline of processor 301 if it finds a gap in the pipeline. This ensures that processor 301 remains unaffected by processor 350.
[0057] According to another embodiment, the selection unit 334 may use a ratio to determine when a request should be inserted into the adjacent pipeline. This ratio may be based on the priority of the processors, communicated to the selection unit 334 by the system administrator, or defined in a parameter of a job submitted to the computer system. The ratio may, for example, define a maximum possible number of adjacent requests that can be inserted based on clock cycles—i.e., one inserted request every four clock cycles—or a maximum possible number of adjacent requests that can be inserted for each set of own requests—i.e., one request from processor 301 every five own requests from processor 350. The latter example ignores gaps or breaks within the pipeline.The selection unit 334 uses the ratio to determine when to insert a request into the pipeline of the processor 350.
[0058] Additionally, in some data processing systems with multiple processors on a chip, one or more processors may be disabled such that one or more pipeline execution stages are disabled. Even if a portion of the pipeline of processor 350 is disabled, processor 301 may use the illustrated portion (i.e., execution stages 356C through 356F) to retrieve data from the L1 cache of processor 350.
[0059] In step 410, the selection unit 334 determines whether the criterion(s) are met. The criteria may be one of the criteria discussed above or a combination thereof. Furthermore, this invention is not limited solely to the criteria discussed above.
[0060] The adjacent queue 332 may be organized according to the FIFO (first-in / first-out) principle, with the selection unit 334 applying the criteria during each clock cycle to determine whether the request should be inserted at the beginning of the queue 332.
[0061] After determining that the criteria are met, the select unit 334 may, in step 415, drive the select line of the multiplexer 362 (as well as any other necessary lines) to insert the request into the pipeline of the processor 350. The request is then treated as a separate request originating from an instruction executing in the processor 350, as discussed above.
[0062] However, if the criteria are not met, the selection unit 334 proceeds to store the request in the adjacent queue 332 in step 420. The selection unit 334 may re-evaluate the criteria during each clock cycle or wait a predetermined number of cycles before re-checking whether the criteria are met.
[0063] According to one embodiment, the adjacent queue 332 may include a clock cycle counter to record how long each request has been stored in the queue 332. The selector 334 may use the clock cycle counter to determine whether to continue storing the request in the adjacent queue 332 or forward the request to the L2 queue 330. According to one embodiment, the processor 301 may include a data path (not shown) connecting the selector 334 to the L2 queue 330. When a request is stored in the adjacent queue 332 for a predefined number of clock cycles, the selector 334 may forward the request to the L2 queue 330 instead of waiting for the criteria to be met, allowing the request to be inserted into the pipeline of the processor 350.For example, any request stored in the adjacent queue 332 for ten clock cycles may be forwarded to the L2 queue to retrieve the data from higher levels of the cache hierarchy.
[0064] Additionally, the selection unit 334 may use a different clock cycle threshold depending on the memory position of the requests within the neighboring queue 332. For example, the threshold may be higher for a request at the top of the queue 332, but lower for requests at lower positions within the queue 332. This may prevent a queue 332 from becoming congested, particularly if the selection unit 334 is configured to insert a request only when there is a gap in the neighboring processor's pipeline.
[0065] Furthermore, the selection unit 334 may specify a maximum number of requests allowed in the queue 332 to prevent congestion. Once the maximum number is reached, the selection unit 334 may automatically forward received requests to the L22 queue 330.
[0066] Of course, the criteria for inserting requests can be applied selectively. According to one embodiment, the selection units 334, 384 can insert the request into an adjacent pipeline immediately after it is received. An exemplary data processing system
[0067] The Fig. 5A to 5B are block diagrams illustrating a networked system according to embodiments of the invention for executing jobs submitted by a client in a multi-node system. Fig. 5A shows a block diagram illustrating a networked system for executing jobs submitted by a client in a multi-node system. In the illustrated embodiment, system 500 includes a client system 520 and a multi-node system 570 interconnected by a network 550. Generally, client system 520 submits jobs over network 550 to a file system executing on multi-node system 570. Regardless, any requesting entity may submit jobs to multi-node system 570. For example, jobs may be submitted by software applications (e.g., an instruction executing on client system 520), operating systems, subsystems, other multi-node systems 570, and, at the highest level, by the user. The term "job" refers to a set of commands for requesting resources from multi-node system 570 and utilizing those resources.Any object-oriented programming language such as Java, Smalltalk, C++, or the like may be used to format the set of commands. Additionally, a multi-node system 570 may employ a specialized programming language or provide a specific template. These jobs may be predefined (i.e., hard-coded as part of an instruction) or generated in response to input (e.g., user input). Upon receiving the job, the multi-node system 570 executes the request and then returns the result.
[0068] Fig. 5B is a block diagram of a networked data processing system configured to execute jobs submitted by a client in a multi-node system, according to an embodiment of the invention. It shows that system 500 includes a client system 520 and a multi-node system 570. Client system 520 includes a computer processor 522, storage media 524, memory 528, and a network interface 538. Computer processor 522 may be any processor capable of performing the functions described herein. Client system 520 may connect to network 550 using network interface 538. Furthermore, those skilled in the art will appreciate that any computer system capable of performing the functions described herein may be used.
[0069] In the illustrated embodiment, memory 528 includes an operating system 530 and a client instruction 532. Although memory 528 is shown as a single unit, memory 528 may include one or more storage units having blocks of memory with associated physical addresses, such as random access memory (RAM), read-only memory (ROM), flash memory, or other types of volatile and / or permanent storage. Client instruction 532 is generally capable of generating job requests. Once client instruction 532 generates a job, it may be submitted over network 550 to file system 572 for execution. Operating system 530 may be any operating system capable of performing the functions described herein.
[0070] The multi-node system 570 includes a file system 572 and at least one node 590. Each job file 574 contains data necessary for the nodes 590 to complete a submitted job. The update unit 582 maintains a log of the pending job files, i.e., the job files currently being executed by a node 590. The network interface 584 connects to the network 550 and receives the job files 574 sent by the client system 520. Furthermore, it will be apparent to those skilled in the art that any computer system capable of performing the functions described herein may be used.
[0071] Nodes 590 include a computer processor 592 and a memory 594. Computer processor 592 may be any processor capable of performing the functions described herein. In particular, computer processor 522 may be a plurality of processors as described in Fig. 1. Alternatively, the computer processor 522 may be a multi-core processor having a plurality of processor cores with the structure shown in the processors 110 of Fig. 1. Memory 594 includes an operating system 598. Operating system 598 may be any operating system capable of performing the functions described herein. Memory 594 may include both the cache memory located within processor 592 and one or more memory devices having memory blocks with associated physical addresses, such as random access memory (RAM), read-only memory (ROM), flash memory, or other types of volatile and / or non-volatile memory.
[0072] Fig. Figure 6 illustrates a 4x4x4 ring (torus) 601 of data processing nodes 590, with the inner nodes omitted for clarity. Although Fig. 6 shows a 4x4x4 ring with 64 nodes, it is clear that the actual number of computing nodes in a parallel computing system is typically much larger; for example, a Blue Gene / L system includes 65,536 computing nodes. Each computing node in the circular ring 601 includes a set of six inter-node data links 605A through 605F to allow each computing node in the ring 601 to exchange data with its six immediately neighboring nodes, two nodes each in the x, y, and z coordinate dimensions. According to one embodiment, the parallel computing system 570 may establish a separate ring network for each job executing in the system 570. Alternatively, all computing nodes may be interconnected in a ring.
[0073] The term "ring" as used herein includes any regular pattern of nodes and data transmission paths between the nodes in more than one dimension, such that each node has a defined group of neighboring nodes and it is possible for each individual node to determine the group of neighboring nodes of that node. A "neighboring node" of a particular node is understood to mean any node that is connected to the particular node by a direct data transmission path between the nodes - i.e., a path that does not need to pass through any other node. The data processing nodes can be arranged in a three-dimensional ring 601 according to Fig. 6 interconnected, but can also be configured with more or fewer dimensions. Furthermore, it is not necessary that the neighboring nodes for a given node be the physically closest nodes to the given node, although it is generally desirable to arrange the nodes in this way wherever possible.
[0074] According to one embodiment, the data processing nodes in any dimension x, y, or z form a ring in that dimension because the point-to-point data transmission links are logically arranged in a circular fashion. This is shown, for example, in Fig. 6 by connections 605D, 605E, and 605F, which encircle from a last node in the x, y, and z dimensions to a first node. Thus, although node 610 appears to be a "corner" of the ring, node-to-node connections 605A through 605F connect node 610 to nodes 611, 612, and 613 in the x, y, and z dimensions of ring 601. Conclusion
[0075] Parallel computing environments, where threads executing on adjacent processors can access the same data set, can be designed and configured to share one or more levels of cache. Before a processor forwards a request for data to a higher level of cache after a cache miss, the processor can determine whether the data is stored in a local cache of a neighboring processor. If so, the processor can forward the request to the neighboring processor to retrieve the data. Because cache access is shared between the two processors, the effective size of the cache increases.This can be beneficial in that it reduces the number of cache misses for each level of shared cache without increasing the size of each individual cache on the processor chip.
Claims
[1] Method comprising: Searching a first directory (314) located within a first processor (301) to determine whether a first cache associated with the first processor (301) contains data necessary to execute an instruction executed by the first processor (301), the first directory (314) comprising an index of the data stored in the first cache; Searching a second directory (316) located within the first processor (301) to determine whether a second cache associated with a second processor (350) contains the required data, the second directory (316) comprising an index of the data stored in the second cache; Sending a request from the first processor (301) to the second processor (350) after determining that the data is located in the second cache to retrieve the data from the second cache; Determining whether the request to retrieve the data from the second cache should be inserted into an execution unit of the second processor (350) based on insertion criteria, comprising: Storing the request in a queue after determining that the insertion criteria are not met and therefore the request should not be inserted into the execution unit, and Inserting the request into the execution unit of the second processor (350) after determining that the insertion criteria are met and thereby the request should be inserted into the execution unit, so that the data is retrieved from the second cache memory and sent to an execution unit of the first processor (301), and Sending a request to another memory associated with the first processor (301) after determining that the data is not located in either the first or second cache memory to retrieve the data. [2] The method of claim 1, wherein the first cache memory is an L1 cache memory of the first processor (301) and the second cache memory is an L1 cache memory of the second processor (350), and wherein the other memory is an L2 cache memory of the first processor (301) or the main memory. [3] The method of claim 1, wherein the first processor (301) is unable to write data to the second cache memory. [4] The method of claim 1, wherein the first and second processors (301, 350) are located on one and the same semiconductor chip (100). [5] The method of claim 4, further comprising sending update data to keep the second directory (316) coherent with a local second directory (364) associated with the second processor (350) such that the second directory (316) contains the same index data as the local directory (364), wherein the second processor (350) sends the update data to the first processor (301) after determining that the local directory (364) needs to be updated. [6] The method of claim 1, wherein the steps of searching the first directory (314) and searching the second directory (316) occur simultaneously. [7] Computer program product comprising: a computer-readable storage medium having computer-readable program code stored thereon, wherein the computer-readable program code comprises computer-readable program code configured to: Searching a first directory (314) located within a first processor (301) to determine whether a first cache associated with the first processor (301) contains data necessary to execute an instruction executed by the first processor (301), the first directory (314) comprising an index of the data stored in the first cache; Searching a second directory (316) located within the first processor (301) to determine whether a second cache associated with a second processor (350) contains the required data, the second directory (316) comprising an index of the data stored in the second cache; Sending a request from the first processor (301) to the second processor (350) after determining that the data is in the second cache to retrieve the data from the second cache; and Determining whether the request to retrieve the data from the second cache should be inserted into an execution unit of the second processor (350) based on insertion criteria, comprising: Storing the request in a queue after determining that the insertion criteria are not met and therefore the request should not be inserted into the execution unit, and Inserting the request into the execution unit of the second processor (350) after determining that the insertion criteria are met and thereby the request should be inserted into the execution unit, so that the data is retrieved from the second cache and sent to an execution unit of the first processor (301); and Sending a request to another memory associated with the first processor (301) after determining that the data is not located in either the first or second cache memory to retrieve the data. [8] System that has: a first processor (301) and a second processor (350); a first cache memory associated with the first processor (301); a second cache memory associated with the second processor (350); a first directory (314) in the first processor (301) having an index of the data stored in the first cache, wherein searching the index of the first directory (314) indicates whether the first cache contains data required to execute an instruction executed by the first processor (301); and a second directory (316) in the first processor (301) having an index of the data stored in a second cache of the second processor (350), wherein searching the index of the second directory (316) indicates whether the second cache contains the required data, wherein the first processor (301) is configured to send a request to the second processor (350) after determining that the data is located in the second cache to retrieve the data from the second cache; wherein the first processor (301) is configured to determine whether the request to retrieve the data from the second cache should be inserted into an execution unit of the second processor (350) based on insertion criteria, the first processor (301) being configured to: to store the request in a queue after determining that the insertion criteria are not met and therefore the request should not be inserted into the execution unit, and inserting the request into the execution unit of the second processor (350) after determining that the insertion criteria are met and thereby the request should be inserted into the execution unit, so that the data is retrieved from the second cache memory and sent to an execution unit of the first processor (301); and wherein the first processor (301) is configured to send a request to another memory associated with the first processor (301) after determining that the data is not located in either the first or second cache memory to retrieve the data. [9] The system of claim 8, wherein the first and second processors (301, 350) are located on the same semiconductor chip (100). [10] The system of claim 9, wherein the second directory (316) is located on the same semiconductor chip (100) as the first processor (301), and the system further comprises sending update data to keep the second directory (316) coherent with respect to a local directory (364) associated with the second processor (350) such that the second directory (316) contains the same index data as the local directory (364), the second processor (350) sending the update data to the first processor (301) after determining that the local directory (364) needs to be updated.