Prefetching for a parent core in a multi-core chip
Patent Information
- Application Number
- DE112014000336
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2013-03-05
- Filing Date
- 2014-02-13
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2034-02-13
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates generally to multi-core chips having a parent core and a scout core, and more particularly to prefetching for a parent core in a multi-core chip. BACKGROUND OF THE INVENTION
[0002] The performance increase of single-threaded processors is limited by the power consumption of single-threaded operation. Doubling the power consumption of a processor due to higher frequencies and / or features does not necessarily deliver a performance increase that exceeds or equals the increased power consumption. This is due to a significant deterioration in the ratio of performance improvement to the increase in power consumption. To increase chip performance, significant portions of the available power can be used for additional cores on a chip.Sharing caches and main memory prevents the performance improvement from being equal to the relative increase in the number of cores, but on the other hand, the performance improvement by increasing the number of cores on the chip can result in a greater performance improvement / energy savings than simply increasing the performance of a single-core processor.
[0003] One approach to improving single-threaded performance involves using a second core on the same chip as a primary or parent core as a scout core. Specifically, the scout core can be used to prefetch data from a shared cache into the private cache of a parent core. This approach is particularly beneficial when a cache miss occurs in the parent core. A cache miss occurs when searching for a given line of data requires searching a directory of the parent core and the requested cache line is not present. A typical approach to finding the missing cache line is to trigger a fetch operation at a higher level of the cache. The scout core provides a mechanism used to prefetch data required by the parent core.
[0004] It should be noted that different programs behave differently, so a prefetching algorithm or approach may not always improve the latency for accessing cache contents. According to one data prefetching approach for the parent core, a relatively small and simple algorithm, which is a step counting routine, can be provided to speculatively prefetch data based on a step size observed between consecutive cache misses. To capture more complex patterns, additional hardware is required, which may be more complex, larger in size, and more power-consuming. However, given the throughput, latency, and power trade-offs for the chip, the amount of dedicated prefetching hardware can be limited to the single core.In addition, the area and memory required to monitor and detect cache misses may be too large to use hardware alone.
[0005] US 2004 / 0 148 491 A1 describes a sideband scout thread processing technique. The sideband scout thread processing technique uses sideband information to identify a subset of processor instructions for execution by a scout thread processor. The sideband information identifies instructions that must be executed to "warm up" a cache memory shared with a main processor that executes the full set of processor instructions. In this way, the main processor has fewer cache misses and lower latency. In one example, a system includes a first processor for executing a sequence of processor instructions, a second processor for executing a subset of the sequence of processor instructions, and a cache shared by the first processor and the second processor.The second processor includes a sideband circuit configured to identify the subset of the sequence of processor instructions to be executed according to sideband information associated with the sequence of processor instructions.
[0006] US 2011 / 0 296 431 A1 describes a method and system that can enable fast, hardware-assisted communication of values between threads in a producer-consumer style. In one example, the method uses a dedicated hardware buffer as a temporary storage for transferring values from registers in one thread to registers in another thread. The method can provide a generic, programmable solution that can transfer any subset of register values between threads in any order, where the source and destination registers may or may not be correlated. The method can also enable fixed access times because it completely bypasses the memory hierarchy. Furthermore, the method is designed to be lightweight and focused on communication, with synchronization capabilities remaining orthogonal to the communication mechanism.For example, it can be applied by a helper thread that performs data prefetching for an application thread to initialize open-ended reads in the address calculation slice of the helper thread's code. SUMMARY
[0007] The objects underlying the invention are achieved by the features of the independent patent claims. Embodiments of the invention are the subject of the dependent patent claims.
[0008] Aspects of the invention include a method, a system, and a computer program product for prefetching data on a chip having at least one scout core, at least one parent core, and a shared cache memory shared equally between the at least one scout core and the at least one parent core. Prefetch code is executed by the scout core to monitor the parent core. The prefetch code executes independently of the parent core. The scout core determines, based on the monitoring of the parent core, that a predetermined data pattern has occurred in the parent core. A prefetch request is sent from the scout core to the shared cache memory. The prefetch request is sent based on the at least one predetermined pattern detected by the scout core.A record indicated by the prefetch request is sent by the Scout core to the parent core. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Embodiments of the invention will now be described, by way of example only, with reference to the accompanying drawings, in which: Fig. 1 illustrates multi-core chips according to an embodiment; Fig. 2 illustrates the central processing (CP) chip according to one embodiment; Fig. 3 illustrates a CP chip according to another embodiment; Fig. 4 illustrates a CP chip according to yet another embodiment; Fig. 5 is a flowchart illustrating a method for prefetching data for a parent core by a scout core according to one embodiment; and Fig. 6 illustrates a computer program product according to an embodiment. DETAILED DESCRIPTION
[0010] An embodiment for prefetching data for a parent core by a scout core in a multi-core chip is disclosed. According to an exemplary embodiment, the multi-core chip includes at least one parent core, at least one scout core, and a shared cache. The scout core monitors the activity of the parent core for at least one type of predetermined pattern generated by the parent core and determines whether to send a prefetch request from the scout core to the shared cache. Upon receiving the prefetch request from the scout core, the data requested by the prefetch request is sent to the parent core. The data requested by the prefetch request is not accepted by the scout core, but only by the parent core.The Scout core monitors the parent core for various types of predefined data patterns found in the parent core. However, some types of dedicated hardware prefetchers currently available are only capable of monitoring the parent node for a predefined subset of data patterns. Furthermore, due to the amount of hardware shared by the Scout core prefetcher, the Scout processor is capable of analyzing more data than a typical hardware prefetcher.
[0011] Fig. 1 illustrates an example of a data processing system 10 according to one embodiment. The data processing system 10 includes at least one central processing (CP) chip 20. In the Fig. In the exemplary embodiment shown in Figure 1, three CP chips 20 are illustrated, however, it should be understood that any number of CP chips 20 may be used. Each CP chip 20 communicates with a shared cache memory 22 and a system memory 24.
[0012] According to the Fig. 1 to 2, each CP chip 20 contains several cores 30 for reading and executing instructions. Fig. For example, in the exemplary embodiment shown in Figure 2, each CP chip 20 includes a master core 32 and a scout core 34, but it is understood that any number of cores 30 may be used as well, and in the Fig. 3 to 4 illustrate alternative embodiments of the CP chip. According to Fig. 2, each core 30 also includes a corresponding I-cache 40 and a D-cache 42. According to the Fig. 2, each of the cores 30 includes only a Level 1 (L1) cache, however, it is understood that in various embodiments, the cores 30 may also include a Level 2 (L2) cache. Each core 30 is operatively connected to a shared cache 50. In the embodiment shown in Fig. 2, the shared cache 50 is an L2 cache, but it is understood that the shared cache 50 may equally well be a level 3 (L3) cache.
[0013] A data return bus 60 is provided between the parent core 32 and the shared cache 50, and a data return bus 62 is provided between the scout core 34 and the shared cache 50. The parent core 32 is connected to the shared cache 50 by a fetch request bus 64, over which data is sent from the parent core 32 to the shared cache 50. The scout core 34 is connected to the shared cache 50 by a fetch monitor bus 66, over which the scout core 34 monitors the shared cache 50. A fetch request bus 68 is arranged between the scout core 34 and the shared cache 50 to send various fetch requests from the scout core 34 to the shared cache 50. The fetch request bus 68 can also be used for typical fetch operations like the fetch request bus 64.Such a fetch may be necessary to load prefetch code into the Scout core 34 if additional data needs to be loaded for analysis, if the data to be analyzed does not fit entirely into the local data cache 42 and / or if the prefetch code does not fit entirely into the local instruction memory 40. .
[0014] At the Fig. 2, the shared cache 50 serves as a hub or link so that the scout core 34 can monitor the parent core 32. The scout core 34 monitors the parent core 32 for at least one predetermined data pattern occurring in the parent core 32. More specifically, the scout core 34 executes prefetch code used to monitor the parent core 32. The prefetch code determines whether one or more predetermined data patterns have occurred in the parent core 32 and sends a fetch request to the shared cache 50 based on the determined data pattern. Furthermore, the prefetch code executes independently of any code executed by the parent core 32. The Scout core 34 generally stores the prefetch code in the L1 I-cache 40 located in the Scout core 34.
[0015] The predetermined data pattern may be a content request leaving the parent core 32 (e.g., a request for a predetermined cache line that is not present in the I-cache 40 and a D-cache 42 of the parent core 32), or alternatively, a checkpoint address of the parent core 32. For example, the parent core 32 may fetch a memory address from either the I-cache 40 or the D-cache 42. If the I-cache 40 or the D-cache 42 does not contain a predetermined cache line requested by the parent core 32, a cache miss has occurred. The scout core 34 detects the cache miss by monitoring the parent core 32 through the shared cache 50 via the fetch monitor bus 66.According to one embodiment, scout core 34 determines whether the cache miss occurred in I-cache 40 or D-cache 42 (or any other type of cache in parent core 32 where a cache miss occurs). Upon detecting a cache miss, a prefetch of a future missing cache line may be sent by scout core 34 to shared cache 50 via fetch request bus 68. According to one approach, scout core 34 may perform a check to determine whether the cache line in question is stored in the cache of parent core 32 (e.g., I-cache 40 or D-cache 42). If the cache line in question is present in the parent core 32, the data already in the cache of the parent core 32 no longer needs to be accessed in advance.
[0016] According to another approach, the checkpoint address of the parent core 32 may be communicated between the parent core 32 and the scout core 34 via the shared cache 50. Certain checkpoint addresses may be representative of predetermined events. The predetermined event may be, for example, a cleanup function or a context switch. According to an exemplary embodiment, the checkpoint address may correspond to a predetermined cache line in either the I-cache 40 or the D-cache 42 of the parent core 32; however, it should be understood that the checkpoint address may not necessarily correspond to a predetermined prefetch address.The scout core 34 monitors the parent core 32, and after the specified event completes, the scout core 34 sends a prefetch request to the shared cache 50 to acquire a cache line associated with the specified event.
[0017] After receiving the prefetch request from the scout core 34, the shared cache 50 sends the data requested by the prefetch request to the parent core 32 via the data return bus 60. The shared cache 50 sends the data requested by the prefetch request to the parent core 32 as a function of the prefetch request. The data requested by the prefetch request is not accepted by the scout core 34, but only by the parent core 32.
[0018] According to one approach, the scout core 34 notifies the parent core 32 that a prefetch has been made on behalf of the parent core 32. Alternatively, according to another approach, after sending the data requested by the prefetch request, the shared cache 50 also notifies the parent core 32 that a prefetch has been made on behalf of the parent core 32. Thus, the scout core 34 tells the shared cache 50 how to forward the data requested by the prefetch request and store it on the parent core 32, just as if the parent core 32 itself had initiated the prefetch request (although the scout core 34, not the parent core 32 itself, initiated the prefetch request).Thus, the data requested by the prefetch request is stored in the I-cache 40 or the D-cache 42 of the parent core 32.
[0019] Fig. Figure 3 is an alternative representation of a CP chip 120 with a single scout core 134, but at least two parent cores 132. Note that in Fig. 3, two higher-level cores 132 are shown, but any number of higher-level cores 132 can equally be used. In the case of Fig. In the embodiment shown in Figure 3, a data return bus 160 is provided between the parent cores 132 and the shared cache 150, and a data return bus 162 is provided between the scout core 134 and the shared cache 150. A fetch request bus 164 is provided for each of the parent cores 132, with the fetch request bus 164 connecting the parent cores 132 to the shared cache 150. A fetch monitor bus 166 connects the scout core 134 to the shared cache 150. A fetch request bus 168 is arranged between the scout core 134 and the shared cache 150 for sending various prefetch requests from the scout core 134 to the shared cache 150.
[0020] Fig. 4 is an alternative representation of a CP chip 224 with at least two scout cores 234 and one parent core 232. Note that in Fig. 4, two Scout cores 234 are shown, but equally several (e.g. more than two) Scout cores 232 can be used. In the Fig. In the embodiment shown in Figure 4, a data return bus 260 is provided between the parent core 232 and the shared cache 250. A data return bus 262 is provided for each of the scout cores 234, which serves to connect one of the scout cores 234 to the shared cache 250. A fetch request bus 264 connects the parent core 232 to the shared cache 250. A fetch monitor bus 266 is provided for each of the scout cores 234, which serves to connect one of the scout cores 234 to the shared cache 250. A fetch request bus 268 is provided for each of the scout cores 234, which serves to connect one of the scout cores 234 to the shared cache 250.
[0021] According to the Fig. 4, each of the scout cores 234 may monitor the parent core 232 for a different predetermined data pattern. For example, according to one approach, one of the scout cores 234 may monitor and analyze the behavior of an L1 I cache 240 of the parent core 232, while the other scout core 234 may monitor and analyze the behavior of an L1 D cache 242 of the parent core 232. Thus, additional data may be monitored and analyzed within a given period of time.
[0022] Fig. 5 is a flowchart of a method 300 for prefetching data for the parent core 32 by the scout core 34, which will be discussed below. Referring to the Fig. 1 to 5, the method 300 starts in block 302, where the scout core 34 monitors the parent core 32 via the shared cache 50. The method may then proceed to block 304.
[0023] At block 304, the scout core 34 monitors the parent core 32 for the predetermined data pattern present in the parent core 32. As discussed above, the predetermined data pattern may be either a content request leaving the parent core 32 (e.g., a request for a predetermined cache line that is not present in either the I-cache 40 or the D-cache 42 of the parent core 32) or, alternatively, a checkpoint address. If the predetermined data pattern is not detected, the method 300 may return to block 302.
[0024] If the predetermined data pattern is recognized, the method may continue to block 306.
[0025] At block 306, the scout core 34 sends the prefetch request to the shared cache 50. For example, as discussed above, the prefetch request may be a prefetch of the missing cache line sent by the scout core 34 to the shared cache 50. Then, the method 300 may proceed to block 308.
[0026] In block 308, the parent core 32 is notified that a prefetch has been performed on behalf of the parent core 32. The method 300 may then proceed to block 310.
[0027] In block 310, shared cache 50 sends the data requested by the prefetch request to parent core 32 via data return bus 60. Shared cache 50 sends the data requested by the prefetch request to parent core 32 as a function of the prefetch request. According to one embodiment, blocks 308 and 310 are executed concurrently. Method 300 may then exit.
[0028] Those skilled in the art will appreciate that one or more aspects of the present invention may be implemented as a system, method, or computer program product. Accordingly, one or more aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which may be generally referred to herein as a "circuit," "module," or "system." In addition, one or more aspects of the present invention may take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code stored thereon.
[0029] Any combination of one or more computer-readable media may be used. The computer-readable medium may be a computer-readable storage medium. A computer-readable storage medium may be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or unit, or any suitable combination thereof.Individual examples (a non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer diskette, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the context of this document, a computer-readable storage medium can be any tangible medium that can contain or store a program for use by or in connection with a system, apparatus, or device for executing instructions.
[0030] According to a Fig.6, a computer program product 600 includes, for example, one or more storage media 602, where the media may be tangible and / or non-transitory, for storing computer-readable program code means or program code logic 604 thereon to provide or enable one or more aspects of embodiments described herein.
[0031] Program code generated and stored on a tangible medium (including, but not limited to, electronic memory modules (RAM), flash memory, compact discs (CDs), DVDs, magnetic tapes, and the like) is often referred to as a "computer program product." The computer program product medium is typically readable by processing circuitry, preferably in a computer system for execution of the computer program product by the processing circuitry. Such program code, for example, may be generated using a compiler or assembler to compile instructions that, when executed, implement aspects of the invention.
[0032] The technical effects and advantages of the data processing system 10 described above include developing a program that can be executed by the L1 I cache 40 of the scout core 34. The scout core 34 can monitor the parent core 32 for various types of predetermined data patterns that occur in the parent core 32. Certain types of currently available hardware prefetchers, however, can only monitor the parent core 34 for one predetermined pattern. Furthermore, the number of data patterns that can be monitored and analyzed by the scout processor 34 can be relatively larger than that of a currently available hardware prefetcher because the entire L1 D cache 42 of the scout processor 34 can be used to store the data to be analyzed.
[0033] Computer program code for performing operations for aspects of the embodiments may be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, Smalltalk, C++, or the like, and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code may execute entirely on a user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server.In the latter case, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, over the Internet using an Internet service provider).
[0034] Aspects of embodiments are described above with reference to flowchart and / or schematic diagrams of methods, apparatus (systems), and computer program products according to the embodiments. It will be appreciated that each block of the flowchart and / or block diagrams, and combinations of blocks in the flowchart and / or block diagrams, may be implemented by computer program instructions. These computer program instructions may be supplied to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions executed by the processor of the computer or other programmable data processing apparatus produce means for implementing the functions / acts specified in the block or blocks of the flowchart and / or block diagrams.
[0035] These computer program instructions may also be stored in a computer-readable medium that can cause a computer, other programmable data processing apparatus, or other devices to function in a particular manner such that the instructions stored in the computer-readable medium produce an article of manufacture that includes instructions for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0036] The instructions of the computer program may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a sequence of operations to be performed on the computer, other programmable device, or other devices to produce a computer-based process such that the instructions executing on the computer or other programmable device provide processes for implementing the functions / acts specified in the block or blocks of the flowchart and / or block diagrams.
[0037] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments. Accordingly, each block in the flowchart or block diagrams may represent a module, segment, or portion of code comprising one or more executable instructions for implementing the specified logical function(s). It should be noted that in some alternative implementations, the functions specified in the block may occur in a different order than that shown in the figures. For example, two blocks shown in sequence may actually execute substantially concurrently, depending on the intended functionality, or the blocks may sometimes execute in the reverse order.It should also be noted that each block of the block diagrams and / or flowchart, or combinations of blocks in the block diagrams and / or flowchart, may be implemented by special purpose hardware systems that perform the specified functions or actions, or by combinations of special purpose hardware and computer instructions.
Claims
[1] A computer system (10) for pre-reading data, the system (10) comprising: a chip (20, 120, 224) comprising: at least one scout core (34, 134, 234) located on the chip (20, 120, 224); at least one superordinate core (32, 132, 232) located on the chip (20, 120, 224); and a shared cache memory (50, 150, 250) shared equally between the at least one scout core (34, 134, 234) and the at least one parent core (32, 132, 232), wherein the shared cache memory (50, 150, 250) is located on the chip (20, 120, 224) and the system (10) is configured to perform a method (300), the method (300) comprising: Executing prefetch code by the at least one scout core (34, 134, 234) to monitor (302) the at least one parent core (32, 132, 232), the prefetch code being executed independently of the at least one parent core (32, 132, 232); Determining (304) by the at least one scout core (34, 134, 234) based on the monitoring of the at least one parent core (32, 132, 232) that at least one predetermined data pattern has occurred in the at least one parent core (32, 132, 232); Sending (306) a prefetch request from the at least one scout core (34, 134, 234) to the shared cache (50, 150, 250), the sending being based on the detection; and Sending (310) a data set indicated by the prefetch request through the shared cache memory (50, 150, 250) to the at least one parent core (32, 132, 232). [2] The computer system (10) of claim 1, further comprising notifying (308) the at least one parent core (32, 132, 232) that the prefetch request was made on behalf of the at least one parent core (32, 132, 232). [3] The computer system (10) of claim 1 or 2, wherein the at least one scout core (34, 134, 234) tells the shared cache (50, 150, 250) how to forward the data requested by the prefetch request and store it in a cache located in the at least one parent core (32, 132, 232). [4] Computer system (10) according to one of the preceding claims, wherein the chip (120) contains at least two higher-level cores (132), each of which exchanges data with the shared cache memory (150). [5] The computer system (10) of any preceding claim, wherein the chip (224) includes at least two scout cores (234) that exchange data with the shared cache memory (250), and wherein the scout cores (234) each monitor the at least one parent core (232) for a different predetermined data pattern. [6] The computer system (10) of any preceding claim, wherein the at least one scout core (34, 134, 234) monitors the at least one parent core (32, 132, 232) via a fetch monitor bus (66, 166, 266), the fetch monitor bus (66, 166, 266) connecting the at least one scout core (34, 134, 234) to the shared cache memory (50, 150, 250). [7] The computer system (10) of any preceding claim, wherein the at least predetermined data pattern is a cache miss occurring in a cache located in the at least one higher-level core (32, 132, 232). [8] The computer system (10) of any one of claims 1 to 6, wherein the at least one predetermined data pattern is a checkpoint address of the at least one higher-level core (32, 132, 232). [9] A computer program product (600) for prefetching data on a chip (20, 120, 224) having at least one scout core (34, 134, 234), at least one higher-level core (32, 132, 232), and a shared cache memory (50, 150, 250) shared equally between the at least one scout core (34, 134, 234) and the at least one higher-level core (32, 132, 232), the computer program product (600) comprising: a tangible storage medium (602) readable by a processing circuit and having stored therein instructions (604) for execution by the processing circuit to perform a method (300) comprising: Executing a prefetch code by the at least one scout core (34, 134, 234) to monitor (302) the at least one parent core (32, 132, 232), wherein the prefetch code is executed independently of the at least one parent core (32, 132, 232). Determining (304) by the at least one scout core (34, 134, 234) based on the monitoring of the at least one parent core (32, 132, 232) that at least one predetermined data pattern has occurred in the at least one parent core (32, 132, 232); Sending (306) a prefetch request from the at least one scout core (34, 134, 234) to the shared cache (50, 150, 250), the sending being based on the detection; and Sending (310) a data set indicated by the prefetch request through the shared cache memory (50, 150, 250) to the at least one parent core (32, 132, 232). [10] The computer program product (600) of claim 9, further comprising notifying (308) the at least one parent core (32, 132, 232) that the prefetch request was made on behalf of the at least one parent core (32, 132, 232). [11] The computer program product (600) of claim 9 or 10, wherein the at least one scout core (34, 134, 234) tells the shared cache (50, 150, 250) how to forward the data requested by the prefetch request and store it in a cache located in the at least one parent core (32, 132, 232). [12] The computer program product (600) of any one of claims 9 to 11, wherein the chip (120) includes at least two higher-level cores (132) each exchanging data with the shared cache memory (150). [13] The computer program product (600) of any one of claims 9 to 12, wherein the chip (224) includes at least two scout cores (234) that exchange data with the shared cache memory (250), and wherein the scout cores (234) each monitor the at least one parent core (232) for a different predetermined data pattern. [14] The computer program product (600) of any one of claims 9 to 13, wherein the at least one scout core (34, 134, 234) monitors the at least one parent core (32, 132, 232) via a fetch monitor bus (66, 166, 266), the fetch monitor bus (66, 166, 266) connecting the at least one scout core (34, 134, 234) to the shared cache memory (50, 150, 250). [15] The computer program product (600) of any one of claims 9 to 14, wherein the at least one predetermined data pattern is a cache miss occurring in a cache located in the at least one parent core (32, 132, 232). [16] A computer-assisted method (300) for prefetching data on a chip (20, 120, 224) having at least one scout core (34, 134, 234), at least one higher-level core (32, 132, 232), and a shared cache memory (50, 150, 250) shared equally between the at least one scout core (34, 134, 234) and the at least one higher-level core (32, 132, 232), the method (300) comprising: Executing prefetch code by the at least one scout core (34, 134, 234) to monitor (302) the at least one parent core (32, 132, 232), the prefetch code being executed independently of the at least one parent core (32, 132, 232); Determining (304) by the at least one scout core (34, 134, 234) based on the monitoring of the at least one parent core (32, 132, 232) that the at least one predetermined data pattern has occurred in the at least one parent core (32, 132, 232); Sending (306) a prefetch request from the at least one scout core (34, 134, 234) to the shared cache (50, 150, 250), the sending being based on the detection; and Sending (310) a data set indicated by the prefetch request through the shared cache memory (50, 150, 250) to the at least one parent core (32, 132, 232). [17] The method (300) of claim 16, further comprising notifying (308) the at least one parent core (32, 132, 232) that the prefetch request was made on behalf of the at least one parent core (32, 132, 232). [18] The method (300) of claim 16 or 17, wherein the at least one scout core (34, 134, 234) informs the shared cache (50, 150, 250) how the data requested by the prefetch request should be forwarded and stored in a cache located in the at least one parent core (32, 132, 232). [19] The method (300) of any one of claims 16 to 18, wherein the chip (120) includes at least two higher-level cores (132) that exchange data with the shared cache memory (150). [20] The method (300) of any one of claims 16 to 19, wherein the chip (224) includes at least two scout cores (234) that exchange data with the shared cache memory (250), and wherein the scout cores (234) each monitor the at least one parent core (232) for a different predetermined data pattern.
Citation Information
Patent Citations
Sideband scout thread processor
US20040148491A1
Method and apparatus for efficient helper thread state initialization using inter-thread register copy
US20110296431A1