Prefetching data for a chip with a parent core and a scout core
Patent Information
- Application Number
- DE112014000340
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2013-03-05
- Filing Date
- 2014-02-12
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2034-02-12
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates generally to multi-core chips having a parent core and a scout core, and more particularly to special prefetch algorithms for a parent core in a multi-core chip. BACKGROUND OF THE INVENTION
[0002] Multiple cores can be arranged on a single chip. According to one approach, a second core can be deployed on the same chip as the parent core as a scout core. To deploy or leverage the existing scout core according to one approach, the scout core is used to prefetch data from a shared cache into the private cache of the parent core. This approach is particularly advantageous when a cache miss occurs in the parent core. A cache miss occurs when searching for a specific line of data requires searching a directory of the parent core and the requested cache line does not exist. A typical approach to finding the missing cache line is to trigger a fetch operation at a higher level of the cache.The Scout core provides a mechanism used to prefetch data required by the parent core.
[0003] It should be noted that different applications behave differently, so a prefetching algorithm or approach may not always improve the latency for accessing cache contents. In particular, if the parent kernel is running several different applications, for example, the prefetching algorithm used to monitor the different applications may provide different latencies for accessing cache contents depending on the individual application currently running. For example, an application for searching a sparsely populated database may behave differently (e.g., the prefetching algorithm may provide a longer or shorter latency for accessing cache contents) than an application for color correction of images.
[0004] US 2004 / 0 148 491 A1 describes a sideband scout thread processing technique. The sideband scout thread processing technique uses sideband information to identify a subset of processor instructions for execution by a scout thread processor. The sideband information identifies instructions that must be executed to "warm up" a cache memory shared with a main processor that executes the full set of processor instructions. In this way, the main processor has fewer cache misses and lower latency. In one example, a system includes a first processor for executing a sequence of processor instructions, a second processor for executing a subset of the sequence of processor instructions, and a cache shared by the first processor and the second processor.The second processor includes a sideband circuit configured to identify the subset of the sequence of processor instructions to be executed according to sideband information associated with the sequence of processor instructions.
[0005] US 2011 / 0 296 431 A1 describes a method and system that can enable fast, hardware-assisted communication of values between threads in a producer-consumer style. In one example, the method uses a dedicated hardware buffer as a temporary storage for transferring values from registers in one thread to registers in another thread. The method can provide a generic, programmable solution that can transfer any subset of register values between threads in any order, where the source and destination registers may or may not be correlated. The method can also enable fixed access times because it completely bypasses the memory hierarchy. Furthermore, the method is designed to be lightweight and focused on communication, with synchronization capabilities remaining orthogonal to the communication mechanism.For example, it can be applied by a helper thread that performs data prefetching for an application thread to initialize open-ended reads in the address calculation slice of the helper thread's code. SUMMARY
[0006] The objects underlying the invention are achieved by the features of the independent patent claims. Embodiments of the invention are the subject of the dependent patent claims.
[0007] Aspects of the present invention relate to a method, a system, and a computer program product for prefetching data on a chip having at least one scout core and a parent core. The method includes the parent core storing the starting address of a prefetch code. The starting address of the prefetch code indicates where the prefetch code is stored. The prefetch code is specifically configured to monitor the parent core based on a predetermined application executed by the parent core. The method includes the parent core sending a broadcast interrupt signal to the at least one scout core. The broadcast interrupt signal is sent based on the stored starting address of the prefetch code.The method includes monitoring the parent core with prefetch code executed by at least one scout core. The scout core executes the prefetch code based on receipt of the broadcast interrupt signal. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Embodiments of the invention will now be described, by way of example only, with reference to the accompanying drawings, in which: Fig. 1 illustrates multi-core chips according to an embodiment; Fig. 2A illustrates a central processing (CP) chip according to one embodiment; Fig. 2B illustrates a CP chip according to an alternative embodiment; Fig. 3 represents an architecturally defined preload instruction of the Scout core; Fig. 4 is a flowchart illustrating an example method for loading prefetch code from a parent core to a scout core; Fig. 5 is a flowchart illustrating an example method for loading prefetch code during a task switch involving a parent kernel operating system; Fig. 6 is a flowchart illustrating an exemplary method for loading prefetch code by a given application executing in the parent core; and Fig. 7 illustrates a computer program product according to an embodiment. DETAILED DESCRIPTION
[0009] An embodiment for prefetching data by a scout core in a multi-core chip with more efficient prefetching is disclosed. According to an exemplary embodiment, the multi-core chip includes a parent chip and at least one scout core. A starting address of a prefetch code is stored in the parent chip. The starting address of the prefetch code indicates where a particular prefetch code is stored. The prefetch code is specifically configured to monitor the parent core based on a predetermined application executed by the parent core. The scout core monitors the parent core by executing the predetermined prefetch code. The predetermined prefetch code may correspond to a predetermined application selectively executed by the parent core (e.g.,If an application does not have a predefined prefetch code associated with it, the Scout core may instead execute de facto or default prefetch code. Note that different applications executed by the parent core behave differently, and thus a general prefetch algorithm (e.g., a prefetch algorithm not tailored to a specific application) may not always improve latency depending on the predefined application executed by the parent core. The approach disclosed in example embodiments enables the Scout core to monitor the parent core using predefined prefetch code that is precisely tailored to monitor a predefined application executed by the parent core.The parent core may send a broadcast interrupt signal to the scout core to switch the prefetch code based on the specified application executed by the parent core.
[0010] Fig. 1 illustrates an example of a data processing system 10 according to one embodiment. The data processing system 10 includes at least one central processing (CP) chip 20. In the Fig. In the exemplary embodiment shown in Figure 1, three CP chips 20 are illustrated, however, it should be understood that any number of CP chips 20 may be used. According to one approach, data processing system 10 may include, for example, eight CP chips 20. According to another approach, data processing system 10 may include up to twelve or sixteen CP chips 20. Each CP chip 20 communicates with a shared cache memory 22 and system memory 24.
[0011] According to the Fig. 1 to 2A, each CP chip 20 contains several cores 30 for reading and executing instructions. Fig. 2A, each CP chip 20 includes a parent core 32 and a scout core 34, however, it is understood that multiple parent cores 32 and scout cores 34 may be accommodated on the CP chip 20. According to one approach, the CP chip 20 may, for example, include four parent cores 32 that each exchange data with a scout core 34 (i.e., a total of eight cores). According to an embodiment shown in Fig. In the alternative embodiment shown in Figure 2B, which illustrates a CP chip 120, a parent core 132 may exchange data with multiple scout cores 134. According to one approach, the CP chip 120 may, for example, be equipped with two parent cores 132, each of which exchanges data with three scout cores 134 (i.e., a total of eight cores).
[0012] Each core 30 in Fig. 2A also includes a corresponding instruction I cache 40 and a data D cache 42. In the Fig. 2A, the cores 30 each include only a Level 1 (L1) cache, however, it is understood that, according to various embodiments, the cores 30 may also include a Level 2 (L2) cache. Each core 30 is operatively connected to a shared cache 50. In the embodiment shown in Fig. In the embodiment shown in Figure 2A, the shared cache is an L2 cache, but it is understood that the shared cache 50 may also be a level 3 (L3) cache.
[0013] A data return bus 60 is provided between the parent core 32 and the shared cache 50, and a data return bus 62 is provided between the scout core 34 and the shared cache 50. A fetch request bus 64 connects the parent core 32 to the shared cache 50 and the scout core 34, with data being sent from the parent core 32 to the shared cache 50 and the scout core 34. A fetch request bus 66 connects the scout core 34 to the shared cache 50, with the scout core 34 monitoring the shared cache 50 via the fetch request bus 66. The fetch request bus 66 can also be used for fetching for the scout core 34. This behavior is similar to the fetch request bus 64, which fetches for the parent core 32.Such a fetch may be necessary to load one or more prefetch algorithms into the scout core 34 and possibly to load additional data to be analyzed if the data to be analyzed does not fit entirely within the local D-cache 40. A load prefetch bus 68 is disposed between the parent core 32 and the scout core 34. The parent core 32 tells the scout core 32 via the load prefetch bus 68 to load a prefetch code and a predetermined prefetch code starting address that indicates where the prefetch code is stored. The prefetch code may be stored in a variety of different memory locations in the data processing system 10 that may be accessed via the memory address, such as the Scout core's L1 I cache 40, the shared cache 50, the shared cache 22 (. Fig. 1) or the system memory 24 ( Fig. 1).
[0014] According to Fig. 2B, a data return bus 160 is arranged between the parent core 132 and a shared cache 150, and a data return bus 162 is arranged between the scout cores 134 and the shared cache 150. A fetch request bus 164 connects the parent core 132 to the shared cache 150, with data being sent from the parent core 132 to the shared cache 150. A fetch request bus 166 is provided for each scout core 134, connecting the scout core 134 to the shared cache 150. The data transferred over the fetch request buses 166 is different for each of the scout cores 134. Each Scout core 134 has a load prefetch bus 168 connected thereto, which is disposed between the parent core 132 and each Scout core 134.For each Scout core 134, a fetch monitor bus 170 is provided, which is arranged between the shared cache memory 150 and one of the Scout cores 134. Unlike the fetch request bus 166, the data transferred over the fetch monitor buses 170 is not necessarily different for each of the Scout cores 134.
[0015] According to Fig. 2A, the shared cache 50 serves as a node or link so that the scout core 34 can monitor the parent core 32. The scout core 34 monitors the parent core 32 for a predetermined data pattern occurring in the parent core 32. In particular, the scout core 34 executes prefetch code used to monitor the parent core 32. The prefetch code determines whether one or more predetermined data patterns have occurred in the parent core 32 and sends a fetch request to the shared cache 50 based on the predetermined data pattern. The scout core 34 generally stores the prefetch code in the I-cache 40 located on the scout core 34.
[0016] The specified data pattern can be either a content request leaving the parent core 32 (e.g., a request for a specific cache line that is not located in the I-cache 40 or a D-cache 42 of the parent core 32) or a checkpoint address of the parent core 32. For example, if the specified data pattern is a cache miss (e.g., a missing cache line in the I-cache 40 or the D-cache 42 of the parent core 32), the scout core 34 can send a prefetch request for a sought-after missing cache line to the shared cache 50 via the fetch request bus 66.If the given data pattern is the checkpoint address of the parent core 32, the scout core 34 monitors the parent core 32 and, upon completion of a given event (e.g., a cleanup function or a context switch), sends a prefetch request to the shared cache to acquire a cache line associated with the given event.
[0017] The scout core 34 is configured to selectively execute a predetermined prefetch code based on the predetermined application executed by the parent core 32. For example, when the parent core 32 executes an application "A", an application "B", and an application "C" sequentially (e.g., the parent core 32 first executes application "A", then application "B", and then application "C"), the scout core 34 may execute a prefetch code "A" to monitor application "A", a prefetch code "B" to monitor application "B", and a prefetch code "C" to monitor application "C". That is, the predetermined prefetch code is specifically configured to monitor the parent core 32 while it executes a corresponding application (e.g., the prefetch code "A" monitors application "A").This is because default prefetch codes may behave differently depending on the default application executed by the parent core 32. For example, an application designed to search a sparsely populated database may behave differently (e.g., the prefetch algorithm may provide a longer or shorter latency for accessing cache contents) than an application designed to color correct images. Note that if an application does not have default prefetch code associated with it, the Scout core 34 may instead execute de facto or default prefetch code. Architecturally defined state within the Scout core 34 provides the memory location where the default prefetch code for the Scout core 34 is stored.
[0018] The parent core 32 generally executes various applications sequentially (i.e., one application at a time) and switches from one application to another at a relatively high speed (e.g., up to a speed of several hundred times per second), which is referred to as multitasking. More specifically, the parent core 32 can execute a number of applications. When a given application passes control through the parent core 32 to another application, this is referred to as a task swap. While an operating system of the parent core 32 is affected by a task switch, the parent core 32 stores a (in Fig. 2B, prefetch address 48 within parent core 32 associated with the currently executing application. Prefetch address 48 indicates where the starting address of the prefetch code associated with the currently executing application has been stored. Parent core 32 may then load a new application and update prefetch address 48 with a starting address of the prefetch code associated with the new application. Parent core 32 may then send a broadcast interrupt signal to scout core 34 via load prefetch bus 68. The broadcast interrupt signal provides the starting address of the prefetch code and an interrupt notification.If the new application loaded by the parent core 32 does not have a default prefetch code associated with it, the parent core 32 sends the broadcast interrupt signal to the scout core 34, indicating that the scout core 34 should load the default prefetch code.
[0019] In addition to the task switch by the parent core 32, the broadcast interrupt signal may also be triggered by the specified application executing in the parent core 32. That is, the specified application executing in the parent core 32 may issue an instruction indicating that the scout core 34 should load a specified prefetch code. According to Fig. 3, an architected scout prefetch instruction 70 may be issued by the specified application executing in the parent core 32. The architected scout prefetch instruction 70 indicates the starting address of the prefetch code associated with the specified application executing in the parent core 32. For example, if the parent core 32 is executing an application "A," the architected scout prefetch instruction issued by the application "A" indicates where the architected prefetch code corresponding to the application "A" has been stored.
[0020] The architected Scout prefetch instruction 70 includes an operation code 72 for specifying the operation to be performed, as well as a base register 74, an index register 76, and a relative address 78 for specifying a memory location of a starting address at which the specified prefetch code has been stored. The architected Scout prefetch instruction 70 also indicates the number 80 of the specified Scout core into which the prefetch code is to be loaded (for example, the number 80 of the specified Scout core may be Fig. 2B may display either Scout Core 1, Scout Core 2, or Scout Core 3). Each field (e.g., operation code 7072, base register 74, index register 76, relative address, and Scout Core number 80) may be a multi-bit field. The number of bits may be different for each field. The parent core 32, according to the Fig. 2A and Fig. 3 executes the architected Scout prefetch instruction 70. Then, the parent core 32 stores the prefetch address 48 with the starting address of the prefetch code indicated by the architected Scout prefetch instruction 70.
[0021] According to Fig. 2A, upon receipt of the broadcast interrupt signal, an instruction pipeline of the scout core 34 is flushed. The scout core 34 can then execute the prefetch code indicated by the prefetch code starting address sent from the parent core 32 via the load prefetch bus 68 to restart a sequence of instructions.
[0022] Fig. 4 now shows a flowchart illustrating an exemplary method 200 for loading prefetch code from the parent core 32 to the scout core 34. According to the Fig. 2A through 4, the method 200 generally begins with block 202, where the broadcast interrupt signal is sent from the parent core 32 to the scout core 34 via the load prefetch bus 68 (or the load prefetch bus 168 if the scout cores 134 are connected to the parent core 132). The broadcast interrupt signal provides the memory location of the prefetch code starting address and the interrupt message. The method 200 may then proceed to block 203.
[0023] At block 204, the scout core 32 receives the broadcast interrupt signal indicating the interrupt notification via the load prefetch bus 68. The method 200 may then proceed to block 206.
[0024] In block 206, the instruction pipeline of scout core 34 is flushed, and scout core 34 executes the prefetch code indicated by the prefetch code starting address sent by parent core 32 over load prefetch bus 68. Method 200 may then exit.
[0025] Fig. 5 shows an exemplary method 300 illustrating a triggering of the broadcast interrupt signal by a task switch involving an operating system of the parent core 32. According to the Fig. 2A through 2B and 5, the method 300 generally begins with block 302, where it is determined whether the parent core 32 is affected by a task switch. If the parent core 32 is affected by a task switch, the method 300 may proceed to block 304.
[0026] In block 304, the parent core 32 stores the prefetch address 48 associated with the currently executing application. The method 300 may then proceed to block 306.
[0027] At block 306, the parent core 32 loads a new application and updates the prefetch address 48 with a prefetch code start address associated with the new application. The method 300 may then proceed to block 308.
[0028] In block 308, the master core 32 sends the broadcast interrupt signal to the scout core 34 via the load prefetch bus 68. The broadcast interrupt signal provides the memory location of the prefetch code starting address and the interrupt message. Method 300 may then exit.
[0029] Fig. 6 shows an exemplary method 400 illustrating triggering of the broadcast interrupt signal by the predetermined application executing in the parent core 32. According to the Fig. 2A to 2B, 3 and 6, the method generally begins with block 402, where the parent core 32 executes a predetermined application that issues the architected Scout preload instruction 70. ( Fig. 3). The method may then continue with block 404.
[0030] In block 404, the parent core 32 executes the architected Scout preload instruction 70. The method may then proceed to block 406.
[0031] In block 406, the parent core 32 then stores the prefetch address 48 with the prefetch code starting address specified by the architected scout prefetch instruction 70. The method 400 may then proceed to block 408.
[0032] In block 408, the parent core 32 sends the broadcast interrupt signal and the starting address of the prefetch code to the scout core 34 via the load prefetch bus 68. The method 400 may then exit.
[0033] Those skilled in the art will appreciate that one or more aspects of the present invention may be implemented as a system, method, or computer program product. Accordingly, one or more aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which may be generally referred to herein as a "circuit," "module," or "system." In addition, one or more aspects of the present invention may take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code stored thereon.
[0034] Any combination of one or more computer-readable media may be used. The computer-readable medium may be a computer-readable storage medium. A computer-readable storage medium may be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or unit, or any suitable combination thereof.Individual examples (a non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer diskette, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the context of this document, a computer-readable storage medium can be any tangible medium that can contain or store a program for use by or in connection with a system, apparatus, or device for executing instructions.
[0035] According to a Fig.7, a computer program product 700 includes, for example, one or more storage media 702, where the media may be tangible and / or non-transitory, for storing computer-readable program code means or program code logic 704 thereon to provide or enable one or more aspects of embodiments described herein.
[0036] Program code that is generated and stored on a tangible medium (including, but not limited to, electronic memory modules (RAM), flash memory, compact discs (CDs), DVDs, magnetic tapes, and the like) is often referred to as a "computer program product." The computer program product medium is typically readable by processing circuitry, preferably in a computer system for execution of the computer program product by the processing circuitry. Such program code, for example, may be generated using a compiler or assembler to compile instructions that, when executed, implement aspects of the invention.
[0037] Embodiments relate to a method, a system, and a computer program product for prefetching data on a chip having at least one scout core and a parent core. The method includes the parent core storing a starting address of the prefetch code. The starting address of the prefetch code indicates where a prefetch code is stored. The prefetch code is specifically configured to monitor the parent core based on a predetermined application executed by the parent core. The method includes the parent core sending a broadcast interrupt signal to the at least one scout core. The broadcast interrupt signal sent based on the starting address of the prefetch code is stored.The method includes monitoring the parent core by prefetch code executed by at least one scout core. The scout core executes the prefetch code based on receipt of the broadcast interrupt signal.
[0038] According to one embodiment, the method further includes storing the starting address of the prefetch code by the parent core based on a task switch occurring within the parent core.
[0039] According to one embodiment, the method further includes storing, by the parent core, the starting address of the prefetch code based on the given application issuing an instruction.
[0040] According to one embodiment, the method further includes indicating, via the broadcast interrupt signal, that a default prefetch code is to be loaded by the at least one Scout core. An architected state within the Scout core provides a storage location of the default prefetch code.
[0041] According to one embodiment, the method further includes providing a storage location of the start address of the prefetch code and an interrupt signal through the broadcast interrupt signal.
[0042] According to one embodiment, the method further includes a load prefetch bus arranged between the parent core and the at least one scout core. The broadcast interrupt signal is sent over the load prefetch bus.
[0043] According to one embodiment, the method further includes a shared cache connecting the parent core to the at least one scout core. A fetch request bus is provided to connect the parent core to both the shared cache and the at least one scout core.
[0044] Technical implications and advantages include that the scout core 34 can monitor the parent core 32 using prefetch code specifically tailored to monitor a given application executed by the parent core 32. The parent core 32 sends the broadcast interrupt signal over the load prefetch bus 68 to the scout core 34 to change the given prefetch code based on the application executed by the parent core 32. Thus, the approach discussed above increases the efficiency of the data processing system 10.
[0045] The terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting of the embodiments. The singular forms "a," "an," and "the" are intended to equally include the plural forms, unless the context indicates otherwise. Further, it is to be understood that the terms "comprises" or "having," when used in this specification, refer to the presence of specified features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0046] The respective structures, materials, acts, and equivalents of all means or steps plus functional elements in the following claims are intended to include any structure, material, or act for performing the function in combination with other expressly claimed elements. The description of the embodiments has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the embodiments in the form disclosed. Many modifications and variations will be apparent to those skilled in the art without departing from the scope or spirit of the embodiments. The embodiments were chosen and described in order to best explain the principles and practical application and to enable others skilled in the art to understand the embodiments with various modifications as are suitable for the intended use.
[0047] Computer program code for performing operations for aspects of the embodiments may be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, Smalltalk, C++, or the like, and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code may execute entirely on a user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server.In the latter case, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, over the Internet using an Internet service provider).
[0048] Aspects of embodiments are described above with reference to flowchart and / or schematic diagrams of methods, apparatus (systems), and computer program products according to the embodiments. It will be appreciated that each block of the flowchart and / or block diagrams, and combinations of blocks in the flowchart and / or block diagrams, may be implemented by computer program instructions. These computer program instructions may be supplied to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions executed by the processor of the computer or other programmable data processing apparatus produce means for implementing the functions / acts specified in the block or blocks of the flowchart and / or block diagrams.
[0049] These computer program instructions may also be stored in a computer-readable medium that can cause a computer, other programmable data processing apparatus, or other devices to function in a particular manner such that the instructions stored in the computer-readable medium produce an article of manufacture that includes instructions for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0050] The instructions of the computer program may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a sequence of operations to be performed on the computer, other programmable device, or other devices to produce a computer-based process such that the instructions executing on the computer or other programmable device provide processes for implementing the functions / acts specified in the block or blocks of the flowchart and / or block diagrams.
[0051] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments. Accordingly, each block in the flowchart or block diagrams may represent a module, segment, or portion of code comprising one or more executable instructions for implementing the specified logical function(s). It should be noted that in some alternative implementations, the functions specified in the block may occur in a different order than that shown in the figures. For example, two blocks shown in sequence may actually execute substantially concurrently, depending on the intended functionality, or the blocks may sometimes execute in the reverse order.It should also be noted that each block of the block diagrams and / or flowchart, and combinations of blocks in the block diagrams and / or flowchart, can be implemented by special purpose hardware systems that perform the specified functions or actions, or by combinations of special purpose hardware and computer instructions.
Claims
[1] A computer system (10) for prefetching data on a chip (20, 120), the system (10) comprising: a master core (32, 132) disposed on the chip (20, 120), the master core (32, 132) being configured to selectively execute a plurality of applications, one of the plurality of applications being a particular application; and at least one scout core arranged on the chip (20, 120), wherein the system (10) is configured to perform a method (200) comprising: Storing, by the parent core (32, 132), a starting address of a prefetch code, wherein the starting address of the prefetch code indicates where the prefetch code is stored, and wherein the prefetch code is specifically configured to monitor the parent core (32, 132) based on the particular application being executed; Sending (202, 308, 408) a broadcast interrupt signal by the master core (32, 132) to the at least one scout core (34, 134), the broadcast interrupt signal being sent (202, 308, 408) based on the prefetch code start address having been stored (304, 406); and Monitoring the parent core (32, 132) by the at least one scout core (34, 134), wherein the at least one scout core (34, 134) executes (206) the prefetch code to monitor the parent core (32, 132), wherein the execution (206) of the prefetch code is based on receiving (204) the broadcast interrupt signal. [2] The computer system (10) of claim 1, wherein the parent core (32, 132) stores (304) the starting address of the prefetch code based on a task switch occurring within the parent core (32, 132). [3] The computer system (10) of claim 1, wherein the parent core (32, 132) stores the starting address of the prefetch code based on the particular application issuing an instruction. [4] The computer system (10) of any preceding claim, wherein the broadcast interrupt signal indicates that a default prefetch code is to be loaded by the at least one scout core (34, 134), wherein an architected state located within the at least one scout core (34, 134) provides a storage location of the default prefetch code. [5] The computer system (10) of any preceding claim, wherein the broadcast interrupt signal provides a storage location of the start address of the prefetch code and an interrupt signal. [6] The computer system (10) of any preceding claim, wherein a load prefetch bus (68, 168) is disposed between the parent core (32, 132) and the at least one scout core (34, 134), and wherein the broadcast interrupt signal is sent (202, 308, 408) over the load prefetch bus (68, 168). [7] The computer system (10) of any preceding claim, further comprising a shared cache (50) connecting the parent core (32) to the at least one scout core (34), wherein a fetch request bus (64) is provided to connect the parent core (32) to both the shared cache (50) and the at least one scout core (34). [8] A computer program product (700) for prefetching data on a chip (20, 120) having at least one scout core (34, 134) and a superordinate core (32, 132), the computer program product (700) comprising: a tangible storage medium (702) readable by a processing circuit and having stored therein instructions (704) for execution by the processing circuit to perform a method (200) comprising: Storing, by the parent core (32, 132), a prefetch code starting address, wherein the prefetch code starting address indicates where a prefetch code is stored, and wherein the prefetch code is specifically configured to monitor the parent core (32, 132) based on a particular application executed by the parent core (32, 132); Sending (202, 308, 408) a broadcast interrupt signal by the master core (32, 132) to the at least one scout core (34, 134), the broadcast interrupt signal being sent (202, 308, 408) based on the prefetch code start address having been stored (304, 406); and Monitoring the parent core (32, 132) by the at least one scout core (34, 134), wherein the at least one scout core (34, 134) executes (206) the prefetch code to monitor the parent core (32, 132), and wherein executing (206) the prefetch code is based on receiving (204) the broadcast interrupt signal. [9] The computer program product (700) of claim 8, wherein the parent core (32, 132) stores (304) the starting address of the prefetch code based on a task switch occurring within the parent core (32, 132). [10] The computer program product (700) of claim 8, wherein the parent core (32, 132) stores the starting address of the prefetch code based on the particular application issuing an instruction. [11] The computer program product (700) of any one of claims 8 to 10, wherein the broadcast interrupt signal indicates that a default prefetch code is to be loaded by the at least one scout core (34, 134), wherein an architected state located within the at least one scout core (34, 134) provides a storage location of the default prefetch code. [12] The computer program product (700) of any one of claims 8 to 11, wherein the broadcast interrupt signal provides a storage location of the start address of the prefetch code and an interrupt signal. [13] The computer program product (700) of any one of claims 8 to 12, wherein a load prefetch bus (68, 168) is disposed between the parent core (32, 132) and the at least one scout core (34, 134), and wherein the broadcast interrupt signal is sent (202, 308, 408) over the load prefetch bus (68, 168). [14] A computer-assisted method (200) for prefetching data on a chip (20, 120) having at least one scout core (34, 134) and a superordinate core (32, 132), the method (200) comprising: Storing, by the parent core (32, 132), an address of a prefetch code, wherein the starting address of the prefetch code indicates where a prefetch code is stored, and wherein the prefetch code is specifically configured to monitor the parent core (32, 132) based on a particular application executed by the parent core (32, 132); Sending (202, 308, 408) a broadcast interrupt signal by the master core (32, 132) to the at least one scout core (34, 134), the broadcast interrupt signal being sent based on the start address of the prefetch code being stored (304, 406); and Monitoring the parent core (32, 132) by the at least one scout core (34, 134), wherein the at least one scout core (34, 134) executes (206) the prefetch code to monitor the parent core (32, 132), and wherein executing (206) the prefetch code is based on receiving (204) the broadcast interrupt signal. [15] The method (200) of claim 14, wherein the parent core (32, 132) stores (304) the starting address of the prefetch code based on a task switch occurring within the parent core (32, 132). [16] The method (200) of claim 14, wherein the parent core (32, 132) stores the starting address of the prefetch code based on the particular application issuing an instruction. [17] The method (200) of any one of claims 14 to 16, wherein the broadcast interrupt signal indicates that a default prefetch code is to be loaded by the at least one scout core (34, 134), wherein an architected state located within the at least one scout core (34, 134) provides a memory location of the default prefetch code. [18] The method (200) of any one of claims 14 to 17, wherein the broadcast interrupt signal provides a storage location of the start address of the prefetch code and an interrupt signal. [19] The method (200) of any one of claims 14 to 18, wherein a load prefetch bus (68, 168) is arranged between the parent core (32, 132) and the at least one scout core (34, 134), and wherein the broadcast interrupt signal is sent (202, 308, 408) over the load prefetch bus (68, 168). [20] The method (200) of any one of claims 14 to 19, comprising a shared cache (50) connecting the parent core (32) to the at least one scout core (34), wherein a fetch request bus (64) is provided to connect the parent core (32) to both the shared cache (50) and the at least one scout core (34).
Citation Information
Patent Citations
Sideband scout thread processor
US20040148491A1
Method and apparatus for efficient helper thread state initialization using inter-thread register copy
US20110296431A1