DISTRIBUTION OF INJECTED DATA TO THE CACHE OF A DATA PROCESSING SYSTEM
The described data processing system optimizes cache operations by implementing cache injection and distribution mechanisms to enhance performance through efficient distribution of cache lines across multiple vertical cache hierarchies, addressing inefficiencies in existing multiprocessor systems.
Patent Information
- Application Number
- DE112022003780
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-08-04
- Filing Date
- 2022-08-04
- Publication Date
- 2025-11-13
- Estimated Expiration
- 2042-08-04
AI Technical Summary
Existing cache hierarchies in multiprocessor systems face inefficiencies in optimizing castout operations between different levels, leading to suboptimal performance.
A data processing system with a plurality of processor cores and vertical cache hierarchies that performs cache injection and distribution based on a distribution field in the directory entry, allowing for lateral castouts to optimize cache line distribution among multiple vertical cache hierarchies.
Enhances overall system performance by efficiently distributing cache lines across multiple cache levels, reducing latency and improving resource utilization.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND OF THE INVENTION
[0001] The present invention relates generally to data processing and in particular to the injection and distribution of data onto the cache of a data processing system.
[0002] A conventional multiprocessor (MP) computer system, such as a server computer system, comprises several processing units, all connected by a system interface that typically includes one or more address, data, and control buses. Connected to the system interface is system memory, which represents the lowest level of volatile memory in the multiprocessor computer system and is generally accessible for read and write access by all processing units. To reduce access latency for instructions and data residing in system memory, each processing unit is typically further supported by its own multilevel cache hierarchy, the lower level(s) of which may be shared by one or more processor cores.
[0003] Cache memory is typically used to temporarily store blocks of memory that a processor might access, thereby speeding up processing by reducing access latency introduced by having to load required data and instructions from system memory. In some multiprocessing (MP) systems, the cache hierarchy comprises at least two levels. The level one (L1) or parent cache is usually a private cache belonging to a specific processor core and inaccessible to other cores in the MP system. Typically, in response to a memory access instruction, such as a load or store instruction, the processor core first accesses the directory of the parent cache.If the requested memory block is not found in the higher-level cache, the processor core then accesses lower-level caches (e.g., level two (L2) or level three (L3)) or system memory for the requested memory block.
[0004] Several published documents already exist in this context: Document US 2012 / 0203973A1 describes a method for a selective lateral cache-to-cache castout. In this method, a first processor unit has one cache each at the first upper and first lower levels. The same applies to the respective caches at the second level. A decision is made between a lateral castout between two caches of the same level and a castout to system memory. This distinction is based on a confidence indicator. Document US 2010 / 0262783A1 describes a mode-based castout or a corresponding target selection. Such a method can be successfully used to prevent a castout to system memory, for example, during testing. Document US 2010 / 0235577A1 describes a target selection for a lateral cache castout.This takes into account architectural similarities between different caches. Finally, document US 2008 / 0059707A1 describes the selective storage of data in different levels of a cache memory. This involves a connection between an L1 cache, an L2 cache, and an L3 cache via a corresponding counter value.
[0005] Despite these advances already made, there is still a need to optimize castout operations between different cache levels in order to achieve higher overall performance. SUMMARY
[0006] This task is solved by the subject matter of the independent patent claims. Further details arise from the respective dependent patent claims.
[0007] In at least some embodiments, a data processing system comprises a plurality of processor cores, each supported by one or more vertical cache hierarchies. Based on a system structure receiving a cache injection request that requests data to be injected into a cache row identified by a real target address, the data is written to a cache in a first vertical cache hierarchy among the plurality of vertical cache hierarchies. Based on a value in a field of the cache injection request, a distribution field in a directory entry of the first vertical cache hierarchy is set. After the cache row is removed from the first vertical cache hierarchy, it is determined whether the distribution field is set.Based on the determination that the distribution field is set, a lateral castout of the cache row from the first vertical cache hierarchy to a second vertical cache hierarchy is performed. Brief description of the multiple views of the drawings Fig. Figure 1 is an overview block diagram of an exemplary data processing system according to one embodiment; Fig. Figure 2 is a more detailed block diagram of an exemplary processing unit according to one embodiment; Fig. Figure 3 is a more detailed block diagram of a level two (L2) cache according to one embodiment; Fig. Figure 4 is a more detailed block diagram of a level three (L3) cache according to one embodiment; Fig. Figure 5 illustrates an exemplary cache directory entry according to one embodiment; Fig. Figures 6A to 6B each represent an exemplary direct memory access (DMA) write request and an exemplary cache injection request according to one embodiment; Fig. Figures 7A to 7B each illustrate an exemplary castout (CO) requirement and an exemplary lateral castout (LCO) requirement according to one embodiment; Fig. Figure 8 represents a distribution instruction for an exemplary data cache block zero (DCBZ) according to one embodiment; Fig. Figure 9 is a logical overview flowchart of an exemplary procedure by which a working memory area is initialized for storing injected data according to an embodiment; Fig. Items 10 to 13 together form a logical overview flowchart of an exemplary method for injecting at least one partial cache row of data into a cache of a subordinate level (e.g. L3 cache) according to an embodiment; Fig. Figure 14 is a logical overview flowchart of an exemplary procedure for processing a DMA write request according to one embodiment; Fig. 15 is a logical overview flowchart of an exemplary procedure by which a cache of a higher level (e.g. L2) performs a castout according to an embodiment; Fig. 16 is a logical overview flowchart of an exemplary procedure by which a cache of a subordinate level (e.g. L3) processes a castout from a cache of a superior level (e.g. L2) according to an embodiment; Fig. Figure 17 is a logical overview flowchart of an exemplary procedure by which a cache of a subordinate level (e.g. L3) performs a castout according to an embodiment; Fig. Figure 18 is a logical overview flowchart of an exemplary procedure by which a cache of a lowest level (e.g. L4) performs a castout according to an embodiment; Fig. 19 is a logical overview flowchart of an exemplary procedure by which a cache of a higher level (e.g. L2) retrieves data into its array according to an embodiment; Fig. Figure 20 is a data flow diagram that illustrates a design process. DETAILED DESCRIPTION
[0008] With the following reference to the figures, in which the same reference numerals consistently refer to similar and corresponding parts, and in particular with reference to Fig. Reference 1 illustrates a block diagram depicting an exemplary data processing system 100 according to one embodiment. In the illustrated embodiment, the data processing system 100 is a cache-coherent symmetric multiprocessor (SMP) data processing system comprising several processing nodes 102a, 102b for processing data and instructions. The processing nodes 102 are connected via a system link 110 for transmitting address, data, and control information. The system link 110 can be implemented, for example, as an interconnection with buses, a switched connection, or a hybrid connection.
[0009] In the illustrated embodiment, each processing node 102 is implemented as a multi-chip module (MCM) containing several (e.g., four) processing units 104a to 104d, each preferably implemented as a separate integrated circuit. The processing units 104 in each processing node 102 are connected to each other and to the system connection 110 via a local connection 114 for data transmission. This local connection 114 can be implemented like the system connection 110, for example, with one or more buses and / or switches. The system connection 110 and the local connections 114 together form a connection structure.
[0010] As explained in more detail below with reference to Fig. As described in Figure 2, each processing unit 104 comprises a memory controller 206 connected to the local connection 114 to provide an interface to a respective system memory 108. Data and instructions located in the system memory 108 can generally be accessed, cached, and modified by a processor core in each processing unit 104 of each processing node 102 in the data processing system 100. The system memory 108 thus forms the lowest level of memory in the distributed, shared memory system of the data processing system 100. In alternative embodiments, one or more memory controllers 206 (and system memory 108) can be connected to the system connection 110 instead of a local connection 114.
[0011] The expert will understand that the SMP data processing system is 100% of Fig. 1. Many additional components, not illustrated, may be included, such as connecting bridges, non-volatile memory, connectors for network or connected units, etc. Since such additional components are not necessary for understanding the described embodiments, they are omitted in Fig. 1 is not illustrated or discussed further herein. However, it should also be clear that the improvements described herein are applicable to cache-coherent data processing systems of various architectures and are by no means limited to the generalized architecture of the data processing system described in Fig. 1 is illustrated.
[0012] With the following reference to Fig. Figure 2 shows a more detailed block diagram of an exemplary processing unit 104 according to one embodiment. In the illustrated embodiment, each processing unit 104 is an integrated circuit comprising two or more processor cores 200a, 200b for processing instructions and data. In at least some embodiments, each processor core 200 is capable of executing multiple simultaneous hardware threads of an execution independently of one another.
[0013] As shown, each processor core 200 comprises one or more execution units, such as a load-memory unit (LSU) 202, for executing instructions. The instructions executed by the LSU 202 include memory access instructions that request a load or memory access to a memory block in the distributed, shared memory system, or that cause the generation of a request for a load or memory access to a memory block in the distributed, shared memory system. Memory blocks obtained from the distributed, shared memory system through load accesses are cached in one or more register files (RFs) 208, and memory blocks updated through memory accesses are written from the one or more register files 208 to the distributed, shared memory system.
[0014] The operation of each processor core 200 is supported by a multi-level memory hierarchy, which at its lowest level has a shared system memory 108 accessed via an integrated memory controller 206. As indicated by the dashed line, the system memory 108 can optionally include a collection of D-bits 210, including a plurality of bits, each belonging to one of the memory blocks in the system memory 108. A D-bit is set (e.g., to 1) to indicate that the associated memory block belongs to a data set in which data is to be distributed to the various vertical cache hierarchies of the data processing system 100, and is otherwise reset (e.g., to 0).At its higher levels, the multi-level memory hierarchy comprises one or more levels of cache memory. In the illustrated embodiment, the cache hierarchy includes a level one (L1) memory-permeable cache 226 within and private for each processor core 200, a respective level two (L2) ingest cache 230a, 230b for each processor core 200a, 200b, a respective level three (L3) lookaside victim cache 232a, 232b for each processor core 200a, 200b, which is filled with cache rows removed from one or more L2 caches 230, and optionally a level four (L4) cache 234, which caches data written to or read from the system memory 108. If present, the L4 cache 234 includes an L4 array 236 for temporarily storing cache rows of data and an L4 directory 238 of the contents of the L4 array 236.In the illustrated embodiment, the L4 cache 234 only temporarily stores copies of memory blocks corresponding to those stored in the associated system memory 108. In other embodiments, the L4 cache 234 can alternatively be configured as a general-purpose last-level cache that temporarily stores copies of memory blocks corresponding to those stored in any of the system memory 108. A person skilled in the art will understand from the following discussion the modifications to the disclosed embodiments that would be necessary or desirable if the L4 cache 234 were instead configured to serve as a general-purpose last-level cache.As shown in detail for the L2 cache 230a and the L3 cache 232a, each L2-L3 cache interface comprises a number of channels, including a read (RD) channel 240, a cast-in (CI) channel 242, and a write-injection (WI) channel 244. Each of the L2 cache 230 and L3 cache 232 is further connected to the local link 114 and a structure controller 216 to enable the caches 230 and 232 to participate in the coherent data transmission of the data processing system 100.
[0015] Although the illustrated cache hierarchies comprise three or four cache levels, it will be clear to the person skilled in the art that alternative embodiments may include additional levels of chip-integrated or chip-external, private or shared, in-line or lookaside cache, which may be fully inclusive, partially inclusive, or non-inclusive with respect to the contents of the higher-level caches.
[0016] Each processing unit 104 further comprises an integrated and distributed structure controller 216, which is responsible for controlling the flow of operations on the connection structure that includes the local connection 114 and the system connection 110, and for implementing the coherence data transmission required to implement the selected cache coherence protocol. The processing unit 104 also includes an integrated I / O (input / output) controller 214, which supports the connection of one or more I / O units 218.
[0017] When a hardware thread executing on a processor core 200 during operation includes a memory access instruction (i.e., load or store) that requests a specific memory access operation, the LSU 202 executes the memory access instruction to determine the target address (e.g., an effective address) of the memory access request. After translating the target address to a real address, the L1 cache 226 is accessed using the real target address. If it is determined that the specified memory access cannot be fulfilled solely by referencing the L1 cache 226, the LSU 202 then transfers the memory access request, which includes at least one transaction type (ttype) (e.g., load or store) and the real target address, to its associated L2 cache 230 for servicing.When handling the memory access request, the L2 cache 230 can access its associated L3 cache 232 and / or initiate a transaction that includes the memory access request on the connection structure.
[0018] With the following reference to Fig. Figure 3 shows a more detailed block diagram of an exemplary embodiment of an L2 cache 230 according to one embodiment. As in Fig. As shown in Figure 3, the L2 cache 230 comprises an L2 array 302 and an L2 directory 308 containing the contents of the L2 array 302. Although not explicitly illustrated, the L2 array 302 is preferably implemented with a single read port and a single write port to reduce the chip area required to implement the L2 array 302.
[0019] Assuming that the L2 array 302 and the L2 directory 308 are set-associative as usual, memory locations in the system memory 108 are mapped to specific congruence classes in the L2 array 302 by using predetermined index bits in the (real) addresses of the system memory. The special memory blocks stored in the cache rows of the L2 array 302 are recorded in the L2 directory 308, which contains one directory entry for each cache row. Although in Fig. Although not explicitly shown in Figure 3, it will be clear to the person skilled in the art that each directory entry in the L2 directory 308 comprises various fields, for example a tag field that identifies the real address of the memory block stored in the corresponding cache row of the L2 array 302, a state field that indicates the coherence state of the cache row, a replacement order field (e.g. LRU (Least Recently Used)) that specifies a replacement order for the cache row with respect to other cache rows in the same congruence class, and inclusivity bits that indicate whether the memory block is located in the associated L1 cache 226.
[0020] The L2 cache 230 additionally includes a read request logic 311, which comprises several (e.g., 16) read request (RC) machines 312 for independently and simultaneously handling load (LD) and store (ST) requests received by the associated processor core 200. As should be clear, handling memory access requests by the RC machines 312 may require replacing or invalidating memory blocks in the L2 array 302. Accordingly, the L2 cache 230 also includes a castout logic 309, which comprises several CO (castout) machines 310 that independently and simultaneously manage the removal of memory blocks from the L2 array 302 and the storage of these memory blocks in the system memory 108 (i.e. write-back operations) or an L3 cache 232 (i.e., L3 cast-ins).
[0021] To handle remotely located memory access requests originating from processor cores 200 other than the associated processor core 200, the L2 cache 230 also includes snoop logic 313, which comprises multiple snoop machines 314. The snoop machines 314 can independently and simultaneously handle a remotely located memory access request that has been "spyed on" by the local connection 114 via a monitoring element. As in Fig. As shown in Figure 3, the Snoop logic 313 is connected to the associated L3 cache 232 via a Wl channel 244, which is also in Fig. Figure 2 illustrates this. The Wl channel 244 preferably comprises (from the perspective of the L2 cache 230) several signal lines suitable for transmitting at least one "L2 data ready" signal (L2 D_rdy), (preferably comprising a separate signal line for each SN machine 314 in the L2 cache 230), a multi-bit L2 Snoop machine ID signal (L2 SN-ID), and an "L2 injection OK" signal (L2 I_OK), and for receiving an "L3 injection OK" signal (L3 I_OK), a multi-bit L3 injection machine ID signal (L3 WI ID), and an "L3 done" signal (preferably comprising a separate signal line for each W machine 314 in the L3 cache 232).
[0022] The L2 cache 230 further includes an allotment 305, which controls the multiplexers M1 to M2 to arrange the processing of local memory access requests received by the associated processor core 200 and remote memory access requests intercepted on the local connection 114. Such memory access requests, including local load and store requests and remote load and store requests, are routed to logic allocation according to the allocation policy implemented by the allotment 305. This allocation pipeline 306 processes each memory access request with respect to the L2 directory 308 and the L2 array 302 and, if necessary and if the required resource is available, allocates the memory access request to the appropriate state machine for processing.
[0023] The L2 cache 230 also includes an RC queue (RCQ) 320 and a castout push intervention (CPI) queue 318, each of which caches data that is inserted into and removed from the L2 array 302. The RCQ 320 comprises a number of buffer entries, each individually corresponding to a specific RC machine 312, so that each allocated RC machine 312 retrieves data only from the specified buffer entry. Likewise, the CPI queue 318 includes a number of buffer entries, each individually corresponding to a specific one of the Castout machines 310 and Snoop machines 314, so that the CO machines 310 and Snoopers 314 transfer data from the L2 array 302 (e.g. to another L2 cache 230, to the associated L3 cache 232 or to a system memory 108) directly only via their respective specified CPI buffer entries.
[0024] Each RC machine 312 is also assigned one of several RC data (RCDAT) buffers 322 to temporarily store a working memory block read from the L2 array 302 and / or received from the local connection 114 via the reload bus 323. The RCDAT buffer 322 assigned to each RC machine 312 is preferably constructed with connections and functionality that meet the working memory access requirements that can be served by the associated RC machine 312. The RCDAT buffers 322 have an associated memory data multiplexer M4, which selects data bytes from its inputs for temporary storage in the RCDAT buffer 322 in response to selection signals (not illustrated) generated by the allocator 305.
[0025] During operation, a processor core 200 transmits memory requests, which include a transaction type (ttype), a real destination address, and memory data, to a memory queue (STQ) 304. From STQ 304, the memory data is transferred to a multiplexer M4 via a data path 324 for data storage, and the transaction type and destination address are passed to multiplexer M1. Multiplexer M1 also receives as inputs processor load requests from processor core 200 and directory write requests from the RC machines 312. In response to selection signals (not illustrated) generated by the allotment 305, multiplexer M1 selects one of its input requests to forward to multiplexer M2, which additionally receives as input a remote memory access request from the local connection 114 via a remote request path 326.The allotment 305 schedules local and remote memory access requests for processing and generates a sequence of selection signals 328 based on the schedule. In response to the selection signals 328 generated by the allotment 305, the multiplexer M2 selects either the local memory access request received by the multiplexer M1 or the remote memory access request spied on by the local link 114 as the next memory access request to be processed.
[0026] The memory access request selected for processing by the allotment 305 is placed in the allocation pipeline 306 by the multiplexer M2. The allocation pipeline 306 is preferably implemented as a fixed-duration pipeline in which each of the several possible overlapping requests is processed for a predetermined number of clock cycles (e.g., 4 cycles). During the first processing cycle in the allocation pipeline 306, a directory read operation is performed using the request address to determine whether the request address in the L2 directory 308 matches or misses, and if the memory address matches, to determine the coherence state of the target memory block. The directory information, which includes a hit / miss information and the coherence state of the memory block, is returned from the L2 directory 308 to the allocation pipeline 306 in a subsequent cycle.As should be clear, generally no action is taken in an L2 cache 230 in response to a missed request on a remote memory access request; such remote memory requests are accordingly discarded from the allocation pipeline 306. However, in the case of a hit or missed request on a local memory access request, or a hit on a remote memory access request, the L2 cache 230 will process the memory access request, which, for requests that cannot be fully processed in the processing unit 104, may result in data being transferred over the local connection 114 via the structure controller 216.
[0027] At a predetermined time during the processing of the memory access request in allocation pipeline 306, the allocator 305 transmits the request address to the L2 array 302 via an address and control path 330 to initiate a cache read operation of the memory block specified by the request address. The memory block read from the L2 array 302 is transmitted via a data path 342 to the multiplexer M4 for insertion into the corresponding RCDAT buffer 322. For processor load requests, the memory block is also transmitted to load the data multiplexer M3 via a data path 340 for forwarding to the associated processor core 200.
[0028] In the final cycle of processing a memory access request in Allocation Pipeline 306, Allocation Pipeline 306 makes an allocation determination based on a set of criteria that include (1) the presence of an address collision between the request address and a previous request address currently being processed by a Castout Machine 310, a Snoop Machine 314, or an RC Machine 312, (2) the directory information, and (3) the availability of an RC Machine 312 or Snoop Machine 314 to process the memory access request. If Allocation Pipeline 306 makes an allocation decision that the memory access request must be allocated, the memory access request is allocated by Allocation Pipeline 306 to an RC Machine 312 or a Snoop Machine 314.If the allocation of the memory access request is unsuccessful, the requester (e.g., the local or remote processor core 200) is notified of the error by a "Retry" response. The requester can then retry the failed memory access request if necessary.
[0029] While RC machine 312 is processing a local memory access request, it is in an occupied state and is unavailable to handle another request. While RC machine 312 is in an occupied state, it can perform a directory write operation to update the relevant entry in the L2 directory 308, if necessary. Additionally, RC machine 312 can perform a cache write operation to update the relevant cache row of the L2 array 302. Directory writes and cache writes can be scheduled by the allocator 305 during any interval in which the allocation pipeline 306 is not already processing other requests according to the established directory read and cache read schedule.When all operations for the specified requirement are completed, the RC machine 312 returns to an unoccupied state.
[0030] Associated with the RC machines 312 is a data processing circuit, different sections of which are used when handling various types of local memory access requests. For example, for a local load request involving the L2 directory 308, a copy of the target memory block is forwarded from the L2 array 302 to the associated processor core 200 via data path 340 and the load data multiplexer M3, and additionally to the RCDAT buffer 322 via data path 342. The data forwarded to the RCDAT buffer 322 via data path 342 and the memory data multiplexer M4 is then forwarded from the RCDAT 322 to the associated processor core 200 via data path 360 and multiplexer M3.For a local memory request, memory data in the RCDAT buffer 322 is received from the STQ 304 via data path 324 and the memory data multiplexer M4. The memory is merged with the RAM block read operation via the L2 array 302 through the multiplexer M4, and the merged memory data is then written from the RCDAT buffer 322 to the L2 array 302 via data path 362. In response to a local load failure or local memory failure, the target RAM block, captured by issuing a RAM access operation on local connection 114, is loaded into the L2 array 302 via a reload bus 323, the memory data multiplexer M4, and the RCDAT buffer 322 (with memory merge for a memory failure).
[0031] With the following reference to Fig. Figure 4 shows a more detailed view of an L3 cache 232 according to one embodiment. As in Fig. As shown in Figure 4, the L3 cache 232 comprises an L3 array 402 and an L3 directory 408 containing the contents of the L3 array 402. Assuming that the L3 array 402 and the L3 directory 408 are set-associative as usual, memory locations in the system memory 108 are mapped to specific congruence classes in the L3 array 402 by using predetermined index bits in the (real) addresses of the system memory. The special memory blocks stored in the cache rows of the L3 array 402 are recorded in the L3 directory 408, which contains one directory entry for each cache row. Although in Fig. Although not explicitly shown in Figure 4, it will be clear to the person skilled in the art that each directory entry in the L3 directory 408 comprises different fields, for example a tag field that identifies the real address of the memory block located in the corresponding cache row of the L3 array 402, a state field that indicates the coherence state of the cache row, a replacement order field (e.g. LRU) that specifies a replacement order for the cache row with respect to other cache rows in the same congruence class.
[0032] The L3 cache 232 additionally includes various state machines to handle different types of requests and transfer data to and from the L3 array 402. For example, the L3 cache 232 includes several (e.g., 16) read (RD) machines 412 for independently and simultaneously handling read (RD) requests received by the associated L2 cache 230 via the RD channel 240. The L3 cache 232 also includes several snoop (SN) machines 411 for handling remote memory access requests snooped from the local connection 114, originating from the L2 cache 230 supporting the remote processor cores 200. As is known in the prior art, serving spied-upon requests can, for example, include invalidating cache rows in the L3 directory 408 and / or sourcing cache rows of data from the L3 array 402 by cache-to-cache intervention.Additionally, the L3 cache 232 includes several cast-in (CI) machines 413 for handling cast-in (CI) requests received by the associated L2 cache 230 via CI channel 242. As should be clear, handling cast-in requests by CI machines 413 by storing cache lines in the L3 array 402, for which a castout from the associated L2 cache 230 has occurred, may require the replacement of memory blocks in the L3 array 402. Accordingly, the L3 cache 232 also includes castout (CO) machines 410, which manage the removal of memory blocks from the L3 array 402 and, if necessary, the writing of these memory blocks back to system memory 108. Data removed from the L3 cache 232 by the CO machines 410 and SN machines 411 are temporarily stored in a Castout Push Intervention (CPI) queue 418 before being transferred to the local connection 114.Furthermore, the L3 cache 232 comprises a plurality of write injection (WI) machines 414 that serve requests received on the local connection 114 to inject partial or complete cache rows of data into the L3 array 402 of the L3 cache 232. Write injection data received in connection with write injection requests is temporarily stored in a write injection queue 420 (WIQ), which preferably comprises one or more entries, each with the width of a complete cache row (e.g., 126 bytes). In a preferred embodiment, write injection requests are served exclusively by the L3 cache 232 to avoid introducing additional complexity into higher-level caches with lower access latency requirements, such as the L2 cache 230.One or more of the SN machines 411, the Cl machines 413, and the WI machines 414 additionally process lateral castout (LCO) requests from other L3 caches 232, which were spied on from the local connection 114, and in this way install cache rows with data received in connection with the LCO requests in the L3 array 402. Again, handling LCO requests by storing cache rows in the L3 array 402, for which a castout from other L3 caches 232 has occurred, may require the replacement of cache rows already located in the L3 array 402.
[0033] The L3 cache 230 further includes an allocation unit 404, which sequences the processing of CI and RD requests received from the associated L2 cache 230, as well as remote memory access requests, LCO requests, and write injection requests intercepted from the local connection 114. These memory access requests are routed according to the allocation policy implemented by the allocation unit 404 to allocate logic, such as an allocation pipeline 406, which processes each memory access request with respect to the L3 directory 408 and the L3 array 402 and, if necessary, assigns the memory access requests to the appropriate state machines 411, 412, 413, or 414 for processing.If necessary, at a predetermined time during the processing of the memory access request in the allocation pipeline 406, the allotment 404 transmits the actual destination address of the request to the L3 array 402 via an address and control path 426 to initiate a cache read operation of the memory block specified by the actual destination address of the request.
[0034] The allotment 404 is further connected to a heuristic lateral castout (LCO) logic 405, which, based on a variety of factors such as workload characteristics, hit rates, etc., indicates whether victim cache rows to be removed from the L3 array 402 should undergo a vertical castout to a lower-level memory (e.g., L4 cache 234 or system memory 108) or a lateral castout to another L3 cache 232. As further discussed herein, the allotment 404 generally determines, based on the information provided by the heuristic LCO logic 405, whether a cache row should undergo a vertical or lateral castout. For the subset of cache rows marked with a set distribution field (D) in the L3 directory 408, the allocator 404 preferably does not determine whether a vertical or lateral castout should be performed solely on the basis of the heuristic LCO logic 405.For such cache rows, the allocator 404 causes the L3 cache 232 to distribute the cache rows across vertical cache hierarchies instead, for example according to the process described below with reference to . Fig. 16.
[0035] Fig. Figure 4 further illustrates data processing logic and data paths used to handle various types of memory access requests in the illustrated embodiment. In the illustrated embodiment, the data processing logic comprises multiplexers M5, M6, and M7 and a write injection buffer 422. The data paths include a data path 424, which can forward data read from the L3 array 402, for example, in response to an RD request or a Wl request, to multiplexer M5 and RD channel 240. The L3 cache 232 further includes a data path 428, which can forward CI data received from the associated L2 cache 230 to multiplexers M5 and M7. L3 also includes data path 428, which can forward WI data located in WIQ 420 to the multiplexer M6.The selection signals generated by the allocator 404 (not illustrated) select which data, if applicable, is written to the L3 array 402 by the multiplexer M7.
[0036] In data processing systems such as Data Processing System 100, it is common practice for I / O units like I / O Unit 218 to issue requests on the system structure to write data to the memory hierarchy. When data from an I / O unit is to be written directly to cache memory instead of system memory 108, such a request is called a "cache injection" request. When the data from the I / O unit is to be written to system memory 108 (or its associated memory cache, such as L4 cache 234), the request is called a direct memory access (DMA) write request. Generally, it is preferred to write I / O data to the cache hierarchy rather than system memory 108 or L4 cache 234 due to the lower access latency of cache memory.
[0037] In some cases, however, the data set to be written to main memory by an I / O unit is large compared to the storage capacity of a single cache, and the volume of cache injection requests associated with writing such a large data set can overwhelm the resources (e.g., the WI machines 414 and cache rows in the L3 array 402) in any cache required to handle the cache injection requests. Therefore, the present application recognizes that it would be useful and desirable to enable the selective distribution of cache injection request data across multiple vertical cache hierarchies when it is first written to the main memory system of a data processing system.
[0038] Furthermore, the present application recognizes that, because a cache injection request is by definition a request to update a cache row, a cache injection request can only be successful if the cache row targeted by the cache injection request exists in a cache in a coherence state, which means that the cache in which the cache row resides has write permission for the cache row (i.e., the HPC, as discussed below).Accordingly, the present disclosure provides a statement that specifies a vertical cache hierarchy that receives an injected cache row, that enables a cache row stored in the specified vertical cache hierarchy to be initialized to a corresponding coherence state that allows a successful cache injection request, and that additionally specifies the injected cache row as belonging to a data record that is to be distributed across multiple vertical cache hierarchies.
[0039] The present disclosure further recognizes that the data set written to the memory system by cache injection is often claimed by a single processor core 200 or by processor cores 200 of a single processing unit 104. Since the processor core(s) 200 claim and potentially update the data set, the data set is centralized in the vertical cache hierarchy or hierarchies of a small number of processor cores 200. As the cache lines of the data set begin to age, a castout of the cache lines from higher levels of the memory hierarchy to lower levels of the memory hierarchy occurs.The present disclosure again recognizes that, as this castout process progresses, it would be useful and desirable for the cache rows in the dataset to distribute them across multiple cache hierarchies, rather than concentrating them on one or a few cache hierarchies.
[0040] In the embodiments disclosed herein, the distribution of the cache rows containing a data set of injected data is supported by the implementation of a distribution field (D). This field is stored in conjunction with submodules of the data set at various levels of the main memory hierarchy and is transferred to the system structure in conjunction with requests targeting the data set. A D field is set (e.g., to 1) to indicate that the associated data belongs to a data set in which data must be distributed to the various vertical cache hierarchies of the data processing system 100, and is otherwise reset (e.g., to 0). For example, the system main memory 108, again briefly referring to Fig. 1 optionally provide memory for D-bits 210, which can be implemented, for example, as a respective 65th bit for each 64-bit double word of data in system memory 108. Likewise, as in Fig. As shown in Figure 5, a directory entry 500 of a cache memory (e.g., in the L2 directory 308, L3 directory 408, or L4 directory 238) may, in addition to a possibly conventional valid field 502, an address tag field 504, and a coherence state field 506, include a D field 508 that indicates whether the associated cache row of data is a member of a record in which the data is to be distributed across the various vertical cache hierarchies of the data processing system 100. In the following discussion, it is assumed that each directory entry of each L2 directory 308 and L3 directory 408 contains a D-field 508, and that the entries of the L4 directories 238, if present, may optionally include a D-field 508, but will contain a D-field 508 if the system memory stores 108 optional D-bits 210.
[0041] With the following reference to Fig. Figures 6A to 6B each illustrate an exemplary direct memory access (DMA) write request 600 and an exemplary cache injection request 610 according to one embodiment. Requests 600 and 610 can be issued on the system structure of the data processing system 100 by the I / O controller 214, for example, in response to receiving corresponding requests from the I / O unit 218. As shown, requests 600 and 610 are similarly structured and include, for example, a valid field 602 or 612 indicating that the contents of the request are valid, a transaction type (ttype) field 604 or 614 identifying the type of request (i.e., DMA write or cache injection), and an address field 608 or 618 specifying the actual destination address of the request.In this example, each of the requests 600 and 610 additionally includes a distribution (D) field 606 or 616, which indicates whether the associated data submodule (which is transferred to a separate associated permanent data store on the system structure) belongs to a data record that should be distributed across the various vertical cache hierarchies of the data processing system 100. As above, the D field is set (e.g., to 1) to indicate that the data submodule should be distributed and is reset (e.g., to 0) otherwise. In embodiments where the D bits 210 are omitted from the system memory 108, the D field 606 in the DMA write request 600 can also be omitted.
[0042] With the following reference to Fig. Figures 7A to 7B illustrate an exemplary L2 castout (CO) request 700 and an exemplary lateral L3 castout (LCO) request 720 according to one embodiment. The CO request 700 can, for example, be issued from a higher-level cache to a lower-level cache or to the system memory 108 to release a cache row in the cache. The LCO request 720 can, for example, be issued on the system structure of the data processing system 100 by an L3 source cache 232 to release a cache row in the L3 array 402 of the L3 source cache 232, or to distribute a victim cache row of the L3 source cache 232 to another L3 cache 232, which was received as a castout from the associated L2 cache 230.As indicated, requests 700 and 720 are similarly structured and include, for example, a valid field 702 or 722 indicating that the contents of the request are valid, a transaction type (ttype) field 704 or 724 identifying the type of request, an address field 708 or 728 specifying the actual destination address of the castout cache line, and a state field 710 or 730 indicating a coherence state of the castout cache line. Additionally, each of requests 700 and 720 includes a distribution (D) field 706 or 726 indicating whether the associated data submodule (which is transferred to a separate associated permanent data store on the system structure) belongs to a record that should be distributed across the various vertical cache hierarchies of the data processing system 100. As above, the D field is set (e.g. to 1) to indicate that the data sub-module should be distributed, and is otherwise reset (e.g. to 0).Finally, the LCO request 720 includes a target ID field 732, which specifies the L3 target cache 232 that is to accept the castout cache row of data.
[0043] With the following reference to Fig. Figure 8 presents an exemplary distribution instruction of a data cache block zero (DCBZ_D) according to one embodiment. The DCBZ_D instruction 800 can be executed by an execution unit (e.g., the LSU 202) of each processor core 200 to create a specified cache row in a cache memory of the associated vertical cache hierarchy of that processor core 200, set the data of the cache row to zero, set the associated state field 506 to a coherence state indicating write permission (i.e., an HPC state), and set its D field 508 to a desired state. As shown in the illustrated example, the DCBZ_D instruction 800 includes an opcode field 804 that identifies the type of instruction (i.e., a DCBZ instruction) and operand fields 808 and 810, which specify the write permission and the write permission, respectively.The operands are used to calculate the effective target address of the memory location to be set to zero. Additionally, the DCBZ_D instruction 800 includes a distribution (D) field 806, which is set (e.g., to 1) to indicate that the data submodule to be set to zero belongs to a record that should be distributed across the various vertical cache hierarchies of the data processing system 100 and will otherwise be reset (e.g., to 0).
[0044] With the following reference to Fig. Figure 9 illustrates a logical flowchart of an exemplary procedure for initializing a memory area for storing injected data according to one embodiment. The illustrated process can be performed, for example, by executing a sequence of DCBZ_D instructions 800 by a hardware thread on one of the processor cores 200. It should be clear that the illustrated process, which is optional, increases the likelihood that subsequent cache injection requests targeting addresses in the initialized memory area will be successful.
[0045] The process of Fig. 9 begins at block 900 and then proceeds to block 902, which illustrates processor core 200 selecting the next address of a memory location to be set to zero. The selection shown in block 902 may involve advancing one or more pointers or other variable values to appropriately set the value(s) of operand fields 808 and / or 810 of a DCBZ_D instruction 800. At block 904, processor core 200 executes the DCBZ_D instruction 800, which causes processor core 200 to calculate the effective target address of a memory location to be set to zero based on operand fields 808 and 810, to translate the effective target address to obtain a real target address, and to issue a DCBZ request to the associated L2 cache 230 with a specification of the value in D field 806.In response to receiving the DCBZ request from processor core 200, the associated L2 cache 230, if necessary, obtains write permission for the real target address (e.g., by issuing one or more requests on the system structure of the data processing system 100), assigns a cache row belonging to the real target address in the L2 array 302, and sets it to zero (removing any existing cache row if necessary), and creates a corresponding entry in the L2 directory 308 with a D field 508, which is set or reset according to the value specified in the D field 806 of the DCBZ_D instruction 800.
[0046] At block 906, the hardware thread of processor core 200 executes one or more instructions to determine whether the initialization of the memory area is complete. If not, the process returns to block 902, which was previously described. However, if at block 906 it is determined that all addresses in the memory area to be initialized have been allocated and zeroed in the associated L2 cache 230, the process terminates. Fig. 9 on a block 908.
[0047] Following reference to Fig. As will be clear to a person skilled in the art, a sequence of DCBZ_D 800 instructions can be executed to initialize a collection of cache rows and prepare these cache rows as targets for subsequent cache injection requests. The target cache rows are initialized to a coherence state that designates the vertical cache hierarchy(s) that will receive the various injected cache rows and enable the cache injection request to be executed successfully. Furthermore, assuming that the D fields 508 of the DCBZ_D 800 instructions in the sequence are all set, all cache rows are marked via the D fields 806 as belonging to a single record, which, instead of being cast out to system memory 108, is to be distributed across multiple vertical cache hierarchies.Given the small size of the L2 array 302 and the L3 array 402 relative to many I / O records, marking cache rows as belonging to the record can significantly improve performance by keeping the marked cache rows in a low-latency cache instead of allowing them to be cast out to the high-latency system memory 108. For example, if the record is 16 MB in size, the capacity of an L2 array 302 is 128 KB, and the capacity of an L3 array 402 is 1 MB, initializing the cache rows in the record exceeds the capacity of a given vertical cache hierarchy, resulting in a castout. Marking the cache rows as belonging to the data record (by setting the D fields 806) causes the cache rows to be distributed across the various vertical cache hierarchies during their castout, as detailed below with reference to . Fig. 15 to 16 is described.
[0048] With the following reference to Fig. Figures 10 to 13 illustrate a logical overview flowchart of an exemplary method for injecting at least one partial cache row of write injection data into a cache at a subordinate level (e.g., an L3 cache) according to one embodiment. The process begins at a block 1000 of Fig. The process then proceeds to block 1002, which illustrates a determination of whether a cache injection request 610 on local connection 114 has been received by a pair of associated L2 and L3 caches 230 and 232. As noted above, the cache injection request might be issued, for example, by an I / O controller 214 for an attached I / O unit 218 to write a record from the I / O unit 218 into the memory hierarchy of the data processing system 100. If no cache injection request is received at block 1002, the process continues with a retry at block 1002. However, in response to receiving a cache injection request at block 1002, the process continues from block 1002 to blocks 1004 and 1006, which each illustrate determinations of whether the L3 cache 232 or the L2 cache 230 is the highest point of coherence (HPC) for the actual destination address of the write injection request.
[0049] As used herein, a lowest coherence point (LPC) is defined herein as a memory unit or I / O unit that serves as the repository for a memory block. In the absence of an HPC for the memory block, the LPC stores a faithful image of the memory block and has the authority to allow or deny requests to generate an additional cached copy of the memory block. For a typical requirement in the embodiment of the data processing system of Fig. 1 to 4, the LPC is the memory controller 206 for the system memory 108, which stores the memory block being referenced. An HPC is defined here (throughout the entire data processing system 100) as a unique entity that stores a faithful image of the memory block (which may be consistent with the corresponding memory block on the LPC) and that has the authority to allow or deny a request to modify the memory block. Thus, at most one of the L2 cache 230 and the L3 cache 232 is the HPC for the memory block belonging to a given real target address. Descriptively, the HPC can also provide a copy of the memory block to a requesting entity. Thus, for a typical memory access request in the embodiment of the data processing system of Fig. 1 to 4, if any, an L2 cache 230 or L3 cache 232. Although other indicators may be used to specify an HPC for a memory block in a preferred embodiment, the HPC, if any, for a memory block is specified by one or more cache coherence states in the directory of an L2 cache 230 or L3 cache 232.
[0050] In response to a determination at block 1004 that the L3 cache 232 of the HPC is the actual target address of the cache injection request, the process proceeds via the reference arrow A with Fig. 11 continues, which is described below. In response to a determination that the L2 cache 230 of the HPC for this real target address of the cache injection request, the process proceeds via the reference arrow B with Fig. 12 continues, which is described below. However, if there is no L2 cache 230 or L3 cache 232 of the HPC for the actual destination address of the cache injection request, the cache injection request is preferably downgraded to a DMA write request and served by the relevant memory controller 206 or the associated L4 cache 234 (if present), as shown at blocks 1008 to 1020. In this case, each L2 cache 230 or L3 cache 232 that stores a valid, shared copy of the destination cache line of the cache injection request, or is in the process of invalidating a valid copy of the destination cache line (at block 1008), responds to the cache injection request by beginning to move any modified data associated with the real destination address, if any, to system memory 108 and / or invalidating the copy of the destination cache line (at block 1010).Furthermore, the L2 cache 230 or L3 cache 232 provides a "Retry" coherence response on the system structure, indicating that the cache injection request cannot be completed successfully (at block 1014). After block 1014, the process returns from... Fig. 10 returns to block 1002, as described. It should be clear that the I / O controller 214 can reissue the cache injection request 610 on the system structure of the data processing system 100 several times until all modified data belonging to the actual target address of the cache injection request has been written back to the system memory 108 and all cache copies of the target cache line have been invalidated, resulting in a negative determination at block 1008.
[0051] Based on a negative determination at block 1008, the process continues from Fig. Block 10 continues with block 1012, which illustrates an additional determination of whether the relevant memory controller 206, to which the actual target address of cache injection request 610 or the associated L4 cache 234 (if present) has been assigned, is capable of handling cache injection request 610. If not, memory controller 206 and / or the associated L4 cache 234 provide a "Retry" coherence response on the system structure, which prevents cache injection request 610 from completing successfully (at block 1014), and the process returns to block 1002.However, if an affirmative determination is made at block 1012, the cache injection request 610 is served by the relevant L4 cache 234 (if present) or the relevant memory controller 206 by writing the permanent data storage associated with the cache injection request 610 to the L4 array 236 or the system memory 108 (at a block 1016). As specified by blocks 1018 to 1020, when the D-bits 210 are mapped in system memory 108 or a D-field 508 is mapped in L4 directory 238, the L4 cache 234 and / or the memory controller 206 additionally load the relevant D-bit 210 and / or D-field 508 with the value of D-field 616 into the cache injection request 610. Afterwards, the process returns from... Fig. 10 back to block 1002.
[0052] The following will be discussed Fig. Reference is made to reference 11, which describes in detail an embodiment of a process in the case where an L3 cache 232 was identified as the HPC of the real target address of a cache injection request 610. The process of Fig. Section 11 begins at reference arrow A and then proceeds to block 1100, which illustrates the L3 cache 232. This cache is the HPC for the actual destination address of cache injection request 610, determining whether it is currently capable of handling cache injection request 610. For example, the determination shown at block 1100 might include determining whether all resources, including a WI machine 414, required to serve cache injection request 610 are currently available for allocation to cache injection request 610. In response to a determination that the HPC L3 cache 232 is currently unable to process the cache injection request 610, the L3 cache 232 provides a "Retry" coherence response for the cache injection request 610 on the link structure (at a block 1102), requesting that the source of the cache injection request 610 (e.g.,The I / O controller (214) reissues the cache injection request at a later time. Afterwards, the process returns to block 1002 via the reference arrow D. Fig. 10 back.
[0053] By returning to block 1100, in response to a determination that the L3 cache 232 is currently capable of handling the cache injection request 610, the process branches from Fig. 11 and continues in parallel with blocks 1104 to 1108 and 1110 to 1116. At block 1104, the L3 cache 232 determines whether any shared copies of the target cache line exist in the data processing system 100, for example, by referencing the coherence state of the real target address in its directory 408 and / or a single or system-wide coherence response to the cache injection request 610. If this is not the case, the process simply rejoins the other branch of the process. However, if L3 cache 232 determines that at least one shared copy of the destination cache line may exist in data processing system 100, L3 cache 232 invalidates each shared copy(ies) of the destination cache line by issuing one or more termination requests on local connection 114 (at block 1106). Once the other copy or copies are no longer available, the L3 cache 232 will invalidate each shared copy or copies of the destination cache line by issuing one or more termination requests on local connection 114 (at block 1106).If copies of the target cache line have been invalidated (at block 1108), the process continues to the other branch of the process.
[0054] With the following reference to the other branch of the in Fig. In the process shown in Figure 11, the allocator 404 instructs the L3 array 402 via control path 426 to read the target cache line of the cache injection request (at block 1110). The target cache line is forwarded via data path 424 and multiplexer M5 to the WI buffer 422. Simultaneously, the L3 cache 232 determines at block 1112 whether the cache line containing at least part of the write injection data from the cache injection request has been received in the WIQ 420 in a persistent data storage on the local connection 114. In various embodiments, the cache injection data can be received from the source of the cache injection request simultaneously with, or at a different time than, the write injection request. If this is not the case, the process waits at block 1112 until the cache injection data has been received.Once the destination cache line at block 1110 has been read into the Wl buffer 422 and the cache injection data has been received in the WIQ 420 at block 1112, the allocator 404 controls the selection of bytes of data by the multiplexer M6 to merge the partial or complete cache line of write injection data into the destination cache line (at block 1114). The updated destination cache line is then written to the L3 array 402 via the multiplexer M7. Based on the update of the destination cache line, a WI machine 414 also writes the corresponding entry in the L3 directory 408 to the corresponding modified coherence state, which preferably indicates that the L3 cache 232 remains the HPC of the destination memory block (at block 1116). In addition, WI machine 414 updates the D-field 508 in the entry in the L3 directory 408 according to the D-field 616 of the cache injection request 610 (at a block 1120).After the termination of both branches of the in . Fig. In the process shown in Figure 11, the servicing of cache injection request 610 is complete, the WI machine 414, which is assigned to cache injection request 610, is released to return to the unoccupied state, and the process returns to block 1002 via the reference arrow D. Fig. 10 back.
[0055] With the following reference to Fig. Sections 12 to 13 describe an example of the processing that occurs when the L2 cache 230 is the HPC for the target cache line of the cache injection request 610. The processing starts at the reference arrow B and then branches out and proceeds in parallel with the one described in Fig. 12 specified process, which illustrates the processing carried out by the L3 cache 232, and via the reference arrow C with the in Fig. The process specified in section 13 continues, representing the processing carried out by the L2 cache 230.
[0056] With initial reference to Fig. 12. From reference arrow B, the process continues with block 1200, which illustrates the L3 cache 232 determining whether it is currently capable of handling cache injection request 610. For example, the determination shown at block 1200 might include determining whether all resources, including a WI machine 414, required to serve cache injection request 610 are currently available for allocation to cache injection request 610. In response to a determination at block 1200 that L3 cache 232 is currently unable to process cache injection request 610, L3 cache 232 provides a "Retry" coherence response for cache injection request 610 (at block 1202), requesting that the source of cache injection request 610 reissue another cache injection request at a later time. The process then returns to block 1002 via the reference arrow D. Fig. 10 back.
[0057] Referring again to block 1200, L3 cache 232, in response to a determination that it is currently capable of handling cache injection request 610, allocates the resources necessary to serve the request, including a WI machine 414. The allocated WI machine 414 then uses Wl channel 244 to signal to the associated L2 cache 230 that L3 cache 232 can serve cache injection request 610 by acknowledging L3 I_OK and by providing the WI ID of the allocated WI machine 414 to L2 cache 230 (at block 1204). The Wl-ID informs the L2 cache 230 which of the L3 "Done" signal lines should be monitored to determine when an L2 SN machine 314, assigned to cache injection request 610, can be released.At block 1206, the assigned WI machine 414 then determines whether the L2 cache 230 has indicated that it can also handle the cache injection request 610 by acknowledging L2 I_OK. If not, the processing of the cache injection request 610 by the L3 cache 232 ends, the WI machine 414 assigned to the cache injection request 610 is released to return to an unoccupied state, and the process returns to block 1002 via the reference arrow D. Fig. 10 back.
[0058] In response to an affirmative determination at block 1206, indicating that both L2 cache 230 and L3 cache 232 are capable of handling cache injection request 610, the process branches again and proceeds in parallel through blocks 1208 to 1212 and 1218 to 1228. At block 1208, L3 cache 232 determines whether any shared copies of the target cache row exist in data processing system 100, for example, by referencing the coherence state information provided by the associated L2 cache 230 and / or a single or system-wide coherence response to cache injection request 610. In response to a determination at block 1208 that no shared copies of the target cache line are cached in data processing system 100, the process simply rejoins the other branch of the process.However, if L3 cache 232 determines that at least one shared copy of the target cache line may be cached in data processing system 100, WI machine 414 of L3 cache 232 invalidates each shared copy(ies) of the target cache line by issuing one or more termination requests on local connection 114 (block 1210). Once the other copy(ies) of the target cache line has been invalidated (at block 1212), the process rejoins the other branch of the process.
[0059] With reference to block 1218, the L3 cache 232 determines whether the partial or complete cache line of cache injection data from cache injection request 610 has been received from local connection 114 in WIQ 420. As noted above, in various embodiments, the write injection data from the source of cache injection request 610 can be received simultaneously with or at a different time than the cache injection request 610. If this is not the case, the process waits at block 1218 until the write injection data has been received. Simultaneously, the L2 cache 230 and the L3 cache 232 work together to transfer the destination cache line of data from the L2 cache 230 to the L3 cache 232.In the illustrated embodiment, the L3 cache 232 determines, for example at block 1220, whether the L2 cache 230 has indicated, by confirmation from one of the L2 D_rdy signals of the Wl channel 244, that the target cache line has been read from the L2 array 302 into the L2 CPI buffer 318 (the relevant of the L2 D_rdy signals is identified by the SN ID provided by the L2 cache 230 at a block 1304, as described below). If not, the process is repeated at block 1220. In response to the L2 cache 230 indicating that the destination cache line has been read into the L2 CPI buffer 318, the SN machine 314 of the L2 cache 230, which is assigned to handle the cache injection request 610, causes the destination cache line to be transferred from the L2 CPI buffer 318 to the Wl buffer 422 of the L3 cache 232 via the CI channel 242, the data path 426 and the multiplexer M5 (block 1222).In response to receiving the destination cache line in WI buffer 422, WI machine 414, which is assigned to handle cache injection request 610, acknowledges the corresponding one of the L3 "done" signals over Wl channel 244 (at block 1224), releasing SN machine 314, which is assigned to handle cache injection request 610, to return to an unoccupied state in which it is available for assignment to a subsequent request that was spied on local link 114.
[0060] After the completion of the process shown in blocks 1218 and 1224, the allotment 404 controls the selection of bytes of data by the multiplexer M6 to merge the partial or complete cache line of write injection data into the destination cache line (block 1226). The updated destination cache line is then written to the L3 array 402 via the multiplexer M7. Based on the update of the destination cache line, the allotment 404 also writes the corresponding entry in the L3 directory 408 to the appropriate modified coherence state, which preferably indicates that the L3 cache 232 is the HPC of the destination memory block (at block 1228). Furthermore, the allocator 404 updates the D-field 508 in the entry in the L3 directory 408 according to the D-field 616 of the cache injection request 610 (at a block 1230).After the processing, illustrated in blocks 1230 and 1208 / 1212, is completed, the processing of cache injection request 610 is finished, the WI machine 414, which is assigned to cache injection request 610, is released to return to the unoccupied state, and the in . Fig. 12. The illustrated process returns to block 1002 via the reference arrow D. Fig. 10 back.
[0061] With the following reference to Fig. 13. The process continues from reference arrow C with a block 1300, which illustrates the L2 cache 230 determining whether it is currently capable of handling cache injection request 610. For example, the determination shown at block 1300 might include determining whether all resources, including an SN machine 314, required to serve cache injection request 610 are currently available for allocation to cache injection request 610. In response to a determination at block 1300 that L2 cache 230 is currently unable to process cache injection request 610, L2 cache 230 provides a "Retry" coherence response for cache injection request 610 (at block 1302), requesting that the source of cache injection request 610 reissue the request at a later time. The process then terminates. Fig. 13 at block 1320. It should be recalled that in this case, the L2 cache 230 does not acknowledge L2 I_OK to indicate that it can process the cache injection request 610, which also causes the associated L3 cache 232 to stop processing the cache injection request 610, as discussed above with reference to block 1206.
[0062] Referring again to block 1300, L2 cache 230, in response to a determination that it is currently capable of handling cache injection request 610, allocates the resources necessary to serve the request, including an SN machine 314. The allocated SN machine 314 then uses Wl channel 244 to signal to the associated L3 cache 232 that L2 cache 230 can serve the cache injection request 610 by acknowledging L2 I_OK and by providing the SN ID of the allocated SN machine 314 to L3 cache 232 (at block 1304). The SN ID provided by L2 cache 230 identifies which of the L2 D_rdy signals the WI machine 414 at block 1220 of Fig. 12 is monitored. At block 1306, SN machine 314 then determines whether the associated L3 cache 232 has indicated that it can also handle the cache injection request 610 by confirming L3 I_OK. If this is not the case, SN machine 314, which is assigned to handle the cache injection request 610, is released to return to an unoccupied state, and the process of Fig. 13 ends at block 1320.
[0063] In response to a confirmatory determination at block 1306, indicating that both the L2 cache 230 and the L3 cache 232 are capable of processing cache injection request 610, the process branches and continues in parallel through blocks 1310 and 1312 to 1316. At block 1310, the SN machine 314 of the L2 cache 230 updates the entry in the L2 directory 308, which corresponds to the target cache line of cache injection request 610, to an invalid coherence state. Additionally, at block 1312, the allotment 305 instructs the L2 array 302, via control path 330, to read the target cache line into the L2 CPI buffer 318.In response to the destination cache line being placed in the L2 CPI buffer 318, the SN machine 314 acknowledges the corresponding L2 D_rdy signal at block 1314 to indicate to the L3 WI machine 414 in the associated L3 cache 232 that the destination cache line of data is ready for transfer to the L3 cache 232 (see, for example, block 1220 of ). Fig. 12) The SN machine 314 then waits for acknowledgment from the L3 WI machine 414 of the associated L3 "done" signal to indicate that the destination cache line has been successfully transferred from the L2 CPI buffer 318 to the L3 cache 232 (at block 1316). In response to the completion of processing, represented in blocks 1310 and 1316, the assignment to SN machine 314 is reversed (it returns to an unassigned state), and the process of Fig. 13 ends at block 1320.
[0064] With the following reference to Fig. Figure 14 presents a logical flowchart of an exemplary procedure for processing a DMA write request 600 according to one embodiment. The process begins at block 1400 and then proceeds to block 1402, which illustrates a determination of whether a DMA write request 600 has been issued on the system structure of the data processing system 100. As noted above, the DMA write request 600 can be issued, for example, by an I / O controller 214 to an attached I / O unit 218 to write a portion of a data record from the I / O unit 218 to a system memory 108 or the L4 cache 234 of the data processing system 100. If no DMA write request is issued at block 1402, the process continues with a retry at block 1402. However, in response to an output of a DMA write request 600 at block 1402, the process continues with block 1404.
[0065] Block 1404 illustrates the various actions taken in the data processing system 100 based on whether any cache 230, 232, or 234 stores a valid copy of the destination cache line of the DMA write request 600 or is in the process of invalidating a valid copy of the destination cache line. If so, each cache that stores a valid copy of the destination cache line or is in the process of invalidating a copy of the destination cache line of the DMA write request 600 responds to the DMA write request by beginning to push any modified data, if any, belonging to the actual destination address into the relevant system memory 108 and / or invalidating the copy of the destination cache line (at a block 1410).Furthermore, each of these caches provides a "Retry" coherence response on the system structure, indicating that the DMA write request could not be completed successfully (at block 1411). After block 1411, the process returns from... Fig. 14 returns to block 1402, as described. It should be clear that the I / O controller 214 can reissue the DMA write request 600 on the system structure of the data processing system 100 multiple times until all modified data belonging to the actual destination address of the DMA write request 600 has been written back to the relevant system memory 108 and all cache copies of the destination cache line have been invalidated, resulting in a negative determination at block 1404.
[0066] Based on a negative determination at block 1404, the process continues from Fig. Block 14 continues with block 1412, which illustrates an additional determination of whether the relevant memory controller 206, to which the actual destination address of the DMA write request 600 or the associated L4 cache 234 (if present) has been assigned, is capable of handling the DMA write request 600. If not, the memory controller 206 and / or the associated L4 cache 234 provide a "Retry" coherence response on the system structure, which prevents the DMA write request 600 from completing successfully (at block 1414), and the process returns to block 1402.However, if an affirmative determination is made at block 1412, the DMA write request 600 is served by the relevant L4 cache 234 (if present) or the relevant memory controller 206 by writing the permanent data storage associated with the DMA write request 600 to the L4 array 236 or the system memory 108 (at a block 1416). As specified by blocks 1418 to 1420, when the D bits 210 are mapped in system memory 108 or a D field 508 is mapped in L4 directory 238, the L4 cache 234 and / or the memory controller 206 additionally load the relevant D bit 210 and / or the D field 508 with the value of D field 616 into the DMA write request 600. The process then returns from... Fig. 14 back to block 1402.
[0067] With the following reference to Fig. Section 15 illustrates a logical overview flowchart of an exemplary procedure by which a cache of a higher level (e.g., L2) performs a castout according to one embodiment. The process of Fig. Process 15 begins at block 1500 and then continues at block 1502, for example, in response to an L2 cache 230 that determines to remove a victim cache row from L2 array 302 to accommodate a new cache row to be installed in L2 array 302. The process then continues at block 1502, which illustrates the L2 cache 230, determining whether to cast the victim cache row to the associated L3 cache 232, or instead to the L4 cache 234 or system memory 108. In at least some embodiments, the L2 cache 230 can perform the determination shown at block 1502 based on one or more criteria, including the D field 508 of the directory entry in the L2 directory 308 for the victim cache line.For example, the decision at block 1502 may be more likely to cast the Victim Cache line into the associated L3 Cache 232 if the D field 508 is set, and less likely to cast the Victim Cache line into the associated L3 Cache 232 if the D field 508 is reset.
[0068] In response to a determination at block 1502 to cast the Victim Cache line to the associated L3 Cache 232, the L2 Cache 230 issues a CO request 700 to the associated L3 Cache 232 with the value of the associated D field 508 from the entry in the L2 directory 308 in the D field 706 of the CO request 700. In response to a determination at block 1502 that the victim cache line should be cast to an L4 cache 234 (if available) or system memory 108, the L2 cache 230 issues a CO request 700 to the relevant L4 cache 234 or memory controller 206 via the system structure with the value of the associated D field 508 from the entry in the L2 directory 308 in the D field 706 of the CO request 700 (at block 1510).In response to receiving the CO request 700, the L4 cache 234 (if present) or the memory controller 206 receives and stores the associated victim cache line. If the L4 cache 234 is present, it should be clear that another cache line can be removed from the L4 directory 238 to store the victim cache line. As specified by blocks 1512 to 1514, when the D-bits 210 in system memory 108 are mapped or a D-field 508 in L4 directory 238 is mapped, the L4 cache 234 and / or the memory controller 206 additionally load the relevant D-bit 210 and / or the D-field 508 with the value of D-field 706 into the CO request 700. The process terminates either after block 1504 or blocks 1512 to 1514. Fig. 15 on one block 1520.
[0069] With the following reference to Fig. Figure 16 presents a logical overview flowchart of an exemplary procedure by which a subordinate level cache (e.g., L3) processes a castout received from a superior level cache (e.g., L2) according to an embodiment. The process of Fig. 16 begins at a block 1600, for example in response to a CO request 700 being received by an L3 cache 232, which is returned by the associated L2 cache 230, for example at block 1504 of Fig. 15 is output. In response to receiving CO request 700, the L3 cache 232 determines at block 1602 whether the D field 706 of CO request 700 is set. If not, processing of CO request 700 continues with block 1609 and subsequent blocks, which are described below. However, if the L3 cache 232 determines at block 1602 that the D field 706 of CO request 700 is set, the L3 cache 232 additionally determines at block 1604 whether the state field 710 of CO request 700 indicates an HPC coherence state for the victim cache line. If this is not the case, the L3 cache 232 resets the D field, and the process moves to block 1609, which is described below.
[0070] However, if the L3 cache 232 determines at block 1604 that the state field 710 of CO request 700 indicates an HPC coherence state, the L3 cache 232 also determines at block 1606 whether a lateral castout (LCO) to another L3 cache 232 in the data processing system 100 should occur for the victim cache row of data received from the associated L2 cache 230. As noted above, it is desirable for cache rows belonging to large datasets injected into the cache hierarchy of the data processing system 100 (identified by a set D field) to be distributed across multiple vertical cache hierarchies, rather than being restricted to the vertical cache hierarchy of a single processor core 200. Accordingly, in at least one embodiment, the L3 cache 232 of the data processing system 100 are each assigned to an LCO group that has N (e.g. 4, 8 or 16) L3 cache 232 as members.At block 1606, for example, the L3 cache 232 can decide to perform an LCO on the victim cache line received by the associated L2 cache 230 for N-1 / N of the CO requests 700 received by the associated L2 cache 230, and not to perform an LCO on the victim cache line for 1 / N of the CO requests 700 received by the associated L2 cache 230, in order to distribute the victim cache lines evenly across the LCO group. In response to the L3 cache 232 deciding at block 1606 not to perform a CLO on the victim cache line, the process transitions to block 1610, as described below. If, on the other hand, L3 cache 232 determines at block 1606 to perform an LCO for the victim cache line, L3 cache 232 selects a target L3 cache 232 (e.g.pseudorandomly from the other L3 cache 232 in the same LCO group) and issues an LCO request 720 to the target L3 cache 232 via the system structure (at a block 1608). In LCO request 720, the valid field 722 is set, the ttype field 724 specifies an LCO, the D field 726 is set, the address field 728 specifies the actual address of the victim cache row received in address field 708 of CO request 700, the state field 730 specifies the coherence state indicated by state field 710 of CO request 700, and the target ID field 732 identifies the target L3 cache 232. An additional permanent data store on the system structure transfers the data from the victim cache row to the target L3 cache 232. It should be noted that the source L3 cache 232 does not install the victim cache row in its L3 array 402. The process ends after block 1608. Fig. 16 on one block 1620.
[0071] Referring to block 1609, the L3 cache 232 determines whether to perform an LCO on the victim cache row. As noted above, the L3 cache 232 might, for example, determine to perform an LCO on the victim cache row received from the associated L2 cache 230 based on a statement provided by the heuristic LCO logic 405. In response to an affirmative determination at block 1609, the process proceeds to block 1608, which has been written. In response to a negative determination at block 1609, the L3 cache 232 removes another cache row from the L3 array 402, if necessary, to make room for the victim cache row received from the associated L2 cache 230 (at block 1610).At block 1612, the L3 cache 232 writes the victim cache, received in connection with the CO request 700, to the L3 array 402, creates the corresponding entry in the L3 directory 408, and loads the corresponding value (received from the D field 706) into the D field 508 of the entry in the L3 directory 408. After block 1612, the process ends. Fig. 16 on one block 1620.
[0072] With the following reference to Fig. Figure 17 illustrates a logical overview flowchart of an exemplary procedure by which a cache of a subordinate level (e.g., L3) performs a castout according to one embodiment. The illustrated process can be applied, for example, to a block 1610 of Fig. 16 will be carried out.
[0073] The process of Fig. Process 17 begins at block 1700, for example, based on a determination by an L3 cache 232 that a cache row in its L3 array 402 needs to be removed. The process continues with block 1702, which illustrates the L3 cache 232 determining whether removing the cache row is necessary via an LCO request 720 (which, for example, is made at block 1608 of Fig. 16 is issued), which targets this L3 cache 232. If this is the case, a CO request 700 is used to transfer the L4 cache 234 and / or system memory 108, and the process moves to block 1710 and the following blocks, as described below. If the L3 cache 232 determines at block 1702 that removing the cache line by an LCO request 720 is not necessary, the L3 cache 232 determines at block 1704 whether to perform an LCO of the victim cache line into another L3 cache 232, or a CO of the victim cache line into the L4 cache 234 or system memory 108. The allocator 404 of the L3 cache 232b can perform the determination illustrated at block 1704 based on the information provided by the heuristic LCO logic 405.In response to a request at block 1704 to perform an LCO for the victim cache line, the L3 cache 232 selects a target L3 cache 232 (e.g., pseudorandomly from the other L3 cache 232 in the same LCO group) and issues an LCO request 720 on the system structure (at a block 1706). In LCO request 720, the valid field 722 is set, the ttype field 724 specifies an LCO, the D field 726 has the value in the D field 508 of the relevant entry in the L3 directory 408, the address field 728 specifies the actual address of the victim cache row that was removed from the L3 cache 232, the state field 730 specifies the coherence state, which is indicated by the state field 506 of the relevant entry in the L3 directory 408, and the target ID field 732 identifies the target L3 cache 232. After block 1706, the process ends. Fig. 17 on one block 1720.
[0074] With reference to block 1710, the L3 cache 232 issues a CO request 700 to the relevant L4 cache 234 or the system memory 108 via the system structure, using the value of the corresponding D-field 508 from the entry in the L3 directory 408 in the D-field 706 of the CO request 700. The data of the victim cache line is transferred via the system structure to an additional permanent data storage. In response to CO request 700 and the data from the victim cache line, the L4 cache 234 or the memory controller 206 writes the victim cache line data to the L4 array 236 or system memory 108. If the L4 cache 234 is present, it should be clear that another cache line can be removed from the L4 directory 238 to make room for storing the victim cache line.
[0075] As further specified by blocks 1712 to 1714, when the D-bits 210 are mapped in system memory 108 or a D-field 508 is mapped in L4 directory 238, the L4 cache 234 and / or the memory controller 206 additionally load the relevant D-bit 210 and / or the D-field 508 with the value of D-field 706 into the CO request 700. The process ends either after block 1712 or block 1714. Fig. 17 on one block 1720.
[0076] With the following reference to Fig. Figure 18 illustrates a logical overview flowchart of an exemplary procedure by which a lowest-level cache (e.g., L4), if present, performs a castout according to one embodiment. The process of Fig. 18 begins at a block 1800, for example based on a determination by an L4 cache 234 that a cache row must be removed from its L4 array 236, for example due to a CO request issued at block 1510 or block 1710.
[0077] The process continues from block 1800 to block 1802, which illustrates the L4 cache 234. This cache demonstrates various actions based on whether the L4 directory 238 implements the D field 508. If it does not, the L4 cache 234 issues a CO request 700 (optionally omitting the D field 706) to the memory controller 206 of the associated system memory 108 to cause the removed cache line (which is transferred to a separate permanent data store) to be written to system memory 108 (at block 1810). Afterward, the L4 cache 234 invalidates its copy of the victim cache line in the L4 directory 238 (at block 1808), and the process continues. Fig. 18 ends at a block 1812.
[0078] Referring again to block 1802, when the entries of L4 directory 238 translate D-fields 508 and system memory 108 translates the D-bits 210, the L4 cache 234 issues a CO request 700 to the memory controller 206 of the associated system memory 108 with the value of D-field 508 from the entry in L4 directory 238 in D-field 706 of the CO request 700 (at block 1806). The data of the victim cache line is transferred to memory controller 206 into an additional persistent data store. Afterward, the L4 cache 234 invalidates its copy of the victim cache line in L4 directory 238 (at block 1808), and the process of Fig. 18 ends at a block 1812.
[0079] With the following reference to Fig. Figure 19 illustrates a logical flowchart of an exemplary procedure by which a higher-level cache (e.g., L2) retrieves data into its array according to one embodiment. The process begins at block 1900, for example, in response to an L2 cache 230 determining that a cache row of data requested by the associated processor core 200 is not present in a required state of coherence in either this L2 cache 230 or the associated L3 cache 232. The process then proceeds at block 1902, illustrating an L2 cache 230 issuing a memory access request (e.g., a read request or a read request with intended modification) to the system structure to retrieve an existing copy of a target cache row into the L2 array 302.In response to a memory access request (or a reissuance of the memory access request), the L2 cache 230 in a persistent data store on the system structure receives a cache row of data specified by a real target address in the memory access request. In various operating scenarios, the cache row of data can originate, for example, from the L4 cache 234 (if present), the memory controller 206, an L2 cache 230, or an L3 cache 232. Depending on the data source and whether D fields / D bits are implemented in the L4 cache 234 / system memory 108, the persistent data store or a coherence message associated with the memory access request can transmit a setting of a D field corresponding to the requested cache row of data.
[0080] At block 1904, the requesting L2 cache 230 determines whether the cache line data has been received in conjunction with a D bit. If not, the L2 cache 232 installs the cache line of data in the L2 array 302, creates a corresponding entry in the L2 directory 308, and resets the D field 508 in the directory entry (at blocks 1910 and 1912). The process then terminates. Fig. 19 on one block in 1914.
[0081] Referring again to block 1904, and in response to a determination that a D-bit has been received in conjunction with the cache line data, the L2 cache 230 additionally determines at block 1906 whether the D-bit is set. If not, the process proceeds to blocks 1910 to 1912, which have been described. However, if the L2 cache 230 determines at block 1906 that the D-bit received in conjunction with the cache line data has been set, the L2 cache 232 installs the cache line of data in the L2 array 302, creates a corresponding entry in the L2 directory 308, and sets the D-field 508 in the directory entry (at blocks 1908 and 1912). The process then terminates. Fig. 19 on one block in 1914.
[0082] With the following reference to Fig. Figure 20 presents a block diagram of an exemplary Design Flow 2000, which is used, for example, for the design of semiconductor IC logic, simulation, testing, layout, and fabrication. The Design Flow 2000 comprises processes, machines, and / or mechanisms for processing design structures or units to generate logically or otherwise functionally equivalent representations of the design structures and / or units described and shown herein. The design structures processed and / or generated by the Design Flow 2000 can be encoded on machine-readable transmission or storage media to include data and / or instructions that, when executed or otherwise processed on a data processing system, generate a logically, structurally, mechanically, or otherwise functionally equivalent representation of hardware components, circuits, units, or systems.The machines include, but are not limited to, all machines used in an IC design process, such as those used to design, manufacture, or simulate a circuit, component, unit, or system. For example, the machines may include: lithography machines, machines and / or equipment for generating masks (e.g., electron beam writers), computers or equipment for simulating design structures, any device used in the manufacturing or testing process, or any machine for programming functionally equivalent representations of the design structures into any medium (e.g., a machine for programming a programmable gate array).
[0083] The Design Flow 2000 can vary depending on the type of representation being developed. For example, a Design Flow 2000 for creating an application-specific integrated circuit (ASIC) may differ from a Design Flow 2000 for developing a standard component or from a Design Flow 2000 for instantiating the design in a programmable array, such as a programmable gate array (PGA) or a field-programmable gate array (FPGA) developed by Altera. ® Inc. or Xilinx ® is offered by Inc.
[0084] Fig.Figure 20 illustrates several such design structures, including an input design structure 2020, which is preferably processed by a design process 2010. The design structure 2020 can be a logical simulation design structure that is generated and processed by the design process 2010 to produce a logically equivalent functional representation of a hardware unit. Alternatively, the design structure 2020 can also include data and / or program instructions that, when processed by the design process 2010, generate a functional representation of the physical structure of a hardware unit. Regardless of whether the design structure 2020 represents functional and / or structural design features, it can be generated using electronic computer-aided design (ECAD), as implemented by a core developer / designer.When the Design Structure 2020 is encoded on a machine-readable data transmission, gate array, or storage medium, it can be accessed and processed by one or more hardware and / or software modules in the Design Process 2010 to simulate or otherwise functionally represent an electronic component, circuit, electronic or logic module, device, unit, or system, as shown, for example, herein. The Design Structure 2020 may therefore include files or other data structures, including human- and / or machine-readable source code, compiled structures, and computer-executable code structures, which, when processed by a design or simulation data processing system, functionally simulate or otherwise represent circuits or other levels of hardware logic design.Such data structures may include Hardware Description Language (HDL) design entities or other data structures that are conformal to and / or compatible with lower-level HDL design languages, such as Verilog and VHDL, and / or higher-level design languages such as C or C++.
[0085] The Design Process 2010 preferentially uses and integrates hardware and / or software modules to synthesize, translate, or otherwise process a design / simulation that is functionally equivalent to the components, circuits, units, or logic structures shown herein, in order to generate a Netlist 2080 that may contain design structures such as Design Structure 2020. The Netlist 2080 may, for example, contain compiled or otherwise processed structures that represent a list of wires, discrete components, logic gates, control circuits, I / O units, models, etc., describing the connections to other elements and circuits in the design of an integrated circuit.The Netlist 2080 can be synthesized using an iterative process in which the Netlist 2080 is resynthesized one or more times, depending on the design specifications and parameters for the unit. Just as with other design structure types described herein, the Netlist 2080 can be recorded on a machine-readable storage medium or programmed into a programmable gate array. The medium can be non-volatile storage, such as a magnetic or optical disk drive, a programmable gate array, compact flash memory, or other flash memory. Alternatively, or in addition, the medium can be system or cache memory, or a buffer memory.
[0086] The Design Process 2010 can include hardware and software modules for processing a variety of input data structure types, including the netlist 2080. Such data structure types can be found, for example, in library elements 2030 and comprise a set of commonly used elements, circuits, and units, including models, layouts, and symbolic representations for a particular manufacturing technology (e.g., various technology nodes, 32 nm, 45 nm, 200 nm, etc.). The data structure types can further include design specifications 2040, characterization data 2050, verification data 2060, design rules 2070, and test data files 2085, which may include input test patterns, output test results, and other test information.The Design Process 2010 may further include, for example, standard mechanical design processes such as stress analysis, thermal analysis, mechanical event simulation, and process simulation for operations such as casting, molding, and compression molding. A person skilled in the art of mechanical design can recognize the extent of possible mechanical design tools and applications used in the Design Process 2010 without deviating from the scope of protection of the invention. The Design Process 2010 may also include modules for performing standard circuit design processes, such as timing analysis, verification, design rule checking, and the placement and routing of operations.
[0087] The Design Process 2010 uses and integrates logical and physical design tools, such as HDL compilers and simulation model build tools, to process Design Structure 2020, along with some or all of the supporting data structures shown, and any additional mechanical design or data (where applicable), to generate a second Design Structure 2090. Design Structure 2090 resides on a storage medium or programmable gate array in a data format suitable for exchanging data between mechanical units and structures (e.g., information stored in IGES, DXF, Parasolid XT, JT, DRG, or any other suitable format for storing or reproducing such mechanical design structures).Similar to design structure 2020, design structure 2090 preferably comprises one or more files, data structures, or other computer-encoded data or instructions located on transmission or data storage media, which, when executed by an ECAD system, generate a logically or otherwise functionally equivalent form of one or more of the embodiments of the invention shown herein. In one embodiment, design structure 2090 may include a compiled executable HDL simulation model that functionally simulates the units shown herein.
[0088] The Design Structure 2090 can also use a data format used for exchanging layout data for integrated circuits and / or a symbolic data format (e.g., information stored in a GDSII (GDS2), HL1, OASIS, map files, or any other suitable format for storing such design data structures). The Design Structure 2090 can contain information such as symbolic data, map files, test data files, design content files, manufacturing data, layout parameters, wires, metal planes, vias, shapes, data for routing through the manufacturing line, and any other data requested by a manufacturer or another designer / developer to create a unit or structure as described above and shown herein.The design structure 2090 can then continue with a stage 2095, at which the design structure 2090 continues, for example, with a tape-out, is released for manufacturing, is released for a mask house, is sent to another design house, is sent back to the customer, etc.
[0089] As described, a data processing system in at least one embodiment comprises a plurality of processor cores, each supported by a plurality of vertical cache hierarchies. Based on a cache injection request received on a system structure, requesting the injection of data into a cache row identified by a real target address, the data is written to a cache in a first vertical cache hierarchy of the plurality of vertical cache hierarchies. Based on a value in a field of the cache injection request, a distribution field is set in a directory entry of the first vertical cache hierarchy. After the cache row is removed from the first vertical cache hierarchy, it is determined whether the distribution field is set.Based on a determination that the distribution field is set, a lateral castout of the cache row from the first vertical cache hierarchy to a second vertical cache hierarchy is performed.
[0090] Although various embodiments have been shown and described in detail, it will be clear to a person skilled in the art that various modifications in form and detail can be made to them without deviating from the scope of protection of the claims in the appendix, and that these alternative implementations as a whole fall within the scope of protection of the claims in the appendix. For example, although aspects relating to a computer system that executes program code controlling the functions of the present invention have been described, it should be clear that the present invention can alternatively be implemented as a program product with a computer-readable storage unit that stores program code which can be processed by the data processing system.The computer-readable storage unit may include volatile or non-volatile memory, an optical or magnetic disk, or the like, but does not exclude legally regulated items such as the propagation of signals per se, transmission media per se, and forms of energy per se.
[0091] For example, the program product may include data and / or instructions that, when executed or otherwise processed on a data processing system, generate a logically, structurally, or otherwise functionally equivalent representation (including a simulation model) of hardware components, circuits, units, or systems disclosed herein. Such data and / or instructions may include hardware description language (HDL) design entities or other data structures that are conformal to and / or compatible with lower-level HDL design languages, such as Verilog and VHDL, and / or higher-level design languages, such as C or C++. Furthermore, the data and / or instructions may also use a data format suitable for exchanging layout data of integrated circuits and / or a symbolic data format (e.g.,Information stored in a GDSII (GDS2), HL1, OASIS, map files or any other suitable format for storing such design data structures) is used.
Claims
[1] Method for data processing in a data processing system (100) comprising a plurality of processor cores (200a, 200b) each supported by a plurality of vertical cache hierarchies, wherein the method comprises: based on the fact that a cache injection request is received on a system structure of the data processing system (100), Requesting an injection of data into a cache line that is identified by a real target address, Writing the data to a cache in a first vertical cache hierarchy from the plurality of vertical cache hierarchies; based on a value in a field of the cache injection request, Setting a distribution field in a directory entry of the first vertical cache hierarchy; after removing the cache row from a cache memory in the first vertical cache hierarchy, Determine whether the distribution field is set; and based on determining that the distribution field is set, Performing a lateral castout of the cache row from the first vertical cache hierarchy to a second vertical cache hierarchy from the plurality of vertical cache hierarchies. [2] Method according to claim 1, wherein performing a lateral castout based on determining that the distribution field is set comprises performing a lateral castout based on determining that the distribution field is set only if the cache line is stored by the cache working memory in a coherence state that provides write permission. [3] The method of claim 1, wherein the method further comprises Removing the cache row from a cache of a parent level in the first vertical cache hierarchy; Performing a lateral castout involves retrieving the cache row from a child level in the first vertical cache hierarchy after it has been removed from the parent level's cache and issuing a lateral castout request to the second vertical cache hierarchy without installing the cache row into a data array of the child level's cache. [4] Method according to claim 1, wherein one processor core (200) from the plurality of processor cores (200) is supported by the first vertical cache hierarchy; and the procedure further exhibiting before receiving the cache injection request at the first vertical cache hierarchy, Execute, by the processor core (200), an instruction to cause the cache row in the first vertical cache hierarchy to be installed in a coherence state that provides write permission. [5] Method according to claim 1, wherein performing the lateral castout comprises: Transferred, in a lateral castout request (700) on the system structure of the data processing system, a distribution field that is set; Transferring the cache line to the system structure; and Installing the cache row in a data array in the second vertical cache hierarchy and, based on the fact that the distribution field is set in the lateral castout requirement, Setting a distribution field in a directory entry of the second vertical cache hierarchy. [6] Method according to claim 1, wherein the data processing system (100) has a system working memory (108); the procedure further exhibiting subsequent castout of the data from the cache row of the second vertical cache hierarchy to the system memory (108) and Storing the distribution field in the system memory (108) in conjunction with the data. [7] Processing unit (104a, ..., 104d) for a data processing system, comprising: one processor core (200); a vertical cache hierarchy connected to the processor core (200) and is configured to be connected to a system structure of the data processing system (100), wherein the vertical cache hierarchy has a cache with a data array and a directory and is configured to perform: based on the fact that a cache injection request is received on the system structure of the data processing system, Requesting an injection of data into a cache line that is identified by a real target address, Writing the data to the data array in the vertical cache hierarchy; based on a value in a field of the cache injection request, setting a distribution field in a directory entry of the directory; After removing the cache row from the first vertical cache hierarchy, determine if the distribution field is set; and Based on determining that the distribution field is set, perform a lateral castout of the cache row from the vertical cache hierarchy to another vertical cache hierarchy in the data processing system (100). [8] Processing unit (104a, ..., 104d) according to claim 7, wherein performing a lateral castout based on determining that the distribution field is set comprises performing a lateral castout based on determining that the distribution field is set only if the cache row is stored by the cache memory in a coherence state that provides write permission. [9] Processing unit (104a, ..., 104d) according to claim 7, wherein the cache is a cache of a higher level; the vertical cache hierarchy has a cache of a subordinate level; Performing the lateral castout involves a child level cache that receives the cache row after it has been removed from the parent level cache and issues a lateral castout request directed to the further vertical cache hierarchy without installing the cache row in the child level cache. [10] Processing unit according to claim 7, wherein the processor core (200), prior to receiving the cache injection request, executes an instruction to cause the cache row to be installed in the first vertical cache hierarchy in a coherence state that provides write permission. [11] Processing unit (104a, ..., 104d) according to claim 7, comprising performing the lateral castout: Transferring the cache line to the system structure; and Transferred, in a lateral castout request (700) on the system structure of the data processing system (100), a distribution field that is set, wherein the set distribution field causes a distribution field to be set in a directory entry that is associated with the cache row in the second vertical cache hierarchy. [12] Data processing system (100), comprising: a plurality of processing units (104a, ..., 104d) according to claim 7; a connection structure that connects the majority of processing units (104a, ..., 104d); and a system memory (108) that is connected to the connection structure for data exchange. [13] Data processing system (100) according to claim 12, further comprising a memory controller (206) which, in response to a castout of the data of the cache row from the second vertical cache hierarchy, installs the data in the system memory (108) and stores the distribution field in the system memory (108) in conjunction with the data. [14] Design structure, which is concretely embodied in a machine-readable storage medium for the development, manufacture or testing of an integrated circuit, having the design structure a processing unit (104a, ..., 104d) according to one of claims 7-11. [15] Design structure according to claim 14, wherein the design structure comprises a design structure of a hardware description language (HDL).
Citation Information
Patent Citations
Selective storage of data in levels of a cache memory
US20080059707A1
Victim cache lateral castout targeting
US20100235577A1
Mode-Based Castout Destination Selection
US20100262783A1
Selective cache-to-cache lateral castouts
US20120203973A1