Memory Pipeline Control in a Hierarchical Memory System
By introducing a bypass path mechanism in the memory system, the problem of inefficient write operations is solved, faster write confirmation and consistency management are achieved, and the overall performance of the memory system is improved.
Patent Information
- Application Number
- CN202080038514.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-05-20
- Filing Date
- 2020-05-26
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2040-05-26
AI Technical Summary
In the prior art, memory systems have problems of low efficiency and high latency in writing operation management, especially in hierarchical memory systems, when write requests need to pass through multiple levels of cache, it is easy to cause conflict and consistency problems.
The bypass path mechanism is introduced, allowing certain write requests to bypass some memory pipelines, connect directly to the lower-level memory controller through a higher-level memory controller, determine whether it is a bypass write through the initial scheduling block, and send it immediately upon write acknowledgment to ensure the order and consistency of the write requests.
Improves the efficiency of write operations, reduces write delays, ensures overall performance and consistency of the memory system, allows the processor core to release pauses faster, and improves the system's processing capabilities.
Smart Images

Figure CN113924558B_ABST
Abstract
Description
[0001] This description generally relates to a processing device that can be formed as part of an integrated circuit such as a system-on-chip (SoC). More specifically, this description relates to such systems with improved management of write operations. Background Art
[0002] An SOC is an integrated circuit on a single die having multiple functional blocks such as one or more processor cores, memories, and inputs and outputs.
[0003] Hierarchical memory moves data and instructions between memory blocks having different read / write response times for respective processor cores such as a central processing unit (CPU) or a digital signal processor (DSP). For example, memory that is more local to a respective processor core typically has a lower response time. Hierarchical memory includes a cache memory system having multiple levels such as L1 and L2, where different levels describe different degrees of locality or different average response times of the cache memory to the respective processor cores. Here, a more local or lower response time cache memory (such as an L1 cache) is referred to as a higher level cache memory than a less local or higher response time lower level cache memory (such as an L2 cache or an L3 cache). The associativity of a cache refers to cache storage separation, where set associativity divides the cache into a number of storage groups, and each such group stores a number (way) of blocks, while a fully associative cache is not constrained by group limitations. Thus, for an integer N, each location in the main memory (system memory) can reside in any one of N possible locations in an N-way associative cache.
[0004] A "victim cache" memory cache stores data (such as cache lines) evicted from a cache memory (such as an L1 cache). If an L1 cache read results in a miss (data corresponding to a portion of the main memory is not stored in the L1 cache), then a lookup occurs in the victim cache. If the victim cache lookup results in a hit (data corresponding to the requested memory address exists in the victim cache), then the contents of the victim cache location that produced the hit are swapped with the contents of the corresponding location in the respective cache (in this instance, the L1 cache). Some example victim caches are fully associative. Data corresponding to any location in the main memory can be mapped to (stored in) any location in the fully associative cache. Summary of the Invention
[0005] In the described example, a processor system includes a processor core that generates memory transactions, a lower-level cache memory with a lower memory controller, and a higher-level cache memory with a higher memory controller having a memory pipeline. The higher memory controller is connected to the lower memory controller via a bypass path that skips the memory pipeline. The higher memory controller: determines whether a memory transaction is a bypass write, which is a memory write request that is indicated as not causing the corresponding write to be directed to the higher-level cache memory; if the memory transaction is determined to be a bypass write, then determines whether a memory transaction that prevents propagation is in the memory pipeline; and if it is determined that there is no transaction in the memory pipeline that prevents propagation, then sends the memory transaction to the lower memory controller using the bypass path. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Figure 1 is a block diagram of an example processor that is part of a SoC.
[0007] Figure 2 is for Figure 1 a block diagram of an example memory pipeline of a SoC.
[0008] Figure 3 is for Figure 1 an example of a process for a memory write operation of a SoC.
[0009] Figure 4 is for Figure 1 a block diagram of an example memory pipeline of a SoC.
[0010] Figure 5 is for Figure 1 an example of a process for a memory write operation of a SoC. DETAILED DESCRIPTION
[0011] Figure 1is a block diagram of an example processor 100 that is part of the SoC 10. The SoC 10 includes a processor core 102 that generates new data, such as a CPU or DSP. The processor 100 may include a clock 103, which may be part of the processor core 102 or may be separate from the processor core 102 (a separate clock not shown). The processor core 102 also generates memory read requests that request reads from the data memory controller 104 (DMC) and the streaming engine 106, as well as memory write requests that request writes to the data memory controller 104 (DMC) and the streaming engine 106. In some embodiments, the processor core 102 generates one read request or write request per processor core clock cycle. The memory controller returns a write acknowledgment to the processor core 102 to confirm that the requested memory write has been performed. The processor core 102 is also coupled to receive instructions from the program memory controller 108 (PMC). The streaming engine 106 facilitates the processor core 102 in sending certain memory transactions and other memory-related messages that bypass the DMC 104 and the PMC 108.
[0012] SoC 10 has a hierarchical memory system. Each cache at each level can be unified or partitioned into separate data and program caches. For example, DMC 104 can be coupled to a level 1 data cache 110 (L1D cache) to control data writes to and data reads from the L1D cache 110. Similarly, PMC 108 can be coupled to a level 1 program cache 112 (L1P cache) to read instructions for execution by the processor core 102. (In this example, the processor core 102 does not generate writes to the L1P cache 112.) The L1D cache 110 can have an L1D victim cache 113. A unified memory controller 114 (UMC) for a level 2 cache (L2 cache 116, e.g., L2 SRAM) is communicatively coupled to receive read and write memory access requests from DMC 104 and PMC 108, and read requests from the streaming engine 106, PMC 108, and memory management unit 117 (MMU). (The example L2 controller UMC 114 is referred to as a "unified" memory controller in the example system because UMC 114 can store both instructions and data in the L2 cache 116.) UMC 114 is communicatively coupled to pass read data and write acknowledgments (from outside the level 1 caches) to DMC 104, the streaming engine 106, and PMC 108, which then pass them to the processor core 102. UMC 114 is also coupled to control writes to and reads from the L2 cache 116, and pass memory access requests to a level 3 cache controller 118 (L3 controller). The L3 controller 118 is coupled to control writes to and reads from the L3 cache 119. UMC 114 is coupled to receive (via the L3 controller 118) write acknowledgments and data read from the L2 cache 116 and the L3 cache 119. UMC114 is configured to control pipelining of memory transactions for program content and data content (read and write requests for instructions, data transfers, and write acknowledgments). The L3 controller 118 is coupled to control writes to and reads from the L3 cache 119, and mediate transactions using an external function 120 located outside the processor 100, such as other processor cores, peripheral functions of the SOC 10, and / or other SoCs (and is also used to control snooping transactions). Thus, the L3 controller 118 is a shared memory controller of the SoC 10, and the L3 cache 119 is a shared cache memory of the SoC 10. Thus, memory transactions related to the processor 100 and the external function 120 pass through the L3 controller 118.
[0013] Memory transactions are generated by the processor core 102 and communicated to a lower-level cache memory, or are generated by the external function 120 and communicated to a higher-level cache memory. For example, a victim write transaction may be initiated by the UMC 114 in response to a read transaction that generates a miss in the L2 cache 116 from the processor core 102.
[0014] The MMU 117 provides address translation and memory attribute information to the processor core 102. It does this by looking up information in a table stored in memory (the connection between the MMU 117 and the UMC 114 enables the MMU 117 to access the memory containing the table using a read request).
[0015] Figure 2 is included in Figure 1 a block diagram of the instance memory pipeline 200 that is within or associated with the UMC 114, and thus, for illustration, Figure 2 also repeats various blocks that communicate with the UMC 114 from Figure 1 The memory pipeline 200 includes an initial scheduling block 202 coupled to an integer M number of pipeline banks 206. Each pipeline bank 206 includes an integer P number of stages 208 and is illustrated as a vertical column below the initial scheduling block 202. The DMC 104 is coupled to the initial scheduling block 202 by a bus 204-1 that is N1 rows wide, enabling the DMC 104 to provide read or write requests that transfer N1 bits of data in one transfer. The streaming engine 106 is coupled to the initial scheduling block 202 by a bus 204-2 that is N2 rows wide, enabling the streaming engine 106 to provide read requests that transfer N2 bits of data in one transfer. The PMC 108 is coupled to the initial scheduling block 202 by a bus 204-3 that is N3 rows wide, enabling the PMC 108 to provide read requests that transfer N3 bits of data in one transfer. The L3 controller 118 is coupled to the initial scheduling block 202 by a bus 204-4 that is N4 rows wide, enabling the L3 118 to provide read or write requests that transfer N4 bits of data in one transfer. The MMU 117 is coupled to the initial scheduling block 202 by a bus 204-5 that is N5 rows wide, enabling the MMU 117 to provide read requests that transfer N5 bits of data in one transfer.
[0016] When a memory controller of the processor 100 (such as DMC 104, streaming engine 106, PMC 108, MMU 117, or L3 controller 118) conveys a request to read from or write to a memory mediated by UMC 114 (such as the memory in L2 cache 116, L3 cache 119, or external function 120) to UMC 114, the initial scheduling block 202 schedules that the request will be handled by an appropriate pipeline library 206 for a specific request. Thus, the initial scheduling block 202 arbitrates read and write requests. The arbitration determines which pipeline library 206 will receive which memory transaction queued at the initial scheduling block 202, and in what order. Often, for example, depending on the memory address of the data being written or requested, the request load of the pipeline library 206, or a pseudo-random function, the read or write request is scheduled into the corresponding one of the pipeline libraries 206. The initial scheduling block 202 schedules read and write requests received from DMC 104, streaming engine 106, PMC 108, and L3 controller 118 by making a selection in the first stage of the pipeline library 206. A memory transaction performed on the L3 cache 119 (or external function 120) is arbitrated and scheduled into the L3 cache pipeline by the L3 cache scheduling block 404 in the L3 controller 118 after passing through the memory pipeline 200 corresponding to the L2 cache 116 (the pipeline library 206 and possible bus snooping related stages, which are not shown). Figure 4 )
[0017] Request scheduling prevents conflicts between read or write requests to be handled by the same pipeline library 206 and maintains memory consistency (described further below). For example, request scheduling maintains the order between memory transactions placed in the memory transaction queue (memory access request queue) of the initial scheduling block 202 by different memory controllers of the processor 100 or different buses of the same memory controller.
[0018] Further, pipeline memory transactions (read or write requests) sent by the DMC 104 or PMC 108 are considered because the memory transaction has passed through the corresponding level-1 cache pipeline (in the DMC 104 for the L1D cache 110 and in the PMC 108 for the L1P cache 112), and is targeted at a lower-level cache or memory endpoint (or external function 120), or has incurred a miss in the corresponding level-1 cache, or bypasses the L1D cache 110 because the corresponding data payload of the write request cannot be cached by the L1D cache 110. Generally, memory transactions that result in a level-1 cache hit and are directed to the DMC 104 or PMC 108 result in a write acknowledgment from the L1D cache 110 or a response with data or instructions read from the L1D cache 110 or L1P cache 112, respectively. Thus, memory transactions that result in a level-1 cache hit generally do not require access to Figure 2 the pipeline library 206 shown in Figure 1 ).
[0019] Figure 2 The pipeline library 206 shown in
[0020] is part of the UMC 114. The L1D cache 110 may hold data generated by the processor core 102. For example, the external function 120 may access data in the L1D cache 110 by writing data to the L2 cache 116 or L3 cache 119, or by reading or evicting data from the L1D cache 110 using snoop transactions controlled by the L3 controller 119 and managed by the UMC 114 (the L2 controller, as an agent).
[0020] Memory consistency means that the content of the memory in the system (or at least what is considered or indicated as valid) is the same as what one or more processors in the system expect based on an ordered stream of read and write requests. Writes that affect a particular data or a particular memory location are prevented from bypassing earlier issued writes or reads that affect the same data or the same memory location. Additionally, certain types of transactions have priorities, such as victim cache transactions and snoop transactions.
[0021] Bus snooping is a scheme by which a coherence controller (snooper) in a cache monitors or snoops bus transactions to maintain memory coherence in a distributed shared memory system (such as SoC 10). If a transaction that modifies a shared cache block appears on the bus, the snooper checks whether its corresponding cache has the same copy of the shared block. If the cache has a copy of the shared block, the corresponding snooper performs an action to ensure memory coherence in the cache. For example, depending on the transaction detected on the bus, this action can be to flush the shared block, invalidate the shared block, or update the shared block.
[0022] A "write stream" refers to a stream of write requests issued by a device (e.g., processor core 102), such as one write request per cycle, without pausing. For example, a write stream may be interrupted by a pause that results from a full buffer or an insufficient number of write request identifiers. The ability to pull write requests out of the memory transaction queue as fast as possible promotes the write stream.
[0023] To make the processor core 102 aware that a write has been completed, it must receive a write acknowledgment. To maintain coherence, the processor core 102 can self-limit to a given number of outstanding write requests by restricting write requests that exceed the limit until a write acknowledgment for an outstanding write request is received. Thus, the processor core 102 and the L1D cache 110 can wait for the write acknowledgment (or "handshake") to proceed while pausing the corresponding write stream process within the processor core 102. A pause that interrupts the write stream can also be caused by the processor core 102 or the DMC 104 waiting for a write acknowledgment from a previous write request. The processor core 102 can also be configured to pause while waiting for a write acknowledgment for certain operations (such as a fence operation). The DMC 104 (L1 cache 110 controller) uses a write acknowledgment (handshake) forwarded by the UMC 114 (L2 cache 116 controller) to detect the completion of a write in a lower-level cache (such as the L2 cache 116 or the L3 cache 119). However, due to various pipeline requirements that include arbitration, ordering, and coherence, a write may take many cycles to complete.
[0024] At a first-level arbitration performed by an initial scheduling block 202, the UMC 114 (the L2 cache 116 controller that includes the initial scheduling block 202) determines whether to allow a memory transaction to continue in the memory pipeline 200 and in which pipeline bank 206. Between (1) the initial arbitration and scheduling and (2) the write completion, writes to the L2 cache 116 typically have few operations. The remaining operations of a scheduled write request may include, for example, checking for errors (such as firewall, addressing, and out-of-range errors), read-modify-write actions (updating the error check code of the data payload of the write request), and committing the data payload of the write request to memory. Typically, each pipeline bank 206 is independent, such that write transactions on a pipeline bank 206 (e.g., writing data from the L1D cache 110 to the L2 cache 116) have no ordering or consistency requirements with respect to write transactions on other pipeline banks 206. Within each pipeline bank, writes to the L2 cache 116 are performed in their scheduled order. For partial writes that trigger read-modify-write transactions, relative ordering is maintained. If a memory transaction causes an addressing hazard or violates an ordering requirement, the transaction is suspended and not issued to the pipeline bank 206. (A partial write is a write request with a data payload that has a length less than the minimum write length of the destination cache memory. A partial write triggers a read-modify-write transaction in which data is read from the destination cache memory to pad the data payload of the write request to the minimum write length of the destination cache memory, and an updated error correction code (ECC) is generated from the resulting padded data payload and appended to the padded data payload. The padded data payload with the updated ECC is the content written to the destination cache memory.)
[0025] Due to these characteristics of the memory pipeline 200, once a write is scheduled within the pipeline bank 206 (e.g., a data write from the L1D cache 110 to the L2 cache 116), it is guaranteed that the write adheres to all ordering requirements and does not violate consistency (thus, meeting the conditions required to avoid disrupting ordering and consistency). Committing the write operation to memory may take a (variable) number of cycles, but a read issued after this write operation will "see" the write. Thus, if a read requests data or a memory location modified by a write, the read will retrieve the version of the data or the content of the memory location specified by the write, rather than the previous version. Write-write ordering is also maintained. Write requests to the L3 cache 119 can also be scheduled by the memory pipeline 200 (by the UMC 114) such that an ordered completion of the write requests to the L3 cache 119 is guaranteed. These guarantees mean that write requests scheduled by the memory pipeline 200 (the L2 cache pipeline) into the pipeline bank 206 can be guaranteed to meet the ordering and consistency requirements and be completed within a finite amount of time. Put another way, this guarantee is that a write transaction to a particular address that is currently being scheduled onto the pipeline bank 206 will "commit" its value to memory (the write will complete and store the corresponding data payload in memory) after a previously scheduled write transaction to the same address and before a later scheduled write transaction to the same address. This guarantee can be based on the pipeline inherently being "ordered" such that once a command enters the pipeline, it will be written to memory (committed) in the order in which it was scheduled. In other words, there are no bypass paths within the pipeline. (The bypass paths described below are handled such that they do not disrupt ordering guarantees for older transactions targeting the same memory address, for example.)
[0026] "Simultaneous with" is defined herein to mean simultaneous or immediately thereafter. Thus, the first event occurring "simultaneous with" the second event can mean that the two events occur on the same cycle of the system clock.
[0027] The UMC 114 (L2 controller) sends a write acknowledgment for a write to the L2 cache 116 or a higher-level cache (e.g., data from the L1D cache 110) to the DMC 104 (L1 controller), while the initial scheduling block 202 (first-level arbitration) schedules the corresponding write request. Thus, the write acknowledgment indicating write completion is sent simultaneously with the write request being scheduled, rather than after the memory pipeline 200 has completed processing the write request. This accelerated acknowledgment is achieved through the guarantee that the scheduled write requests will complete in order and meet the coherence requirements. The UMC 114 creates the illusion that the write request is completed on the cycle on which it is scheduled, rather than on the cycle on which the corresponding data is committed (written) to the memory. From the perspective of the observability of the processor core 102 or the DMC 104, the L2 cache 116 appears to complete the write request immediately when scheduling the write request. This enables the DMC 104 to release the suspension of the processor core 102 faster (or prevent the processor core 102 from suspending) and to draw write requests from the queue with lower latency (faster), thus improving the overall performance. The queue is a transaction queue in the UMC 114 sent from the corresponding "masters" (functional blocks that can send memory transactions to the UMC 114 for queuing), such as the DMC 104, the streaming engine 106, the PMC 108, the MMU 117, and the L3 controller 118. The queue can be implemented as a hold stage where memory transactions reside while waiting to be arbitrated by the initial scheduling block 202 and scheduled to the pipeline bank 206.
[0028] The processor core 102 is typically configured to read data from the memory for processing. This also holds for other processor cores 102 of other processors in the SoC 10 (e.g., the processor 100) with respect to the memory that can be accessed by those other processors. However, other processors in the SoC 10 require the data generated by the processor core 102 to be available outside the data-generating processor 100 in order to be able to access the generated data. This means that the generated data passes through the L3 controller 118 so as to be externally accessible within the shared memory (L3 cache 119) or by being sent to the external function 120.
[0029] Figure 3 is an example of Figure 1 the process 300 for a memory write operation of the SoC 10. In step 302, the initial scheduling block 202 schedules a write request directed to the L2 cache 116 to the pipeline bank 206. In step 304, after the write request is scheduled to the pipeline bank, the UMC 114 (L2 cache controller) immediately sends a write acknowledgment to the processor core 102. In step 306, the suspension of the processor core 102 is released in response to the write acknowledgment.
[0030] Figure 4 is for Figure 1 block diagram of an example memory pipeline 400 of the SoC 10. As Figure 4 shown in Figure 2 the memory pipeline 400 is similar to the memory pipeline 200 shown in
[0031] but also includes a bypass path 402. The bypass path 402 couples the initial scheduling block 202 to the L3 cache scheduling block 404, thus skipping the memory pipeline corresponding to at least one level of the cache. For example, the bypass path 402 enables bypass writes to bypass a portion of the memory pipeline 400 associated with the L2 cache 116, thereby shortening the total processing time of the bypass writes. Thus, writes to certain memory regions can be written to the L3 cache 119 instead of the L2 cache 116 (or other lower-level caches). Such writes (bypass writes) can be performed safely without accessing the memory pipeline levels associated with the L2 cache 116, which includes the pipeline bank 206 and associated bus snooping (not shown). This simplifies the coherence requirements for bypass writes.
[0032] The initial scheduling block 202 may have a specified bypass state (or "bypass mode") where bypass writes may be scheduled to the bypass path 402 instead of to the full memory pipeline 400 (including to the pipeline bank 206). When the initial scheduling block 202 is in the bypass mode, the bypass writes bypass the entire pipeline of the middle-level cache, including the associated internal arbitration. When the initial scheduling block 202 is not in the bypass mode, the bypass writes go through the full memory pipeline 400.
[0033] Figure 5 is an example of a process 500 for Figure 1 memory write operations of the SoC 10. Figure 5 Describes making the initial scheduling block 202 enter and remain in the bypass mode such that bypass writes are allowed to go along Figure 4Conditions for the bypass path 402 to proceed. In step 502, the initial scheduling block 202 determines whether the next memory transaction in the queue is a bypass write, carrying data or related to a memory location that is guaranteed not to be written to the L2 cache 116. (For example, the data payload is too large to fit in the L2 cache 116, or contains a data type not stored by the L2 cache 116.) If not, then the next memory transaction is processed according to step 504 using normal (non-bypass) processing (through the full memory pipeline 400), and the process returns to step 502. In step 506, the initial scheduling block 202 determines whether a write to or from the L1D victim cache is in the pipeline (in the pipeline bank 206). If so, then the victim cache write is prioritized as a memory synchronization event, and in step 504, bypass writes are not allowed to proceed along the bypass path 402. Instead, the bypass write receives normal (non-bypass) processing (although non-bypass processing is prioritized), and the process 500 returns to step 502. In step 508, the initial scheduling block 202 determines whether there is an L1D write (a write from the DMC 104; in some embodiments, any write request generated outside the memory pipeline 200) in the pipeline bank 206 (to which the bypass write will be scheduled). If so, then in step 510, the bypass write is delayed until the L1D write (all L1D writes in the corresponding pipeline bank 206; in some embodiments, all write requests generated outside the memory pipeline 200) is cleared from the corresponding pipeline bank 206 (e.g., the corresponding write is committed to memory, or reaches the L3 cache scheduling block 404 for scheduling into the corresponding memory pipeline of the L3 controller 118). Otherwise, in step 512, the bypass write enters the bypass mode 514, and the bypass write is sent along the bypass path 402 to the L3 controller 118 (and from there to the L3 cache 119). Next, in step 516, the initial scheduling block 202 checks whether the next memory transaction in its queue is a bypass write. If so, then the process 500 remains in the bypass mode by returning to step 512, and the next memory transaction in the queue (the bypass write) is sent along the bypass path 402 to the L3 cache scheduling 404. Thus, after a bypass write has met the conditions for using the bypass path 402, queued bypass writes in order can also be sent to the L3 cache scheduling 404 via the bypass path 402 without having to re-check the conditions processed in steps 506 and 508. If the next memory transaction in the queue of the initial scheduling block 202 is not a bypass write, then in step 518, the initial scheduling block 202 exits the bypass mode, the next memory transaction in the queue receives normal (non-bypass) processing, and the process 500 returns to step 502.
[0034] In the described embodiments, modifications are possible and, within the scope of the claims, other embodiments are possible.
[0035] In some embodiments, the streaming engine passes and returns responses to both read and write requests.
[0036] In some embodiments, the processor may include multiple processor cores (embodiments with multiple processor cores are not shown), which have functional couplings with the DMC, streaming engine, and PMC similar to those shown and described herein. Figure 1 with the DMC, streaming engine, and PMC similar to those shown and described herein.
[0037] In some embodiments, a bus that implements parallel read or write requests may correspond to different types of read or write requests, such as made for different memory blocks or for different purposes.
[0038] In some embodiments, the streaming engine enables the processor core to communicate directly with a lower-level cache (e.g., an L2 cache), skipping a higher-level cache (e.g., an L1 cache), to avoid data synchronization issues. This can be used to help maintain memory consistency. In some such embodiments, the streaming engine may be configured to emit only read requests, rather than both read and write requests.
[0039] In some embodiments, an L3 cache or other lower-level memory may schedule write requests such that a write acknowledgment can be sent to the DMC (or the processor core or another lower-level memory controller) while the write request is being scheduled to the corresponding memory pipeline.
[0040] In some embodiments, different memory access pipeline banks may have different numbers of stages.
[0041] In some embodiments, a processor in an external function may access data stored in the L2 cache; in some such embodiments, consistency between the content stored in the L2 cache and the content cached in other processors in the external function is not guaranteed.
[0042] In some embodiments, if the data being included is too large for a lower-level cache (e.g., an L1D cache or an L2 cache), the write may be a bypass write.
[0043] In some embodiments, if the page attributes mark the write as corresponding to a device type memory region not cached by the UMC (L2 cache controller), the write may be a bypass write.
[0044] In some embodiments, the L1D cache (or other lower-level cache) may cache the data payload of a bypass write.
[0045] In some embodiments, the memory coherence rules of the processor prohibit bypassing memory transactions (memory read requests or memory write requests) other than L1D victim cache writes and L1D writes.
[0046] In some embodiments, the bypass path jumps the bypass write to the final arbitration stage before being scheduled into the memory pipeline bank of the L3 cache (not shown).
[0047] In some embodiments, the guarantee that memory writes are never directed to the L2 cache (corresponding to bypass writes) includes the option that this will never change, or the guarantee that such writes to the L2 cache are impossible.
[0048] In some embodiments, the guarantee that memory writes are never directed to the L2 cache includes that the L2 cache does not have a copy or hash of the corresponding data.
[0049] In some embodiments, the guarantee that memory writes are never directed to the L2 cache may change (if the guarantee has not been in effect recently, then it is activated). In such embodiments, making this guarantee when a line is being written to the L2 cache (e.g., changing the corresponding mode register to make the guarantee) may require flushing the corresponding L2 cache. In this case, cache flushing can prevent a cache copy of the data payload that is now guaranteed not to be written to the L2 cache from remaining in the L2 cache after the guarantee is made.
[0050] In some embodiments, the cache controller initiates a memory transaction only in response to a transaction initiated by the processor core 102 or the external function 120.
[0051] In some embodiments, for example, depending on the memory address of the data being written or requested, the request payload of the pipeline bank 206, or a pseudo-random function, only read or write requests can be scheduled into the corresponding one of the pipeline banks 206.
[0052] Modifications are possible in the described embodiments, and other embodiments are possible within the scope of the claims.
Claims
1. A processor system, comprising: a processor core configured to generate memory transactions; a first-level cache memory; a first memory controller coupled to control the first-level cache memory; and a second-level cache memory, which is a cache memory of a different level from the first-level cache memory; a second memory controller having a memory pipeline configured to process the memory transactions, the second memory controller coupled to control the second-level cache memory, the second memory controller connected to the first memory controller through a bypass path that skips the memory pipeline, the second memory controller configured to: determine whether a first memory transaction to be scheduled is a bypass memory write request, a bypass memory write request being a memory write request having a data payload or corresponding to a memory location indicated not to cause a corresponding write to be directed to the second-level cache memory; and if the first memory transaction is a bypass memory write request, then determine whether a second memory transaction in the memory pipeline prevents passing through the first memory transaction, and if not, then use the bypass path to send the first memory transaction to the first memory controller.
2. The processor system according to claim 1, wherein the second memory controller is configured to: enter a bypass mode if the first memory transaction is sent to the first memory controller using the bypass path; if a third memory transaction sequentially after the first memory transaction is determined to be a bypass memory write request and the second memory controller is in the bypass mode, then skip the action of determining whether the second memory transaction prevents passing, and use the bypass path to send the third memory transaction to the first memory controller.
3. The processor system according to claim 2, wherein the second memory controller is configured to exit the bypass mode if the third memory transaction is not determined to be a bypass memory write request.
4. The processor system according to claim 1, wherein the second memory controller is configured not to cause the data payload of the first memory transaction to be written to the second-level cache memory after sending the first memory transaction to the first memory controller using the bypass path.
5. The processor system according to claim 1, wherein the first memory controller is configured to process the first memory transaction after receiving the first memory transaction through the bypass path.
6. The processor system according to claim 1, further comprising: a third-level cache memory, which is a cache memory of a different level from the first-level cache memory and the second-level cache memory; and a victim cache of the third-level cache memory; Wherein the second memory controller is configured such that a memory write request requesting a memory write from the victim cache to the level-3 cache memory is a memory transaction that is prevented from being passed through the first memory transaction.
7. The processor system according to claim 6, wherein, If, when the first memory transaction is a bypass memory write request, it is determined that a memory write request requesting a memory write from the victim cache to the level-3 cache memory is in the memory pipeline, then the bypass memory write request is scheduled into the memory pipeline and executed by the memory pipeline.
8. The processor system according to claim 1, further comprising a level-3 cache memory, which is a cache memory of a different level from the first-level cache memory and the second-level cache memory; wherein the second memory controller is configured such that a memory write request requesting a write to the level-3 cache memory is a memory transaction that is prevented from being passed through the first memory transaction.
9. The processor system according to claim 8, wherein the second memory controller is configured to delay the first memory transaction until it is determined that there is no level-3 cache write request in the memory pipeline by the determination of whether a second memory transaction prevents a pass-through action, and thereafter send the first memory transaction to the first memory controller using the bypass path.
10. The processor system according to claim 1, wherein the bypass path is directly connected from the memory transaction scheduling portion of the first memory controller to the second memory controller.
11. A method of operating a processor system, the method comprising: wherein the processor system has a first-level cache memory and a second-level cache memory, the second-level cache memory being a cache memory of a different level from the first-level cache memory, the first-level cache memory having a first memory controller, and the second-level cache memory having a second memory controller; using the second memory controller to determine whether a first memory transaction is a bypass memory write request, the first memory transaction being a memory transaction to be scheduled by the second memory controller, the bypass memory write request being a memory write request having a data payload or corresponding to a memory location indicated not to cause a corresponding write to be directed to the second-level cache memory; determining whether a memory transaction that prevents passing is in the memory pipeline of the second memory controller; and sending the first memory transaction from the second memory controller to the first memory controller using a bypass path that skips the memory pipeline of the second memory controller, the sending being performed based on the determination that the first memory transaction is a bypass memory write request and based on the determination that there is no memory transaction that prevents passing in the memory pipeline.
12. The method according to claim 11, wherein the bypass path is directly connected from the memory transaction scheduling part of the second memory controller to the first memory controller.
13. The method according to claim 11, further comprising: Entering a bypass mode in the memory pipeline based on the sending step performed on the first memory transaction.
14. The method according to claim 13, further comprising: Using the second memory controller to determine whether a second memory transaction is a bypass memory write request, the second memory transaction being a memory transaction to be scheduled by the second memory controller; Based on determining that the second memory transaction is a bypass memory write request and based on the memory pipeline being in the bypass mode, not determining whether a memory transaction that prevents the second memory transaction from being transmitted is in the memory pipeline of the second memory controller; Based on determining that the second memory transaction is a bypass memory write request and based on the memory pipeline being in the bypass mode, performing the sending step on the second memory transaction; and Based on determining that the second memory transaction is not a memory write request and based on the memory pipeline being in the bypass mode, exiting the bypass mode and processing the second memory transaction using the memory pipeline.
15. The method according to claim 11, further comprising: After sending the first memory transaction to the first memory controller using the bypass path, the second memory controller does not cause the data payload of the first memory transaction to be written to the second-level cache memory.
16. The method according to claim 11, further comprising: After the first memory controller receives the first memory transaction through the bypass path, processing the first memory transaction using the first memory controller.
17. The method according to claim 11, further comprising: Wherein the processor system has a third-level cache memory, which is a cache memory of a different level from the first-level cache memory and the second-level cache memory, and the third-level cache memory has a victim cache; and Wherein a memory transaction related to the victim cache memory prevents transmission through the first memory transaction; Based on determining that a memory transaction related to the victim cache memory is in the memory pipeline, processing the first memory transaction using the memory pipeline.
18. The method according to claim 11, further comprising: Based on determining that the memory transaction is not a bypass memory write request, processing the first memory transaction using the memory pipeline.
19. The method according to claim 11, Wherein the processor system has a third-level cache memory, which is a cache memory of a different level from the first-level cache memory and the second-level cache memory; and Wherein a memory write request requesting a memory write to the third-level cache memory is a memory transaction that is prevented from being transmitted through the first memory transaction.
20. The method according to claim 19, further comprising: Using the second memory controller, delaying the first memory transaction based on determining that the first memory transaction is a bypass memory write request and based on determining that a memory write request requesting a memory write to the third-level cache memory is in the memory pipeline; And After performing the delaying step, performing the sending step on the first memory transaction based on determining that the memory write request requesting the memory write to the third-level cache memory is no longer in the memory pipeline.
Citation Information
Patent Citations
Common platform for one-level memory architecture and two-level memory architecture
US20150178204A1
Storage subsystem including an error correcting cache and means for performing memory to memory transfers
US6161208A