Write stream in processor
By introducing memory pipelines and arbitration mechanisms into the memory controller of the SOC system, the delay and consistency problems in the memory write operation of the SOC system are solved, and more efficient memory write operation and system performance improvements are achieved.
Patent Information
- Application Number
- CN202080036288.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-05-14
- Filing Date
- 2020-05-26
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2040-05-26
AI Technical Summary
Existing SOC systems have latency and consistency problems in memory write operations, which affect system performance.
By introducing memory pipelines and arbitration mechanisms into the memory controller, memory write requests are scheduled and write acknowledgements are sent at the same time when the write request is scheduled, ensuring the order and consistency of the write operations.
Improves the efficiency of memory write operations, reduces latency, ensures system consistency and stability, thereby improving overall performance.
Smart Images

Figure CN113826085B_ABST
Abstract
Description
Technical Field
[0001] The present description generally relates to processing devices that may be formed as part of an integrated circuit, such as a system on a chip (SoC). More particularly, the present description relates to such systems with improved management of write operations. Background Art
[0002] A SOC is an integrated circuit having multiple functional blocks, such as one or more processor cores, memory, and input and output, on a single die.
[0003] Hierarchical memory moves data and instructions between memory blocks with different read / write response times for respective processor cores, such as central processing units (CPUs) or digital signal processors (DSPs). For example, memory that is more local to a respective processor core typically has a lower response time. Hierarchical memory includes a cache memory system having multiple levels, such as L1 and L2, where the different levels describe different degrees of locality or different average response times of the cache memory to the respective processor core. Here, a cache memory that is more local or has a lower response time, such as an L1 cache, is referred to as a higher level of cache memory than a lower level cache memory that is less local or has a higher response time, such as an L2 cache or an L3 cache. Cache associativity refers to cache storage separation, where set associativity divides the cache into several storage groups, and each such group stores several (way) blocks, while a fully associative cache is not constrained by the set restriction. Thus, for an integer N, each location in the main memory (system memory) can reside in any of the N possible locations in an N-way associative cache.
[0004] A "victim cache" memory caches data (e.g., cache lines) evicted from a cache memory (e.g., an L1 cache). If an L1 cache read results in a miss (the data corresponding to a portion of main memory is not stored in the L1 cache), then a lookup occurs in the victim cache. If a victim cache lookup results in a hit (the data corresponding to the requested memory address is present in the victim cache), then the contents of the victim cache location that produced the hit are swapped with the contents of the corresponding location in the corresponding cache (in this example, the L1 cache). Some example victim caches are fully associative. Data corresponding to any location in main memory can be mapped to (stored in) any location in a fully associative cache. Summary of the invention
[0005] In the described example, a processor system includes a processor core that generates memory write requests, a cache memory, and a memory controller. The memory controller has a memory pipeline. The memory controller is coupled to control the cache memory and is communicatively coupled to the processor core. The memory controller is configured to receive the memory write requests from the processor core; schedule the memory write requests on the memory pipeline; and while scheduling a respective one of the memory write requests on the memory pipeline, send a write acknowledgement to the processor core, the write acknowledgement confirming that the write of the data payload of the respective memory write request to the cache memory has been completed. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Figure 1 is a block diagram of an example processor that is part of a SoC.
[0007] Figure 2 is used for Figure 1 Block diagram of the memory pipeline of an example SoC.
[0008] Figure 3 is used for Figure 1 An example of the process of a SoC memory write operation.
[0009] Figure 4 is used for Figure 1 Block diagram of the memory pipeline of an example SoC.
[0010] Figure 5 is used for Figure 1 An example of the process of a SoC memory write operation. DETAILED DESCRIPTION
[0011] Figure 11 is a block diagram of an example processor 100 as part of SoC 10. SoC 10 includes a processor core 102, such as a CPU or DSP, that generates new data. Processor 100 may include a clock 103, which may be part of processor core 102, or may also be separate from processor core 102 (a separate clock not shown). Processor core 102 also generates memory read requests requesting to read from data memory controller 104 (DMC) and stream engine 106 and memory write requests requesting to write to data memory controller 104 (DMC) and stream engine 106. In some embodiments, processor core 102 generates one read request or write request per processor core clock cycle. The memory controller returns a write confirmation to processor core 102 to confirm that the requested memory write has been executed. Processor core 102 is also coupled to receive instructions from program memory controller 108 (PMC). Stream engine 106 facilitates processor core 102 to send certain memory transactions and other memory-related messages that bypass DMC 104 and PMC 108.
[0012] SoC 10 has a hierarchical memory system. Each cache at each level can be unified or divided into separate data and program caches. For example, DMC 104 can be coupled to level 1 data cache 110 (L1D cache) to control data writing to L1D cache 110 and data reading from L1D cache 110. Similarly, PMC 108 can be coupled to level 1 program cache 112 (L1P cache) to read instructions from L1P cache 112 for execution by processor core 102. (In this example, processor core 102 does not generate writes to L1P cache 112.) L1D cache 110 may have L1D victim cache 113. A unified memory controller 114 (UMC) for a second level cache (L2 cache 116, such as an L2 SRAM) is communicatively coupled to receive read and write memory access requests from the DMC 104 and PMC 108, and to receive read requests from the stream engine 106, PMC 108, and memory management unit 117 (MMU). (The example L2 controller UMC 114 is referred to as a "unified" memory controller in the example system because the UMC 114 can store both instructions and data in the L2 cache 116.) The UMC 114 is communicatively coupled to pass read data and write acknowledgements (outside of the level 1 cache) to the DMC 104, stream engine 106, and PMC 108, which are then passed to the processor core 102. The UMC 114 is also coupled to control writes to and reads from the L2 cache 116, and pass memory access requests to a level 3 cache controller 118 (L3 controller). The L3 controller 118 is coupled to control writes to and reads from the L3 cache 119. The UMC 114 is coupled to receive (via the L3 controller 118) write confirmations and data read from the L2 cache 116 and the L3 cache 119. The UMC 114 is configured to control pipelining of memory transactions for program content and data content (read and write requests for instructions, data transmission and write confirmations). The L3 controller 118 is coupled to control writes to and reads from the L3 cache 119, and mediates transactions using an external function 120, which is located outside the processor 100, such as other processor cores, peripheral functions of the SOC 10, and / or other SoCs (and also used to control snooping transactions). Therefore, the L3 controller 118 is a shared memory controller of the SoC 10, and the L3 cache 119 is a shared cache memory of the SoC 10. Therefore, memory transactions associated with the processor 100 and the external functions 120 pass through the L3 controller 118 .
[0013] Memory transactions are generated by processor core 102 and communicated to lower levels of cache memory, or generated by external functions 120 and communicated to higher levels of cache memory. For example, a victim write transaction may be initiated by UMC 114 in response to a read transaction from processor core 102 that generates a miss in L2 cache 116.
[0014] The MMU 117 provides address translations and memory attribute information to the processor core 102. It does this by looking up the information in a table stored in memory (the connection between the MMU 117 and the UMC 114 enables the MMU 117 to access the memory containing the table using a read request).
[0015] Figure 2 is included in Figure 1 A block diagram of an example memory pipeline 200 within or associated with the UMC 114 is shown, for purposes of illustration, Figure 2 Also repeated from Figure 1 1 and 1 . The memory pipeline 200 includes an initial scheduling block 202 coupled to an integer number M of pipeline banks 206. Each pipeline bank 206 includes an integer number P of stages 208 and is illustrated as a vertical column below the initial scheduling block 202. The DMC 104 is coupled to the initial scheduling block 202 via a bus 204-1 that is N1 rows wide, enabling the DMC 104 to provide read or write requests that transfer N1 bits of data at a time. The stream engine 106 is coupled to the initial scheduling block 202 via a bus 204-2 that is N2 rows wide, enabling the stream engine 106 to provide read requests that transfer N2 bits of data at a time. The PMC 108 is coupled to the initial scheduling block 202 via a bus 204-3 that is N3 rows wide, enabling the PMC 108 to provide read requests that transfer N3 bits of data at a time. The L3 controller 118 is coupled to the initial scheduling block 202 via a bus 204-4 having a width of N4 rows, so that the L3 118 can provide read or write requests that transfer N4 bits of data at a time. The MMU 117 is coupled to the initial scheduling block 202 via a bus 204-5 having a width of N5 rows, so that the MMU 117 can provide read requests that transfer N5 bits of data at a time.
[0016] When a memory controller (e.g., DMC 104, stream engine 106, PMC 108, MMU 117, or L3 controller 118) of processor 100 communicates to UMC 114 a request to read from or write to a memory (e.g., L2 cache 116, L3 cache 119, or memory in external function 120) intervened by UMC 114, initial scheduling block 202 schedules the request to be handled by the appropriate pipeline library 206 for the particular request. Thus, initial scheduling block 202 performs arbitration for read and write requests. Arbitration determines which pipeline library 206 will receive which memory transaction queued at initial scheduling block 202, and in what order. Typically, a read or write request is scheduled to a corresponding one of pipeline libraries 206, for example, depending on the memory address of the data being written or requested, the request load of the pipeline library 206, or a pseudo-random function. The initial scheduling block 202 schedules read and write requests received from the DMC 104, the stream engine 106, the PMC 108, and the L3 controller 118 by selecting in the first stage of the pipeline library 206. Memory transactions requested to be executed on the L3 cache 119 (or external functions 120) are sent by the L3 controller 118 (see the L3 controller 118) after passing through the memory pipeline 200 corresponding to the L2 cache 116 (the pipeline library 206 and possible bus snooping related stages, which are not shown). Figure 4 ) in the L3 cache scheduling block 404 arbitrates and schedules to the L3 cache pipeline.
[0017] Request scheduling prevents conflicts between read or write requests to be handled by the same pipeline library 206 and maintains memory consistency (described further below). For example, request scheduling maintains the order between memory transactions placed into the memory transaction queue (memory access request queue) of the initial scheduling block 202 by different memory controllers of the processor 100 or different buses of the same memory controller.
[0018] Further, a pipeline memory transaction (read or write request) sent by the DMC 104 or PMC 108 is requested because the memory transaction has passed through the corresponding level 1 cache pipeline (in the DMC 104 for the L1D cache 110 and in the PMC 108 for the L1P cache 112) and has either generated a miss in the corresponding level 1 cache for a lower level cache or memory endpoint (or external function 120), or bypassed the L1D cache 110 because the corresponding data payload of the write request is not cacheable by the L1D cache 110. Typically, a memory transaction directed to the DMC 104 or PMC 108 that generates a level 1 cache hit results in a write acknowledgement from the L1D cache 110 or a response with data or instructions read from the L1D cache 110 or L1P cache 112, respectively. Therefore, a memory transaction that generates a level 1 cache hit typically does not require access to the L1D cache 110. Figure 2 The pipeline library 206 shown in FIG. 2 controls or intervenes in the memory access to the L2 cache 116, the L3 cache 119, and the external function 120 (see Figure 1 ).
[0019] Figure 2 1 is part of the UMC 114. The L1D cache 110 may hold data generated by the processor core 102. For example, the external function 120 may access data in the L1D cache 110 by writing the data to the L2 cache 116 or L3 cache 119, or reading or evicting data from the L1D cache 110 using snoop transactions controlled by the L3 controller 119 and managed by the UMC 114 (L2 controller, as a proxy).
[0020] Memory consistency is the process of ensuring that the contents of memory in a system (or at least the contents that are deemed or indicated as valid) are the same as expected by one or more processors in the system based on an ordered flow of read and write requests. Writes that affect specific data or a specific memory location are prevented from bypassing earlier issued writes or reads that affect the same data or the same memory location. In addition, certain types of transactions have priority, such as victim cache transactions and snoop transactions.
[0021] Bus snooping is a scheme whereby a coherence controller (snooper) in the cache monitors or snoops bus transactions to maintain memory coherence in a distributed shared memory system such as SoC 10. If a transaction that modifies a shared cache block appears on the bus, the snooper checks whether its corresponding cache has an identical copy of the shared block. If the cache has a copy of the shared block, the corresponding snooper performs an action to ensure memory coherence in the cache. For example, this action may be to refresh the shared block, invalidate the shared block, or update the shared block, depending on the transaction detected on the bus.
[0022] A "write stream" refers to a stream of write requests issued by a device (e.g., processor core 102), e.g., one write request per cycle, without pauses. For example, a write stream may be interrupted by pauses due to a full buffer or an insufficient number of write request identifiers. The ability to cause write requests to be pulled from the memory transaction queue as quickly as possible facilitates write streaming.
[0023] In order for the processor core 102 to know that the write has completed, it must receive a write confirmation. To maintain consistency, the processor core 102 can limit itself to a given number of outstanding write requests by limiting the write requests that exceed the limit until the write confirmation of the outstanding write request is received. Therefore, the processor core 102 and the L1D cache 110 can wait for the write confirmation (or "handshake") to proceed while pausing the corresponding write flow process within the processor core 102. The pause that interrupts the write flow can also be caused by the processor core 102 or DMC 104 waiting for the write confirmation from the previous write request. The processor core 102 can also be configured to pause while waiting for a write confirmation for certain operations (such as fence operations). The write confirmation (handshake) forwarded by the UMC 114 (level 2 cache 116 controller) is used by the DMC104 (level 1 cache 110 controller) to detect the completion of the write in the lower level cache (such as L2 cache 116 or L3 cache 119). However, due to various pipeline requirements including arbitration, ordering, and coherency, writes may take many cycles to complete.
[0024] At the first level of arbitration performed by the initial scheduling block 202, the UMC 114 (L2 cache 116 controller, which includes the initial scheduling block 202) determines whether a memory transaction is allowed to proceed in the memory pipeline 200, and in which pipeline bank 206. A write to the L2 cache 116 typically has very few operations between (1) the initial arbitration and scheduling and (2) the write completion. The remaining operations of the scheduled write request may include, for example, checking for errors (such as firewall, addressing, and out-of-range errors), read-modify-write actions (updating the error checking code of the data payload of the write request), and committing the data payload of the write request to memory. Typically, each pipeline bank 206 is independent, so that write transactions on a pipeline bank 206 (e.g., data writes from the L1D cache 110 to the L2 cache 116) have no ordering or consistency requirements with respect to write transactions on other pipeline banks 206. Within each pipeline bank, writes to the L2 cache 116 proceed in the order in which they were scheduled. Relative ordering is maintained for partial writes that trigger read-modify-write transactions. If a memory transaction causes an addressing hazard or violates ordering requirements, the transaction is suspended and is not issued to the pipeline library 206. (A partial write is a write request with a data payload that is less than the minimum write length of the destination cache memory. A partial write triggers a read-modify-write transaction, where data is read from the destination cache memory to fill the data payload of the write request to the minimum write length of the destination cache memory, and an updated error correction code (ECC) is generated from and appended to the generated pad data payload. The pad data payload with the updated ECC is what is written to the destination cache memory.)
[0025] Due to these characteristics of the memory pipeline 200, once a write is scheduled within the pipeline library 206 (e.g., a write of data from the L1D cache 110 to the L2 cache 116), the write is guaranteed to comply with all ordering requirements and not violate consistency (thus, satisfy the conditions that need to be satisfied to avoid breaking ordering and consistency). It may take a (variable) number of cycles to commit a write operation to memory, but read operations issued after this write operation is issued will "see" the write operation. Therefore, if a read requests data or a memory location modified by a write, the read will retrieve the version of the data or the contents of the memory location specified by the write, rather than the previous version. Write-write ordering is also maintained. L3 cache 119 write requests can also be scheduled by the memory pipeline 200 (by the UMC 114) so that the ordered completion of the L3 cache 119 write requests is guaranteed. These guarantees mean that write requests scheduled into the pipeline library 206 by the memory pipeline 200 (L2 cache pipeline) can be guaranteed to comply with ordering and consistency requirements and to complete within a bounded amount of time. Stated another way, this guarantee is a guarantee that a write transaction to a particular address and currently being scheduled onto pipeline store 206 will have its value "committed" to memory (the write will complete and the corresponding data payload stored in memory) after a previously scheduled write transaction to the same address and before a later scheduled write transaction to the same address. This guarantee may be based on the pipeline being inherently "in-order", so that once commands enter the pipeline, they will be written to memory (committed) in the order in which they were scheduled. In other words, there are no bypass paths within the pipeline. (The bypass paths described below are handled so that they do not break ordering guarantees, for example, regarding older transactions targeting the same memory address.)
[0026] "Concurrently with" is defined herein to mean at the same time or immediately thereafter. Thus, a first event occurring "concurrently with" a second event may mean that the two events occur on the same cycle of a system clock.
[0027] The UMC 114 (L2 controller) sends a write acknowledgement of a write of data to the L2 cache 116 or a higher level cache (e.g., data from the L1D cache 110) to the DMC 104 (L1 controller) while the initial scheduling block 202 (first level arbitration) schedules the corresponding write request. Thus, the write acknowledgement indicating the completion of the write is sent simultaneously with the write request being scheduled, rather than after the memory pipeline 200 completes processing the write request. This accelerated acknowledgement is achieved by the guarantee that the scheduled write request will be completed in order and meet the consistency requirements. The UMC 114 creates the illusion that the write request is completed on the cycle in which it is scheduled, rather than on the cycle in which the corresponding data is committed (written) to the memory. From the perspective of observability of the processor core 102 or the DMC 104, the L2 cache 116 appears to complete the write request immediately when the write request is scheduled. This enables the DMC 104 to un-pause the processor core 102 more quickly (or prevent the processor core 102 from being stalled), and to extract write requests from the queue with lower latency (faster), thereby improving overall performance. The queue is a queue of transactions in the UMC 114 that are sent from the corresponding "masters" (functional blocks that can send memory transactions to the UMC 114 for queuing), such as the DMC 104, the stream engine 106, the PMC 108, the MMU 117, and the L3 controller 118. The queue can be implemented as a holding stage where memory transactions reside while waiting to be arbitrated by the initial scheduling block 202 and scheduled to the pipeline library 206.
[0028] The processor core 102 is typically configured to read data from memory to process the data. This also applies to other processor cores 102 of other processors of the SoC 10 (e.g., processor 100) with respect to memory accessible by those other processors. However, the other processors of the SoC 10 require that the data generated by the processor core 102 be available outside the data generating processor 100 in order to be able to access the generated data. This means that the generated data passes through the L3 controller 118 in order to be accessible externally in the shared memory (L3 cache 119) or by being issued to the external function 120.
[0029] Figure 3 is used for Figure 1 10. In step 302, the initial dispatch block 202 dispatches the write request directed to the L2 cache 116 to the pipeline library 206. In step 304, after dispatching the write request to the pipeline library, the UMC 114 (L2 cache controller) directly sends a write confirmation to the processor core 102. In step 306, the processor core 102 is unpaused in response to the write confirmation.
[0030] Figure 4 is used for Figure 1 A block diagram of an example memory pipeline 400 of the SoC 10 is shown. Figure 4 The memory pipeline 400 shown in FIG. Figure 2 116, but also includes a bypass path 402. The bypass path 402 couples the initial scheduling block 202 to the L3 cache scheduling block 404, thereby skipping the memory pipeline corresponding to at least one level of the cache. For example, the bypass path 402 enables bypass writes to bypass the portion of the memory pipeline 400 associated with the L2 cache 116, thereby shortening the overall processing time of the bypass writes. Therefore, writes to certain memory areas can be written to the L3 cache 119 instead of being written to the L2 cache 116 (or other lower level caches). Such writes (bypass writes) can be performed safely without accessing the memory pipeline stages associated with the L2 cache 116, which include the pipeline library 206 and related bus snooping (not shown). This simplifies the consistency requirements of the bypass writes.
[0031] To allow bypass writes to use bypass path 402, memory consistency imposes ordering requirements. For example, writes cannot bypass other writes (write-write ordering); writes cannot bypass reads (read-write ordering); writes cannot bypass victim cache transactions (e.g., write requests from L1D victim cache 113 to L1D cache 110, which can be handled by UMC 114 at the L1 level through a cache miss of the victim cache-related write request); and writes cannot bypass snoop responses (e.g., snoop responses corresponding to a controller request to write from the victim cache to L1D cache 110 when the request is not caused by a cache miss). Victim cache transactions and snoop responses have high priority because they constitute memory synchronization events. In addition, L1D cache 110 victims pass through the full pipeline because, for example, they include updating internal states in UMC 114. However, bypass writes directly following victim cache transactions can also be prioritized, thereby ensuring that bypass writes will not be blocked or paused (similar to slip flow). This prioritization state applies only to a single bypass write directly following a victim cache transaction and ends when the bypass write is sent to the L3 controller 118 for further processing.
[0032] The initial scheduling block 202 may have a designated bypass state (or "bypass mode") in which bypass writes may be scheduled to bypass path 402 instead of to the full memory pipeline 400 (including to pipeline bank 206). When the initial scheduling block 202 is in bypass mode, bypass writes bypass the entire pipeline of the intermediate level cache, including associated internal arbitration. When the initial scheduling block 202 is not in bypass mode, bypass writes pass through the full memory pipeline 400.
[0033] Figure 5 is used for Figure 1 An example of a process 500 for a memory write operation of the SoC 10 is shown. Figure 5 Describes the conditions that cause the initial scheduling block 202 to enter and remain in bypass mode, allowing bypass writes along Figure 4The bypass path 402 continues. In step 502, the initial scheduling block 202 determines whether the next memory transaction in the queue is a bypass write, carrying data or associated with a memory location that is guaranteed not to be written to the L2 cache 116. (For example, the data payload is too large to fit in the L2 cache 116, or contains a data type that is not stored by the L2 cache 116.) If not, then the next memory transaction is processed using normal (non-bypass) processing (through the full memory pipeline 400) according to step 504, and processing returns to step 502. In step 506, the initial scheduling block 202 determines whether a write to or from the L1D victim cache is in the pipeline (in the pipeline library 206). If so, then the victim cache write is prioritized as a memory synchronization event, and in step 504, the bypass write is not allowed to continue along the bypass path 402. Instead, the bypass write receives normal (non-bypass) processing (although the non-bypass processing is prioritized), and the process 500 returns to step 502. In step 508, the initial dispatch block 202 determines whether there is an L1D write (a write from the DMC 104; in some embodiments, any write request generated outside the memory pipeline 200) in the pipeline library 206 (the pipeline library 206 to which the bypass write is to be dispatched). If so, then in step 510, the bypass write is delayed until the L1D write (all L1D writes in the corresponding pipeline library 206; in some embodiments, all write requests generated outside the memory pipeline 200) are cleared from the corresponding pipeline library 206 (e.g., the corresponding write is committed to memory, or reaches the L3 cache dispatch block 404 for dispatch to the corresponding memory pipeline of the L3 controller 118). Otherwise, in step 512, the bypass write enters the bypass mode 514, and the bypass write is sent along the bypass path 402 to the L3 controller 118 (and from there to the L3 cache 119). Next, in step 516, the initial dispatch block 202 checks whether the next memory transaction in its queue is a bypass write. If so, the process 500 remains in the bypass mode by returning to step 512, and the next memory transaction in the queue (bypass write) is sent to the L3 cache schedule 404 along the bypass path 402. Therefore, after the bypass write has satisfied the conditions for using the bypass path 402, the bypass write queued in order may also be sent to the L3 cache schedule 404 via the bypass path 402 without having to recheck the conditions addressed in steps 506 and 508. If the next memory transaction in the queue of the initial scheduling block 202 is not a bypass write, then in step 518, the initial scheduling block 202 exits the bypass mode, the next memory transaction in the queue receives normal (non-bypass) processing, and the process 500 returns to step 502.
[0034] Modifications are possible in the embodiments described, and other embodiments are possible within the scope of the claims.
[0035] In some embodiments, the stream engine communicates and returns responses to both read and write requests.
[0036] In some embodiments, the processor may include multiple processor cores (an embodiment with multiple processor cores is not shown) having the same Figure 1 Functional couplings with the DMC, stream engine, and PMC are similar to those shown and described in .
[0037] In some embodiments, a bus implementing parallel read or write requests may correspond to different types of read or write requests, such as made to different memory blocks or for different purposes.
[0038] In some embodiments, the stream engine enables the processor core to communicate directly with the lower level cache (e.g., L2 cache), skipping the higher level cache (e.g., L1 cache) to avoid data synchronization issues. This can be used to help maintain memory coherency. In some such embodiments, the stream engine can be configured to only transmit read requests, rather than both read and write requests.
[0039] In some embodiments, an L3 cache or other lower level memory may schedule write requests so that write acknowledgements may be sent to the DMC (or processor core or another lower level memory controller) while the write requests are being scheduled to the corresponding memory pipeline.
[0040] In some embodiments, different memory access pipeline libraries may have different numbers of stages.
[0041] In some embodiments, a processor in an external function may access data stored in the L2 cache; in some such embodiments, coherence between the contents stored in the L2 cache and the contents of caches in other processors in the external function is not guaranteed.
[0042] In some embodiments, a write may be a bypass write if the data involved is too large for a lower level cache (eg, an L1D cache or an L2 cache).
[0043] In some embodiments, a write may be a bypass write if the page attributes mark the write as corresponding to a device type memory region that is not cached by the UMC (L2 cache controller).
[0044] In some embodiments, an L1D cache (or other lower level cache) may cache the data payload of a bypass write.
[0045] In some embodiments, memory consistency rules of the processor prohibit bypassing of memory transactions (memory read requests or memory write requests) other than L1D victim cache writes and L1D writes.
[0046] In some embodiments, the bypass path bypasses writes to a final arbitration stage before being scheduled into the memory pipeline bank of the L3 cache (not shown).
[0047] In some embodiments, the guarantee that a memory write will never be directed to the L2 cache (corresponding to a bypass write) includes a guarantee that this is a selection that will never change, or that such a write to the L2 cache is impossible.
[0048] In some embodiments, the guarantee that a memory write is never directed to the L2 cache includes that the L2 cache does not have a copy or hash of the corresponding data.
[0049] In some embodiments, the guarantee that memory writes will never be directed to the L2 cache may be changed (or enabled if the guarantee was not recently enabled). In such embodiments, while a line is being written to the L2 cache, newly making this guarantee (e.g., changing the corresponding mode register to make the guarantee) may require flushing the corresponding L2 cache. In this case, the cache flush may prevent cache copies of data payloads that are now guaranteed not to be written to the L2 cache from remaining in the L2 cache after the guarantee was made.
[0050] In some embodiments, the cache controller initiates memory transactions only in response to transactions initiated by the processor core 102 or the external function 120 .
[0051] In some embodiments, a read or write request can only be dispatched to a corresponding one of the pipeline libraries 206, for example, depending on the memory address of the data being written or requested, the request load of the pipeline library 206, or a pseudo-random function.
[0052] Modifications are possible in the embodiments described, and other embodiments are possible within the scope of the claims.
Claims
1. A processor system, comprising: a processor core configured to generate a memory write request; Cache memory; and a memory controller having a memory pipeline, the memory controller coupled to control the cache memory and communicatively coupled to the processor core, the memory controller configured to: receiving the memory write request from the processor core; scheduling the memory write request on the memory pipeline; and Concurrently with scheduling a respective one of the memory write requests on the memory pipeline, a write acknowledgement is sent to the processor core, the write acknowledgement confirming that writing of the data payload of the respective memory write request to the cache memory has been completed.
2. The processor system of claim 1 , wherein the memory controller is configured to regulate the scheduling of performing the memory write request actions and the sending of the write confirmation actions if conditions are met that include at least the following: the respective memory write requests will not disrupt ordering and consistency, and if scheduled on a memory pipeline will complete within a bounded amount of time.
3. The processor system of claim 2, wherein the condition is satisfied if the memory pipeline is inherently in-order such that once a memory transaction for a corresponding memory enters the memory pipeline, it will read or write to the memory in the order in which it was scheduled.
4. The processor system of claim 2, wherein the condition is satisfied if the memory pipeline includes one or more pipeline libraries, and each pipeline library is independent of the other pipeline libraries, such that: A write transaction on one of the pipeline libraries does not affect ordering or consistency requirements with respect to write transactions on other of the pipeline libraries, and Within each of the pipeline libraries, write and read memory access requests to the cache memory proceed in the order in which they are scheduled.
5. The processor system according to claim 2, further comprising a system clock, wherein the scheduling of the memory write request action and the sending of the write confirmation action to the processor core are performed simultaneously, which means that the two actions are performed on the same cycle of the system clock.
6. The processor system of claim 2, wherein the processor core is configured to, after receiving a write acknowledgement corresponding to the corresponding memory write request, unpause a process dependent on completed processing of the corresponding memory write request.
7. The processor system of claim 2, wherein the memory controller is configured to suspend the memory write request if the condition cannot be guaranteed to be met.
8. The processor system of claim 2, further comprising a clock, wherein the processor core is configured to generate a memory write request on each cycle of the clock.
9. The processor system of claim 2, wherein the cache memory is a lower level cache memory and the memory controller is a lower level memory controller, the processor system further comprising: higher level cache memory; and a higher level cache memory controller coupled to control the higher level cache memory to process memory transactions and communicatively coupled to the processor core and the lower level memory controller; The lower level memory controller receives the memory write request via the higher level memory controller or sends the write confirmation to the processor core.
10. The processor system of claim 9, wherein the higher level cache memory is more local to, or has a lower response time to, memory transactions generated by the processor core than the lower level cache memory.
11. The processor system of claim 9, wherein the higher level cache memory is a level 1 cache (L1 cache) and the lower level cache memory is a level 2 cache (L2 cache).
12. A method of operating a processor system, the method comprising: generating a memory write request using a processor core; receiving, using a memory controller, the memory write request from the processor core; scheduling the memory write request on a memory pipeline of the memory controller; and Concurrently with the scheduling step scheduling respective ones of the memory write requests on the memory pipeline, a write acknowledgement is sent to the processor core, the write acknowledgement confirming that writing of a data payload of the respective memory write request to a cache memory controlled by the memory controller is complete.
13. The method according to claim 12, further comprising, before the scheduling and the sending, determining whether a condition is satisfied, the condition comprising at least: If scheduled on the memory pipeline, the corresponding memory write request will comply with ordering and consistency requirements and will complete in a bounded amount of time; and The scheduling and the sending of the respective memory write requests are adjusted based on whether the determining determines that the condition regarding the respective memory write requests is satisfied.
14. The method of claim 13, wherein the condition is satisfied if the memory pipeline is inherently in-order such that once a memory transaction to the cache memory enters the memory pipeline, it will be read from or written to the cache memory in the order in which it was scheduled.
15. The method of claim 13, wherein the condition is satisfied if the memory pipeline includes one or more pipeline libraries, and each pipeline library is independent of the other pipeline libraries, such that: A write transaction on one of the pipeline libraries does not affect ordering or consistency requirements with respect to write transactions on other of the pipeline libraries, and Within each of the pipeline libraries, write and read memory access requests to the cache memory proceed in the order in which they are scheduled.
16. The method of claim 13, wherein performing the sending step simultaneously with the scheduling step means performing the sending step and the scheduling step on the same cycle of a clock of the processor system.
17. The method of claim 13, further comprising: After the processor core receives a write confirmation corresponding to the corresponding memory write request, a process in the processor core is unpaused depending on completion of processing of the corresponding memory write request.
18. The method of claim 13, further comprising: If the condition cannot be guaranteed to be met, then the corresponding memory write request is suspended at the memory controller.
19. The method of claim 13, further comprising: The generating step generates a memory write request at each cycle of a clock of the processor system.
20. The method according to claim 13, The processor system further includes a higher level memory controller that controls the higher level cache memory; wherein the cache memory is a lower level cache memory and the memory controller is a lower level memory controller; and Wherein the lower level memory controller receives the memory write request via the higher level memory controller or sends the write confirmation to the processor core.
21. The method of claim 20, wherein the higher level cache memory is a level 1 cache (L1 cache) and the lower level cache memory is a level 2 cache (L2 cache).
22. The method of claim 20, wherein the higher level cache memory is more local to, or has a lower response time to, memory transactions generated by the processor core than the lower level cache memory.
23. The method of claim 20, wherein the lower level memory controller performs the sending step by sending the write acknowledgement to the higher level memory controller.
Citation Information
Patent Citations
Configurable cache for multiple clients
CN102640127A
Processor cache with independent assembly line for accelerating prefetching requests
CN107038125A