A cache design method for recording memory access missing information by using cache line space

CN117472444BActive Publication Date: 2026-09-25HUAZHONG UNIV OF SCI & TECH +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311431298.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-31
Publication Date
2026-09-25
Estimated Expiration
2043-10-31

AI Technical Summary

Technical Problem

由于不同应用和工作负载的访存局部性不同,对访存缺失状态保存寄存器的需求量也不同,因此,配置过多将容易导致资源浪费,配置过少则可能影响性能

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117472444B_ABST
    Figure CN117472444B_ABST
Patent Text Reader

Abstract

The application relates to a cache design method for recording access missing information by using cache line space, and the design is as follows: cache data and access missing information are stored in the same storage space in a shared storage mode; the identification of a cache data block is stored in a different static storage from the cache data and the missing request information, wherein the identification is stored in a plurality of independent static storages, and the cache data and the missing request information are stored in a single static storage; a parallel request processing pipeline and a data processing pipeline are constructed and are respectively used for executing access request processing and executing memory return data processing. Compared with the existing non-blocking cache design supporting a large number of access missing state saving registers, the application stores cache data and access missing information in a shared storage space, fully utilizes the dual-port function of the static storage in the FPGA, and respectively designs a pipeline for the processing of the access request and the processing of the memory return data to access the static storage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer system architecture technology, and in particular to a cache design method that utilizes cache line space to record memory access miss information. Background Technology

[0002] Non-blocking caches use a Miss Status Holding Register (MSHR) to record memory miss information when a cache miss occurs, allowing the cache system to continue operating without blocking even before the miss handling is complete. This is an optimization technique to reduce miss overhead and has been widely used in modern processors. Memory miss information includes the identifier of the missing cache data block and the miss request information for accessing that missing data block. A memory miss status holding register includes an identifier field to record the identifier of the missing cache data block, and several slots to record the corresponding miss request information (called sub-entries, including cache block offset, memory access ID, etc.).

[0003] When a miss occurs in the non-blocking cache, it is necessary to check if there is an entry in the memory miss status saver array that matches the identifier of the missing cached data block. If not, a new entry is allocated to record the memory miss information, and the requested missing data is stored in the next lower level. If an entry exists, no new entry needs to be allocated, and no additional data request needs to be sent. After allocating or finding a memory miss status saver, the information of the miss request is recorded in one of its free sub-slots. When the missing data is returned, the corresponding memory miss status saver is searched again and released. A memory access response is generated based on the information of all its sub-slots, completing the miss handling. If there are no free memory miss status savers when a miss occurs, or if the required memory miss status saver has no free sub-slots, blocking will still occur.

[0004] CN115357292A discloses a non-blocking data cache structure, including a data access module and a non-blocking miss handling module. The data access module is used to process memory access instructions passed from an external processor and divides the loading pipeline into 3 stages and the main pipeline into 4 stages. The non-blocking miss handling module is used to process multiple miss instructions and implement pipelined processing for multiple miss instructions to request data from the next level of memory.

[0005] Non-blocking caches play a crucial role in improving memory access throughput. In Field-Programmable Arrays (FPGAs), multiple parallel processing units can be configured to support a large number of incomplete memory access requests, overlapping memory access processing time to maximize memory bandwidth. To address this, related research has implemented thousands of memory miss state save registers in FPGAs using Static RAM (SRAM) to handle misses caused by a large number of memory access requests. Their results show that for applications with limited memory bandwidth and insensitive memory access latency, configuring a large number of memory miss state save registers can achieve performance gains similar to configuring a large-capacity cache, reducing more miss blocking, and consuming less static storage resources than cache lines.

[0006] While configuring a large number of memory miss state savers can improve performance in Field-Programmable Array (FPGA) applications, proper configuration becomes crucial. Unlike caches, which can reside in the cache indefinitely, memory miss state savers remain idle after the missing data is returned. Because different applications and workloads exhibit varying memory locality, their requirements for memory miss state savers also differ. Therefore, configuring too many savers can lead to resource waste, while configuring too few can negatively impact performance. Properly configuring memory miss state savers becomes challenging when the application's memory locality characteristics are unknown or when the application workload is highly variable.

[0007] Furthermore, on the one hand, there are differences in understanding among those skilled in the art; on the other hand, the applicant studied a large number of documents and patents when making this invention, but due to space limitations, not all details and contents were listed in detail. However, this does not mean that the present invention does not possess the features of these prior art. On the contrary, the present invention already possesses all the features of the prior art, and the applicant reserves the right to add relevant prior art to the background art. Summary of the Invention

[0008] To address the shortcomings of existing technologies, this invention provides a cache design method that utilizes cache line space to record memory miss information. This method enables runtime adaptive adjustment of storage space between cache lines and the memory miss state saver register based on application memory locality behavior, thereby achieving automatic adaptation to different memory access characteristics. Simultaneously, it provides a dual-pipeline design method that supports simultaneous processing of memory access requests and memory return data. This fully utilizes the dual-port functionality of the static memory in the FPGA while overcoming memory port contention caused by shared storage space between cache lines and the memory miss state saver register, thus improving memory access throughput.

[0009] To achieve the above objectives, this invention provides a cache design method that utilizes cache line space to record memory access miss information, which can adopt the following design:

[0010] D1. Implement cached data and memory access miss information in the same storage space through shared storage.

[0011] D2. Separate the identifier of the cached data block from the cached data and miss request information and store them in different static memories. Specifically, separate metadata such as the validity bit, the identifier of the cached data block, the transfer flag bit, and the sub-item counter from the cached data block and miss request information and store them in different static memories. The memory storing the identifier and other metadata (called the identifier memory) is implemented using multiplexed independent memory. The memory storing the cached data block and miss request information (called the data memory) is implemented using a single memory.

[0012] D3. Construct a parallelizable dual pipeline: one pipeline (called the request processing pipeline) handles memory access requests, and the other pipeline (called the data processing pipeline) handles memory return data. By setting up two parallel pipelines, the memory port contention problem caused by the shared memory space of the cache line and the Miss Status Holding Register (MSHR) is overcome. Specifically, when the cache line and the MSHR are allowed to share memory space, if only one pipeline is used, it will be impossible to find memory for the memory access request when processing memory return data and looking up the corresponding MSHR, causing blocking. This invention utilizes the dual-port functionality of the static memory in the FPGA and constructs a dual pipeline to overcome this memory port contention problem.

[0013] This invention utilizes cache line space to record memory miss information, allowing the cache line and the memory miss state saver register to share static memory space. This effectively tolerates a large number of memory misses, reduces miss penalties, and improves system throughput. Furthermore, it can adaptively adjust based on application memory access cache hits at runtime, enabling storage space to switch between cache lines and the memory miss state saver register. This allows the cache to automatically adapt to different memory access characteristics, solving the configuration problem of the memory miss state saver register and improving storage efficiency. In addition, it reduces the additional storage overhead required to implement a large number of memory miss state saver registers.

[0014] Currently, many computer systems employ caching techniques to alleviate database pressure and improve system performance and throughput. Caching avoids frequent database accesses and allows for efficient use of cache space to store cached data, significantly improving system performance. However, if the cache space becomes full, update requests will be blocked, severely impacting system performance and stability. Therefore, efficiently managing cache space and achieving data persistence and consistency is a key technical issue. For example, CN115618336A discloses a cache and its operation method, and a computer device including the cache. The cache includes a data array, a tag array, and a marking array, wherein multiple cache rows, multiple tags, and multiple marking rows correspond one-to-one. Each marking row stores a second number of memory tags associated with a first number of tag storage unit items stored in the corresponding cache row. The memory tags stored in the marking row are mapped to the memory addresses of the associated tag storage unit items. In response to a cache hit query using a memory access address, this solution retrieves the second memory tag corresponding to the memory access address from the marking array, and compares the retrieved first memory tag with the retrieved second memory tag to determine if the first and second memory tags match. This technical solution improves the security of memory access in a computer system by marking memory space, but it cannot achieve simultaneous processing of requests and program execution, thus failing to improve memory access throughput. Compared with the prior art, the technical solution of this application overcomes the memory port contention problem by constructing an executable parallel dual-pipeline processing process. Specifically, the parallel processing process handles memory access requests through a request processing pipeline and memory return data through a data processing pipeline. Since the cache line and the memory miss status saver share the same storage space, if only one pipeline is used, it will be impossible to simultaneously search for the corresponding memory for the memory access request when processing memory return data and looking up the corresponding memory miss status saver, thus causing blocking. Therefore, this invention utilizes the dual-port function of static memory to construct dual pipelines for memory access request processing and memory return data processing respectively, which not only overcomes the memory port contention problem but also prevents blocking problems when misses occur.

[0015] Existing technologies have attempted to implement buffering techniques for non-blocking miss handling by using multiple pipelines to process requests in parallel. For example, CN105955711A discloses a buffering method supporting non-blocking miss handling. This method, when the processor encounters a miss request, temporarily stores the missing request in a miss instruction queue through buffering, allowing the pipeline to continue sending subsequent unrelated requests. The cost of the miss is hidden within the normal processing of unrelated requests, thus reducing memory access latency by lowering the buffered miss cost. In this solution, the buffer consists of a control module and the memory bank. Multiple pipelines share a single buffer. Different miss requests are selected by the control module to enter the memory bank. When a miss request returns from DDR, it is returned to the processing module and then sent back to the control module for unified processing of related miss requests, further reducing the actual miss cost and improving pipeline efficiency. However, this solution requires multiple pipelines to share a single port, leading to port contention, which contradicts the memory port contention problem that this invention aims to overcome. Specifically, in this invention, the request processing pipeline and the data processing pipeline share a second port of the identifier memory. If port contention occurs in either identifier memory, this invention executes operations sequentially according to the priority of write operations in the request processing pipeline, write operations in the data processing pipeline, and read operations in the data processing pipeline, or sequentially according to the priority of write operations in the data processing pipeline, write operations in the request processing pipeline, and read operations in the data processing pipeline. If necessary, the data processing pipeline or the request processing pipeline can be paused. This is completely opposite to the prior art's attempt to avoid pipeline pauses. In other words, this invention, through parallel dual-pipeline processing, can fully utilize the dual ports of the static memory in the FPGA for reasonable and efficient allocation, thereby overcoming the memory port contention problem caused by the shared storage space of the cache line and the memory access miss state save register.

[0016] Preferably, D1 is designed as follows:

[0017] D1.1: When a cache miss occurs, the cache line space where the missing data block would have been placed is used as a memory miss status register to record the memory miss information. A cache miss occurs when the processor cannot find the required data or instruction in the cache, requiring it to retrieve the data from main memory or other slower storage levels. This process consumes more time, thus impacting program execution efficiency. Specifically, while using the cache line space to store application data, this space is also used to temporarily store memory miss information, allowing each storage line to function as both a cache line and a memory miss status register. When a miss occurs, the memory miss information is temporarily stored in the storage line; once the missing data is returned, the storage line is then used as a cache line. In this way, the caching system can adaptively adjust the storage space between cache lines and memory miss status registers at runtime based on the application's cache hit status, automatically adapting to different memory locality characteristics. Simultaneously, it eliminates the additional overhead of implementing a large number of memory miss status registers, improving storage efficiency.

[0018] D1.2: A transfer flag is added to each cache line to determine whether the corresponding cache line is being used as a memory miss state saver register when cached data is being retrieved. Specifically, an additional transfer flag (M) is added to each cache line to determine whether the required cached data is being transferred from memory during cached data retrieval, thus determining whether the corresponding cache line is being used as a memory miss state saver register. Furthermore, a sub-item counter is added to each cache line to count the number of missed requests for accessing the missing data block when it is being used as a memory miss state saver register.

[0019] Preferably, D2 is designed as follows:

[0020] D2.1: The identifiers of cached data blocks are stored using multiple independent static memories (called identifier memories), while cached data and miss request information are stored using a single static memory (called data memory). The storage rows of the data memory correspond one-to-one with the storage rows of each identifier memory. Using multiple identifier memories supports cache address mapping methods similar to multi-way set-associative mapping. Preferably, in this invention, the identifiers of cached data blocks are mapped and stored using the Cuckoo hash algorithm to improve the load factor of hash storage. Specifically, each identifier memory performs the mapping conversion from cached data identifiers to cache addresses using multiple different hash functions. Furthermore, a temporary queue is used to store the memory miss status saver register that was evicted due to hash collisions; this temporary queue uses fully associative lookup. Cached data and miss request information are stored using a single data memory, with the same number of storage rows as the sum of the number of rows in all identifier memories, ensuring a one-to-one correspondence between cached data blocks, miss request information, and identifiers. Each static memory has two independent ports (called first port A and second port B), and each port can perform simultaneous read and write operations on the same address.

[0021] D2.2: Establish a redirected data bypass for the identifier memory. This invention optimizes read memory requests by using a read-only cache. However, when updating metadata such as identifiers, write operations to the memory are still required. Therefore, a redirected data bypass needs to be established for the identifier memory to address data consistency issues in the pipeline. Pipeline access to the data memory occurs in the final pipeline stage, at which point the data in the data memory contains the latest values; therefore, a redirected data bypass is not necessary for the data memory.

[0022] Preferably, the request processing pipeline in this invention is designed as follows:

[0023] D3.1: Request the pipeline to exclusively access the first port of each identifier memory for simultaneous lookup and read;

[0024] And / or exclusively use the first port of the data storage to perform cached data reads or write miss request information;

[0025] And / or allocate or update the memory miss state save register through the second port of the identifier memory, the operation being performed on only one identifier memory at a time.

[0026] Specifically, for a request processing pipeline, the runtime can process it according to the following steps:

[0027] S1. Receive the memory access request, and calculate the identifier of the memory access address according to the selected cache address mapping algorithm to obtain the cache address of the corresponding storage line in each identifier memory.

[0028] S2. Based on the cache addresses obtained in S1, read each memory channel through the first port of the identified memory, and obtain the read result after a set clock cycle.

[0029] S3. Check if there is a memory miss status saver in the temporary queue that matches the memory access address identifier.

[0030] S4. Obtain the storage lines of each identifier memory read in S2, match the identifier of each storage line with the identifier of the memory access address, and combine the query result of S3 to obtain the comprehensive matching result, and determine whether the current request is hit and the hit object (i.e., cache line or memory access miss status saver register).

[0031] Furthermore, if the current request misses, it is handled according to different situations: if all the memory lines read by S2 are used as memory access miss status save registers, one of them is evicted to the temporary queue to store the current memory access miss information; if there is a free position in each memory line, the free position is selected for subsequent processing; if there is no free position in each memory line but there is a cache line, one of the cache lines is evicted for subsequent processing.

[0032] S5. Based on the matching results obtained in S4, perform the following processing operations according to the matching conditions:

[0033] C5.1, Cache line hit: Based on the identifier memory number and cache address of the cache line, calculate the address of the corresponding cache data block in the data memory, read the requested data from the first port of the data memory and generate a memory access response.

[0034] C5.2, Memory Miss Status Register Hit: Update its sub-item counter (increment the sub-item counter by 1) and write it back through the second port of its corresponding identifier memory. Further, as described in C5.1, obtain the address of the corresponding miss request information in the data memory, and based on the old value of the sub-item counter, obtain the position of the next free sub-item slot. Record the current request information into that sub-item slot; the write operation is completed through the first port of the data memory. Specifically, if the request hits the memory miss status register in the temporary queue, then the temporary queue is updated instead of the identifier memory and data memory.

[0035] C5.3 Miss: Select the identifier memory corresponding to an idle or evicted memory row, and perform a write operation through the second port of this identifier memory. This includes writing the identifier of the missing data, setting the valid bit and transfer flag, and initializing the sub-item counter. Further, as described in C5.2, write the current request information to the corresponding address through the first port of the data memory. If it is necessary to evict the memory access miss status saver register, then simultaneously read the original miss request information from that address and store it in the temporary queue. In addition, record the identifier of the missing cached data in the memory request queue and send a data request to memory.

[0036] Preferably, in C5.2 and C5.3, if a write operation to the identifier memory is associated with a read operation in the request processing pipeline or the data processing pipeline, that is, the current write address is the same as the address of these read operations, the current write data needs to be redirected to the relevant pipeline stage to ensure that the read operation can obtain the latest data.

[0037] Preferably, the data processing pipeline in this invention is designed as follows:

[0038] D3.2: The data processing pipeline performs simultaneous lookup and read operations through the second port of each identifier memory;

[0039] and / or read miss request information and write cached data through the second port of the data storage;

[0040] And / or release the memory miss state save register through the second port of the identifier memory, this operation is performed only on one identifier memory at a time.

[0041] Specifically, for a data processing pipeline, the following steps can be followed during runtime:

[0042] S6. Receive the corresponding identifiers of the memory return data and the memory request queue dequeue. Based on the selected cache address mapping algorithm, calculate the target storage address of the current data in each identifier memory.

[0043] S7. Based on the addresses obtained in S6, read the respective identifier memories through the second port of the identifier memory, and obtain the reading results after a set clock cycle.

[0044] Preferably, if a write operation is being performed on a certain identifier memory, then reading from that identifier memory is ignored. That is, valid bits read from that memory are considered invalid.

[0045] S8. Check if there is a memory miss status saver register in the temporary queue that matches the current data identifier. If so, release it and read the value of the corresponding sub-item counter and the miss request information for subsequent processing.

[0046] S9. Match the contents of each valid memory line read in S7 with the identifier of the current data, and combine this with the matching result obtained in S8 to obtain a comprehensive matching result. If no memory access miss status saver matching the current data identifier is read in either the validly read identifier memory or the temporary queue, and only one identifier memory was not validly read in S7 due to writing, then it can be determined that the target memory access miss status saver must be in the corresponding position of the identifier memory, and this can still be considered a hit.

[0047] Preferably, for any identifier memory, if the memory address read is the same as the memory address read by the request processing pipeline in S4, the two pipelines may write to the same address of the identifier memory. In this case, it is necessary to pause and release the pipeline to avoid writing to the same address at the same time, or optionally, pause the request processing pipeline.

[0048] S10. Based on the matching result obtained in S9, perform the release operation according to the following matching conditions:

[0049] C10.1 Hit the memory miss status register in the identifier memory: cancel its transfer flag, making the memory line a cache line, and write the contents back to the corresponding identifier memory through the second port of the identifier memory. Further, perform simultaneous read and write operations on the corresponding address through the second port of the data memory, read the sub-entry information (i.e., the miss request information) of the memory miss status register, and write it to the cached data returned from memory. If the write operation to the identifier memory is associated with a read operation in the request processing pipeline or the data processing pipeline (i.e., the current write address is the same as the address of these read operations), to ensure that the read operation obtains the latest data, the current write data needs to be redirected to the relevant pipeline stage.

[0050] C10.2, Memory Miss Status Save Register in Hit Temporary Queue: Obtain the value of the corresponding sub-item counter read by S8 and the miss request information from the temporary queue.

[0051] C10.3, No memory access miss status register was hit: This indicates that due to port contention, the identifier memory where the target memory access miss status register is located could not be effectively read. The current data and its identifier need to be added back to the data processing pipeline for execution, and subsequent operations should be skipped, i.e., S6 should be re-executed.

[0052] S11. Generate a memory access response based on the sub-counter value of the released memory access miss status save register, the read miss request information, and the returned cache data.

[0053] Preferably, D3 further includes:

[0054] D3.3: The request processing pipeline and the data processing pipeline share the second port of the identifier memory. If there is port contention in any identifier memory, the operation is executed in the order of priority of write to the request processing pipeline, write to the data processing pipeline, and read to the data processing pipeline, or in the order of priority of write to the data processing pipeline, write to the request processing pipeline, and read to the data processing pipeline.

[0055] Preferably, when a contention occurs between a request processing pipeline write and a data processing pipeline write, the data processing pipeline is paused or the request processing pipeline is paused.

[0056] Preferably, D3 further includes:

[0057] D3.4: For the second port of the identifier memory, if a read operation of the data processing pipeline is contention with a write operation of any pipeline, the data processing pipeline will still perform read operations on the remaining identifier memories that are not being written to. If the required memory access miss status saver is not read as a result, and only one identifier memory is contentioned and not read, it can be confirmed that the required memory access miss status saver is in the identifier memory that was not read, and it can still be considered a hit. If multiple identifier memories are contentioned and not read, the current release operation will be returned to the release pipeline entry point for re-execution.

[0058] Preferably, the request processing pipeline allocates or updates the memory miss status save register through the second port of the identifier memory, and the data processing pipeline releases the memory miss status save register through the second port of the identifier memory. Each of the request processing pipeline and the data processing pipeline performs the above operations on only one identifier memory per cycle.

[0059] The present invention includes a request processing pipeline for processing memory access requests and a data processing pipeline for processing data returned from memory, each step being completed within one hardware clock cycle (pipeline stage).

[0060] This invention implements a parallel dual-pipeline design, handling memory access requests and memory return data separately. It fully utilizes the dual ports of the static memory in the FPGA and allocates them efficiently, overcoming the memory port contention problem caused by the shared storage space of the cache line and the memory miss state saver register, thus improving the throughput of memory access request processing. Simultaneously, it fully considers issues such as data read / write consistency, memory port contention, and writes to the same address in the dual-pipeline design, ensuring correctness.

[0061] Preferably, the present invention also provides a cache structure, which may include:

[0062] Multiple independent identifier memories are used to store identifiers for data blocks;

[0063] The data storage has multiple storage spaces corresponding to each of the identifier storages, and each storage space is used to store cached data blocks and miss request information;

[0064] The request processing pipeline is used to perform memory access request processing between the multiplexed identifier memory and the data memory;

[0065] A data processing pipeline is used to perform memory-return data processing between multiplexed identifier memory and data memory. Attached Figure Description

[0066] Figure 1 This is a schematic diagram of the hardware module composition of a cache design that utilizes cache line space to record memory access miss information, according to a preferred embodiment of the present invention.

[0067] Figure 2 This is a schematic diagram of a preferred embodiment of the storage structure provided by the present invention;

[0068] Figure 3 This invention provides a preferred embodiment of a dual-pipeline architecture design and processing flow.

[0069] Figure 4 This is a preferred example of performing access to each port of the identifier memory based on a dual-pipeline architecture provided by the present invention. Detailed Implementation

[0070] The preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby providing a clearer and more explicit definition of the scope of protection of the present invention. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited to the embodiments described herein.

[0071] This invention provides a cache design method that utilizes cache line space to record memory access miss information, comprising the following steps:

[0072] D1: Implement cache lines and the Miss Status Holding Register (MSHR) through shared storage, that is, store cache data and memory miss information in the same storage space.

[0073] D2: Separate the identifier of the cached data block from the cached data block and the miss request information and store them in different static storage. The identifier of the cached data block is stored in multiple independent static storage (called the identifier storage), while the cached data block and the miss request information are stored in a single static storage (called the data storage).

[0074] D3: Construct parallelizable request processing and data processing pipelines, where the request processing pipeline is used to perform memory access request processing, and the data processing pipeline is used to perform memory return data processing.

[0075] The cache design method described in this invention, which utilizes cache line space to record memory access miss information, provides a method such as... Figure 1 The cache structure shown includes the following hardware components:

[0076] a. An identifier memory array is used to store identifiers of cached data blocks and missing cached data blocks (i.e., memory miss status savers). When processing an application memory access request, the identifier memory is queried based on the identifier of the required data block to check whether the data block is already in the cache or whether the corresponding memory miss information already exists. When processing a data block returned from memory, the identifier memory is queried based on the identifier of the data block to obtain the corresponding memory miss status saver register for the missed request to access that data block.

[0077] b. Data Memory: Used to store application cached data blocks and miss request information (i.e., a sub-item of memory access miss information). When processing an application memory access request, after finding a valid cached data block in the identifier memory array, the contents of that data block are read from the data memory and returned to the application; or after finding a valid memory access miss status saver, the corresponding miss request information is updated and appended to the data memory. When processing a data block returned from memory, the corresponding miss request information (i.e., the memory access request information for that data block) is read from the data memory, and the data block is filled into the data memory.

[0078] c. Temporary queue: Used to store memory miss status save registers that have been evicted from memory due to hash collisions. When no matching memory row for the current memory access request address is found at the corresponding cache address in any of the identifier memories, and all of these memory rows are valid memory miss status save registers, one of them must be evicted to the temporary queue, and the memory miss information of the current memory access request must be stored in the space of the evicted memory row.

[0079] d. The memory access response generation module is used to parse the memory access missing information and generate the corresponding memory access data response based on the recorded miss request information (i.e., sub-items) when a missing data block is returned.

[0080] According to a preferred embodiment, the cache design method of the present invention, which utilizes cache line space to record memory access miss information, employs the following... Figure 2 The storage organization structure is shown. This storage organization structure includes a multi-way memory array consisting of several identifier memories and a data memory. Each identifier memory corresponds one-to-one with each data storage space in the data memory. Specifically, the identifier memory is used to store the relevant identifier information of cached data, and the data memory is used to store cached data blocks and miss request information (i.e., sub-items of memory access miss information).

[0081] According to a preferred embodiment, such as Figure 2 As shown, cache line space is used to record memory miss information, identifying the relevant identifiers of cached data stored in the memory, and a counter for the sub-items recording memory miss information. The relevant identifiers of the cached data include its validity bit V, transfer flag bit M, and identifier. The data memory stores cached data blocks and miss request information (or sub-items of memory miss information). Specifically, when a miss occurs, the cache line where the missing data will be placed is used as a memory miss status saver. Further, when the missing data is returned, the data is placed back into that line, making it a cache line.

[0082] According to a preferred embodiment, such as Figure 2 As shown, metadata such as the valid bit (V), the identifier of the cached data block, the transfer flag bit (M), and the sub-item counter are stored separately from the cached data block and miss request information in different static memories (i.e., identifier memory and data memory). Similar to multiplexed set-associative mapping, the identifier of the cached data block is mapped and stored using multiple Cuckoo hash functions to improve the load factor of hash storage. Therefore, multiplexed memory is used for storage to enable simultaneous reading. The mapping and conversion from cached data identifier to cache address is performed using multiple different hash functions, with each corresponding to an independent memory.

[0083] Furthermore, cached data blocks and miss request information are stored in a single memory, with the same number of rows as the total number of rows in all identifier memories, ensuring a one-to-one correspondence between data and identifiers. Specifically, the data memory space is divided into the same number of equal parts as the number of identifier memory paths, with one data memory space corresponding to one identifier memory path. In particular, each static memory block has two independent ports (referred to as port A and port B), and each port can perform simultaneous read and write operations on the same address.

[0084] The cache design method for recording memory access miss information using cache line space described in this invention adopts a dual-pipeline architecture design. For example... Figure 3As shown, a dual-pipeline architecture can include a request processing pipeline and a data processing pipeline. Specifically, the request processing pipeline is used to perform memory access request processing. The data processing pipeline is used to perform memory return data processing. In particular, the two pipelines, namely the request processing pipeline and the data processing pipeline, can run in parallel.

[0085] According to a preferred embodiment, such as Figure 3 As shown, the task of handling memory access requests in the request processing pipeline includes the following steps in sequence: calculating the cache address—reading the identifier memory—querying the temporary queue—matching the identifier to determine a hit—processing the matching result based on the preset memory access request processing logic. On the other hand, as... Figure 3 As shown, the memory return data processing task of the data processing pipeline includes the following steps in sequence: calculating the cache address, reading the identifier memory, querying the temporary queue, matching the identifier to determine if a match has occurred, and processing the matching result based on the preset memory data processing logic.

[0086] To more clearly describe the working principle and process of this invention, the hardware modules and their processing flow are explained below in conjunction with practical application scenarios. Assume that the application logic accesses memory data through the cache described in this invention, the size of the data word requested by the application is 1 byte, the size of the cache data block is 16 bytes, and the memory address is addressed byte-wise. The cache uses a 4-way identifier memory array.

[0087] The application accesses a data word at memory address 0xBA75 with memory access ID number 0. The identifier of the data block containing this data word is 0xBA7, and the offset of the accessed data word in the data block is 0x5.

[0088] After the above memory access request enters the request processing pipeline, it can be processed according to the following steps:

[0089] S1, Hash Calculation Cache Address: Upon receiving a memory access request, the identifier 0xBA7 of the access address is calculated using four different hash functions to obtain the cache address of the corresponding storage row in each identifier memory, assuming they are 0x05, 0xB1, 0x3F, and 0x84 respectively.

[0090] S2. Read the identifier memory: Based on the cache addresses obtained in S1, read each memory concurrently through the first port (e.g., port A) of the identifier memory, and obtain the read result after a set clock cycle (e.g., 2 clock cycles).

[0091] S3. Check the temporary queue: Check if there is a memory miss status saver register in the temporary queue that matches the memory address identifier 0xBA7.

[0092] S4. Matching Identifiers to Determine if a Hit Occurs: Obtain the contents of the memory rows (or pipeline redirection contents) of each identifier memory read in S2, and match the identifiers of these memory rows with the identifier 0xBA7 of the memory access address. Combine this with the query result from S3 to obtain a comprehensive matching result, and determine whether the current request has a hit, and whether the hit is a cache line or a memory access miss status saver. If a miss occurs and the memory access miss status saver needs to be evicted, it is placed in a temporary queue.

[0093] S5. Based on the matching result obtained in S4, perform processing operations according to the preset memory access request processing logic. Specifically, the result of the matching identifier judgment can include a request hitting a cache line, a request hitting a memory access miss status saver, and a request missing.

[0094] The different matching results obtained from S4 can be handled as follows:

[0095] C5.1 Requesting a cache line hit: For example, if a valid item with the identifier 0xBA7 is read from address 0x05 of identifier memory 1, and the transfer flag bit M of this item is not set, it indicates a cache line hit. Based on the identifier memory number and cache address of the cache line, calculate the address of the corresponding data in the data memory (e.g., 0x005), read the requested data from data memory port A, select the data word with the requested offset of 0x5, and generate a memory access response.

[0096] C5.2 Requesting a Hit on a Missed Access Register: For example, a valid entry with the identifier 0xBA7 is read from address 0xB1 of identifier memory 2, and the transfer flag bit M of this entry is set, indicating a hit on the missed access register. At this time, its sub-entry counter is incremented by 1, and the updated content is written back through port B of its identifier memory 2. Further, as described in C5.1, the address of the corresponding missed request information in the data memory (e.g., 0x1B1) is obtained, and based on the old value of the sub-entry counter, the position of the next free sub-entry slot is determined. The current missed request information (intra-block offset: 0x5, memory access ID: 0) is recorded in this slot, and the write operation is completed through port A of the data memory. Specifically, if the hit is on the missed access register in the temporary queue, the corresponding entry in the temporary queue is updated without requiring operation on the identifier memory and data memory.

[0097] C5.3, Request Miss: No entry with identifier 0xBA7 is found at the corresponding address of each identifier memory calculated according to S1. Select the identifier memory corresponding to an idle or evicted memory row. Assuming identifier memory 3 is selected, at address 0x3F, allocate a memory access miss status saver register for the current request via port B. Specifically, write the identifier 0xBA7 of the missing data, set the valid bit (V) and transfer flag bit (M), and initialize the sub-item counter to 1. Further, as described in C5.2, write the current miss request information (intra-block offset: 0x5, memory access ID: 0) to the corresponding address (e.g., 0x23F) via data memory port A. Specifically, if it is necessary to evict the existing memory access miss status saver register at address 0x23F, simultaneously read its corresponding miss request information and store it in a temporary queue. In addition, record the identifier of the missing cached data in the memory request queue and send a data request to memory.

[0098] Once the missing data block marked 0xBA7 is returned, it enters the data processing pipeline for processing, which can be done according to the following steps:

[0099] S6, Hash Calculation Cache Address: The corresponding identifier 0xBA7 for receiving data returned from memory and dequeuing from the memory request queue is used to calculate the target storage address of the current data's corresponding storage row in the identifier memory (the same as the result obtained in S1), which are 0x05, 0xB1, 0x3F, and 0x84 respectively.

[0100] S7. Read the identifier memory: Based on the addresses obtained in S6, read each identifier memory concurrently through the second port (port B) of the identifier memory, and obtain the reading result after a set clock cycle (such as 2 clock cycles).

[0101] Specifically, if a certain identifier memory is being written to, i.e., port B is occupied, then reading from that identifier memory is ignored. Figure 4 The illustrated embodiment describes the read / write behavior within a certain clock cycle when a 4-way identifier memory is configured. When two pipelines write to identifier memories 1 and 2 through port B respectively, the data processing pipeline concurrently reads from identifier memories 3 and 4 through port B.

[0102] S8. Query the temporary queue: Query whether there is a memory miss status saver register in the temporary queue that matches the current data identifier 0xBA7. If so, release it and read the value of the corresponding sub-item counter and the miss request information for subsequent processing.

[0103] S9. Matching Identifier to Determine a Hit: Obtain the valid memory content (or pipeline redirection content) read in S7 and match it with the identifier 0xBA7 of the current data. Combine this with the matching result obtained in S8 to obtain the comprehensive matching result. If no memory access miss status register matching the current data identifier 0xBA7 is read in either the validly read identifier memory or the temporary queue, and only one identifier memory was not validly read in S7 due to writing, then it can be determined that the target memory access miss status register with identifier 0xBA7 is in the corresponding position of that identifier memory, and this can still be considered a hit. For example, when reading the identifier memory in S7, identifier memory 2 cannot be read because it is being written to, and no memory access miss status register with identifier 0xBA7 is read from identifier memories 1, 3, 4, or the temporary queue, then it can be determined that the target memory access miss status register should be located at address 0xB1 of identifier memory 2. Specifically, for a certain identifier memory, if the memory address read is the same as the memory address read by the request processing pipeline in S4, the two pipelines may write to the same address of the identifier memory. In this case, it is necessary to pause the request processing pipeline or the data processing pipeline to avoid writing to the same address at the same time.

[0104] S10. Based on the matching result obtained in S9, perform processing operations according to the preset memory data processing logic. Specifically, process according to the following matching cases:

[0105] C10.1 Hitting the memory miss status register in the identifier memory: For example, if the corresponding entry is found at address 0x84 in identifier memory 4, its transfer flag (M) is canceled, and the content is written back to the corresponding identifier memory through port B, making that memory line a cache line. Simultaneous read and write operations are performed on the corresponding address (e.g., 0x384) through port B of the data memory to read the sub-entry information (i.e., the miss request information) of the memory miss status register and write it to the cached data returned from memory.

[0106] C10.2, Memory Miss Status Save Register in Hit Temporary Queue: Obtain the value of the corresponding sub-item counter read by S8 and the miss request information from the temporary queue.

[0107] C10.3, No memory access miss status register was hit: This indicates that due to port contention, the identifier memory where the target memory access miss status register is located could not be effectively read. The current data and its identifier need to be added back to the data processing pipeline for execution, and subsequent operations should be skipped, i.e., S6 should be re-executed.

[0108] S11. The sub-counter value of the released memory miss status saver, the read miss request information, and the cache data block with the identifier 0xBA7 are passed to the memory access response generation module for parsing and processing to generate a memory access response. For example, if the memory miss information recorded in the released memory miss status saver includes: the identifier of the missing cache data block is 0xBA7, the sub-counter is 1 (indicating that only one miss request is recorded), the offset of the data word required by the miss request within the data block is 0x5, and its memory access ID is 0, then the data word with the offset of 0x5 in the data block with the identifier 0xBA7 required by the request is selected, a memory access response with a response ID of 0 is generated, and sent to the application.

[0109] This invention provides a cache design method that uses cache line space to store application data while simultaneously utilizing that space to temporarily store memory miss information. This method is suitable for hardware architectures with high-concurrency memory access characteristics, such as Field-Programmable Arrays (FPGAs). Unlike methods that use a separate MissStatus Holding Register (MSHR) to store memory miss information, this invention allows cache lines and the MSHR to share static RAM space, thereby reducing the total demand for static RAM and enabling support for a large number of MSHRs. It also allows for adaptive adjustment at runtime based on application memory access cache hits, enabling the storage space to switch between cache lines and MSHRs, thus allowing the cache to automatically adapt to different memory access characteristics.

[0110] Compared to existing non-blocking cache designs that support numerous memory miss status savers, this invention provides a cache design method that utilizes cache line space to record memory miss information. While still supporting non-blocking characteristics (i.e., processing numerous misses simultaneously), it can adapt to the locality of memory access behavior of applications, eliminating the need for separate memory miss status saver configurations for different applications and avoiding resource waste or performance limitations caused by improper configuration. Furthermore, this invention improves the throughput of the cache system for handling memory accesses by implementing parallel dual pipelines.

[0111] Furthermore, this invention fully utilizes the dual-port functionality of static memory in FPGAs, designing separate pipelines for processing memory access requests and memory return data, enabling parallel processing of both. This overcomes the memory port contention problem caused by cache lines and memory miss state save registers sharing storage space, thereby improving the throughput of memory access processing.

[0112] Those skilled in the art will understand that, as long as the objectives of the present invention can be achieved, other steps or operations may be included before, after, or between the steps described above, for example, to further optimize and / or improve the method described in the present invention. Furthermore, although the method described in the present invention is shown and described as a series of actions performed sequentially, it should be understood that the method is not limited by the order. For example, some actions may occur in a different order than that described herein. Alternatively, one action may occur simultaneously with another action.

[0113] It should be noted that the specific embodiments described above are exemplary. Those skilled in the art can devise various solutions inspired by the disclosure of this invention, and these solutions all fall within the scope of this invention and its protection. Those skilled in the art should understand that this specification and its accompanying drawings are illustrative and not intended to limit the scope of the claims. The scope of protection of this invention is defined by the claims and their equivalents. This specification contains multiple inventive concepts; terms such as "preferredly," "according to a preferred embodiment," or "optionally" indicate that the corresponding paragraph discloses an independent concept. The applicant reserves the right to file divisional applications based on each inventive concept.

Claims

1. A cache design method that utilizes cache line space to record memory access miss information, characterized in that, The following design is adopted: D1: Cache data and memory access miss information are stored in the same storage space using a shared storage method; wherein, the method adopts the following design: D1.1: If a cache miss occurs, the cache line space used to store the missing data block will be used as a memory miss status saver to record the memory miss information; D1.2: Add a transfer flag bit to each cache line so that when looking up cached data, the transfer flag bit is used to determine whether the corresponding cache line is used as a memory access miss state save register. D2: Separate the identifier of the cached data block from the cached data block and the miss request information and store them in different static storage; wherein, the method adopts the following design: D2.1: The static memory used to store the identifiers of the cached data blocks uses multiple independent identifier memories, and the static memory used to store the cached data blocks and miss request information uses a single data memory, wherein the storage rows of the data memory correspond one-to-one with the storage rows of each identifier memory; D2.2: Establish a redirected data bypass for the identifier memory; D3: Construct parallelizable request processing and data processing pipelines, wherein the request processing pipeline is used to perform memory access request processing, and the data processing pipeline is used to perform memory return data processing; wherein the method adopts the following design: D3.1: Request the pipeline to exclusively access the first port of each of the aforementioned identifier memories for simultaneous lookup and read; And / or exclusively use the first port of the data storage to perform cached data reads or miss request information writes; And / or allocate or update the memory miss state save register through the second port of the identifier memory; D3.2: The data processing pipeline performs simultaneous lookup and read operations through the second port of each of the aforementioned identifier memories; And / or read miss request information and write cached data through the second port of the data storage; And / or release the memory miss state save register through the second port of the identifier memory; D3.3: The request processing pipeline and the data processing pipeline share the second port of the identifier memory. If there is port contention in any identifier memory, the operation is executed in the order of priority of writing to the request processing pipeline, writing to the data processing pipeline, and reading to the data processing pipeline, or in the order of priority of writing to the data processing pipeline, writing to the request processing pipeline, and reading to the data processing pipeline. D3.4: For the second port of the identifier memory, if a read operation of the data processing pipeline is contested by a write operation of any pipeline, the data processing pipeline will perform a read operation on the remaining identifier memories that are not being written to. If the required memory access miss status saver is not found, and only one identifier memory is contested and not read, then it is confirmed that the required memory access miss status saver is in the identifier memory that has not been read, and it is considered a hit. If multiple identifier memories are contested and not read, then the current release operation is returned to the entry point of the data processing pipeline for re-execution.

2. The cache design method according to claim 1, characterized in that, The D1.1 also includes: When the missing data block is returned, a memory access response is generated based on the miss request information recorded in the corresponding memory access miss status save register and released so that it can be used as a cache line, allowing the data block to be placed in the data field of the cache line.

3. The cache design method according to claim 1, characterized in that, The D1.2 also includes: Add a sub-counter to each cache line to count the number of missed request messages for accessing the missing data block when it is used as a memory miss status saver register.

4. The cache design method according to claim 1, characterized in that, The D3.4 also includes: If a port contention occurs between the write operation of the request processing pipeline and the write operation of the data processing pipeline, the data processing pipeline or the request processing pipeline may be paused.

5. The cache design method according to claim 1, characterized in that, The request processing pipeline allocates or updates the memory miss status saver register through the second port of the identifier memory, and the data processing pipeline releases the memory miss status saver register through the second port of the identifier memory. Each of the request processing pipeline and the data processing pipeline performs the above operations on only one identifier memory per cycle.

Citation Information

Patent Citations

  • Buffering method capable of supporting non-blocking missing processing

    CN105955711A

  • Cache line state

    CN109614348A

  • Multi-port configurable cache access method and device for reconfigurable processor

    CN115421899A