Caching method, cache, electronic device, and readable storage medium

By merging requests of the same address in the missed state processing register and sending the merged request to the main pipeline, the problem of request blocking at the cache entry is solved and the cache performance is improved.

WO2025130581A1PCT designated stage expired Publication Date: 2025-06-26BEIJING INSTITUTE OF OPEN SOURCE CHIP

Patent Information

Application Number
PCT/CN2024/136266
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-19
Filing Date
2024-12-03
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

In processors that support prefetching technology, there is a problem of untimely prefetching, which leads to request blocking at the cache entry, affecting overall performance.

Method used

By merging requests for the same address in the missed state processing register, a merged third request is generated and sent to the main pipeline, reducing the blockage of requests in the cache entry.

Benefits of technology

Reduces the blockage of requests in the cache entry, allowing subsequent requests to enter the cache smoothly, improves the overall performance of the cache, and reduces the pressure of the mainstream pipeline and the number of cache read and write times.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024136266_26062025_PF_FP_ABST
    Figure CN2024136266_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present application relate to the technical field of computers. Provided are a caching method, a cache, an electronic device and a readable storage medium. The method comprises: when a miss status handling register has a second request, a request address in which is the same as a first request address, a request buffer transmitting a first request into the miss status handling register, and merging the first request and the second request to obtain a third request; the miss status handling register generating a handling task on the basis of the third request, and sending the handling task to a main pipeline; the main pipeline accessing a cache line of a cache on the basis of the handling task, and generating a handling result; and a grant buffer generating a first grant for the first request and a second grant for the second request on the basis of the handling result, and sending the first grant and the second grant to a first node in parallel. The embodiments of the present application can improve the overall performance of a cache.
Need to check novelty before this filing date? Find Prior Art

Description

Cache method, high-speed cache, electronic device and readable storage medium

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 19, 2023, with application number 202311753900.9 and invention name “A caching method, cache, electronic device and readable storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of computer technology, and in particular to a cache method, a high-speed cache, an electronic device, and a readable storage medium. Background Art

[0003] In processors that support prefetching technology, there is a large proportion of untimely prefetching, that is, although the prefetcher predicts the data needed in the future, the request is sent late. When the cache miss generated by the prefetch is still waiting for data to be returned in the Miss Status Handling Register (MSHR), the acquire request for the same address has already arrived at the cache. In general cache design, requests with the same address as the prefetch task recorded in the MSHR will be blocked at the cache entry and wait for the data of the prefetch task to be backfilled before entering the main pipeline to access the cache line. This will block other subsequent requests with the same address, causing the buffer at the cache entry used to buffer blocked requests to be occupied. If the buffer is full, all subsequent requests to access the cache cannot enter the cache, affecting the overall performance of the cache. Summary of the Invention

[0004] The embodiments of the present application provide a caching method, a high-speed cache, an electronic device, and a readable storage medium, which can solve the problem in related technologies that if a request arriving at the cache has the same address as a request recorded in the MSHR, it will be blocked at the cache entry, thereby blocking other subsequent requests with the same address, affecting the overall performance of the cache.

[0005] To solve the above problems, an embodiment of the present application discloses a caching method, which is applied to a high-speed cache. The high-speed cache includes a request buffer, a miss status processing register, a main pipeline, and a response buffer. The method includes:

[0006] The request buffer receives a first request sent by a first node, where the first request carries a first request address;

[0007] When a second request having the same request address as the first request address exists in the miss status processing register, the request buffer transfers the first request to the miss status processing register, and merges the first request and the second request to obtain a third request; the third request carries the first request information of the first request and the second request information of the second request;

[0008] The miss status processing register generates a processing task according to the third request, and sends the processing task to the main pipeline;

[0009] The main pipeline accesses the cache line of the cache according to the processing task and generates a processing result;

[0010] The response buffer generates a first response to the first request and a second response to the second request according to the processing result, and sends the first response and the second response to the first node in parallel.

[0011] Optionally, when a second request having the same request address as the first request address exists in the miss status processing register, the request buffer transfers the first request to the miss status processing register, and merges the first request and the second request to obtain a third request, including:

[0012] The request buffer queries, according to the first request address, whether there is a target register entry in the miss status processing register, the target register entry being used to record second request information of a second request, the second request address of the second request being the same as the first request address;

[0013] When the target register entry exists in the miss status processing register, the request buffer adds the first request information of the first request to the target register entry to generate a third request.

[0014] Optionally, the first request includes a get request, and the second request includes a prefetch request;

[0015] The miss status processing register generates a processing task according to the third request and sends the processing task to the main pipeline, including:

[0016] The miss status processing register obtains a target data block from a downstream node according to second request information in the third request;

[0017] The miss status processing register generates a processing task according to the third request, and sends the processing task to the main pipeline when receiving the target data block returned by the downstream node.

[0018] Optionally, the main pipeline accesses the cache line of the cache according to the processing task and generates a processing result, including:

[0019] The main pipeline determines a first cache line from the cache and writes the target data block into the first cache line;

[0020] The main pipeline updates the state of the target data block to a first state, where the first state is used to indicate that the acquisition request has been completed.

[0021] Optionally, the main pipeline determines a first cache line from the cache and writes the target data block into the first cache line, comprising:

[0022] The main pipeline sends a read request to a directory of the cache to read status information of each cache line in the cache, and determines a first cache line from each cache line of the cache by using a cache replacement strategy;

[0023] The main pipeline writes the target data block into the first cache line.

[0024] Optionally, the method further includes:

[0025] The main pipeline writes the original data in the first cache line back to the downstream node.

[0026] On the other hand, an embodiment of the present application discloses a cache, the cache comprising a request buffer, a miss status processing register, a main pipeline, and a response buffer;

[0027] The request buffer is configured to receive a first request sent by a first node, the first request carrying a first request address; if a second request having the same request address as the first request address exists in the miss status processing register, transfer the first request to the miss status processing register, and merge the first request and the second request to obtain a third request; the third request carrying first request information of the first request and second request information of the second request;

[0028] The miss status processing register is used to generate a processing task according to the third request and send the processing task to the main pipeline;

[0029] The main pipeline is used to access the cache line of the cache according to the processing task and generate a processing result;

[0030] The response buffer is used to generate a first response to the first request and a second response to the second request according to the processing result, and send the first response and the second response to the first node in parallel.

[0031] On the other hand, an embodiment of the present application further discloses an electronic device, which includes a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; the memory is used to store executable instructions, and the executable instructions enable the processor to execute the aforementioned caching method.

[0032] On the other hand, an embodiment of the present application further provides a computer program, comprising computer-readable code, which, when executed on a computing and processing device, causes the computing and processing device to execute the aforementioned caching method.

[0033] An embodiment of the present application further discloses a readable storage medium in which the aforementioned computer program is stored. When the instructions in the readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the aforementioned caching method.

[0034] The embodiments of the present application include the following advantages:

[0035] An embodiment of the present application provides a cache method. When a second request having the same request address as the first request address of a received first request exists in a miss status processing register, the method combines the first request and the second request recorded in the MSHR into a third request in the MSHR, thereby reducing the blocking of requests at the cache entry, allowing subsequent requests to enter the cache smoothly, and thus improving the overall performance of the cache. Furthermore, in an embodiment of the present application, the MSHR generates a processing task based on the merged third request and sends the processing task to the mainstream pipeline, i.e., it only needs to enter the mainstream pipeline once, and the mainstream pipeline only needs to access the cache once based on the processing task, thereby reducing the pressure on the mainstream pipeline and the number of cache reads and writes, improving the first node's memory access efficiency to the cache, and thus improving the overall performance of the computer system. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0037] FIG1 is a flowchart of a cache method embodiment of the present application;

[0038] FIG2 is a schematic diagram of the structure of a cache of the present application;

[0039] FIG3 is a structural block diagram of an electronic device provided by an example of the present application;

[0040] FIG4 schematically shows a block diagram of a computing and processing device for executing the method according to the present application;

[0041] FIG5 schematically shows a storage unit for storing or carrying program codes for implementing the method according to the present application. Specific embodiments

[0042] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0043] The terms "first", "second", etc. in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, the term "and / or" in the specification and claims is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects before and after the association are in an "or" relationship. In the embodiments of the present application, the term "multiple" refers to two or more, and other quantifiers are similar.

[0044] Method Example

[0045] 1 , a flowchart of a cache method embodiment of the present application is shown. The method may specifically include the following steps:

[0046] Step 101: The request buffer receives a first request sent by a first node, where the first request carries a first request address.

[0047] Step 102: If the request buffer contains a second request with the same request address as the first request in the miss status processing register, the request buffer transfers the first request to the miss status processing register and merges the first request and the second request to obtain a third request; the third request carries the first request information of the first request and the second request information of the second request.

[0048] Step 103: the miss status processing register generates a processing task according to the third request, and sends the processing task to the main pipeline;

[0049] Step 104: the main pipeline accesses the cache line of the cache according to the processing task and generates a processing result;

[0050] Step 105: The response buffer generates a first response to the first request and a second response to the second request according to the processing result, and sends the first response and the second response to the first node in parallel.

[0051] The caching method provided in the embodiments of the present application is applied to a high-speed cache. A high-speed cache is a primary memory located between the main memory (or main memory) and the processor, typically composed of a static random access memory (SRAM), and is used to improve the system's memory access speed. The processor may include, but is not limited to, a CPU, a GPU, a data processing unit (DPU), a field programmable gate array (FPGA), and a processing module or processing unit in an application-specific integrated circuit (ASIC), etc.

[0052] 2, which shows a schematic diagram of the structure of a cache provided by an embodiment of the present application. As shown in FIG2, the cache includes a request buffer, a miss status processing register, a main pipeline, and a response buffer.

[0053] The request buffer is used to determine whether the first request sent by the first node needs to be blocked at the entrance of the cache, and if the first request needs to be blocked, further determine whether the first request meets the merge condition. It can be understood that the first node usually refers to the upstream node of the cache, that is, the node adjacent to the cache in the computer storage hierarchy and closer to the processor side, such as registers, processors, etc. The computer storage hierarchy divides various storage components (including registers, caches, memory, hard disks, etc.) into different levels based on operating speed and unit cost. The closer to the processor end, the faster the operating speed of the storage component, the smaller the capacity, and the higher the cost per unit capacity; the closer to the memory end, the larger the capacity of the storage component, the slower the operating speed, and the lower the cost per unit capacity.

[0054] In the embodiment of the present application, the first request may be any request sent by the first node to the cache. For example, the first request may be a memory access request sent by a processor to the cache, or a get request or read request sent by another upstream node in the computer storage hierarchy to the cache. If the first request address carried in the first request is the same as the request address of the second request recorded in the miss status processing register, the first request may be considered to meet the merge condition.

[0055] It should be noted that the Miss Status Handling Register (MSHR) is a structure used to handle cache misses. When a cache access miss occurs, the relevant data needs to be obtained from a lower-level cache or main memory, which usually takes some time. MSHR exists to optimize this process and ensure that the system can operate more efficiently when processing multiple cache miss requests. When the cache receives an Acquire request or Release request from an upstream node, or a Probe request from a downstream node, it allocates an MSHR for the request and obtains the permission information of the request address in the current cache and the previous cache by reading the directory.

[0056] In general cache design, requests with the same address as prefetch requests or prefetch tasks in MSHR will be blocked at the cache entry, waiting for the data backfill of the prefetch request to be completed before entering the mainstream pipeline to access the cache line. MSHR will block subsequent requests with the same address as the prefetch request, so that the request buffer at the cache entry used to buffer blocked requests is occupied. If the request buffer is full, all subsequent requests to access the cache cannot enter, affecting the overall performance of the cache.

[0057] In an embodiment of the present application, when the request buffer determines that the first request address of the received first request is the same as the request address of the second request recorded in the MSHR, the request buffer directly passes the first request into the MSHR, merges the first request and the second request in the MSHR, and obtains a third request. It can be understood that the third request contains both the first request information of the first request and the second request information of the second request. Exemplarily, the request buffer can write the first request information of the first request into the second request information of the second request in the MSHR, so that the first request information of the first request and the second request information of the second request are recorded in the modified MSHR at the same time, without the need to allocate a new MSHR for the third request. Alternatively, the request buffer can also write the first request information of the first request and the second request information of the second request into a new MSHR, and this new MSHR is used to record the merged third request, and then release the MSHR originally occupied by the second request.

[0058] MSHR generates a processing task based on the third request. It should be noted that if the first request and the second request are not merged, the second request in MSHR needs to enter the mainstream pipeline during the data backfill process, and the mainstream pipeline searches and compares the cache line to determine whether it hits; the parsing process of the first request also needs to enter the mainstream pipeline, and the mainstream pipeline accesses the cache. In the embodiment of the present application, MSHR directly generates a processing task based on the third request and sends the processing task to the mainstream pipeline, thereby merging the two entries of the first request and the second request into one, reducing the pressure on the mainstream pipeline.

[0059] The main pipeline accesses the cache line of the Cache according to the processing task sent by the MSHR, processes the first request information of the first request, and processes the second request information of the second request, and generates a processing result. For example, assuming that the first request is an acquisition request and the second request is a prefetch request, then after receiving the processing task, the main pipeline first obtains the prefetch block required for the prefetch request from the downstream node, writes the prefetch block into the Cache, and determines the data required for the acquisition request based on the prefetch block, and generates the corresponding processing result. Alternatively, assuming that the first request is a release request and the second request is a prefetch request, then after receiving the processing task, the main pipeline can directly update the status of the data block corresponding to the request address of the release request and the prefetch request to the status after the release request and the prefetch request are completed, and so on.

[0060] It should be noted that in the embodiment of the present application, the mainstream pipeline only needs to access the cache once based on the processing task. Compared with before the merger, the mainstream pipeline needs to access the cache once for the parsing and response process of the first request, and needs to access the cache once for the parsing and response process of the second request. The embodiment of the present application reduces the number of cache reads and writes, improves the memory access efficiency of the first node to the cache, and is conducive to improving the overall performance of the computer system.

[0061] The Grant Buffer generates a first response for the first request and a second response for the second request based on the processing results of the main pipeline, and sends the first and second responses in parallel to the first node. As an example, assuming that the first request is a get request and the second request is a prefetch request, the first response is a get response, which indicates that data acquisition is complete and carries the acquired data; the second response is a prefetch response, which indicates that prefetch is complete.

[0062] The cache method provided in the embodiment of the present application reduces the blocking of requests at the cache entry by merging the first request and the second request recorded in the MSHR into a third request in the case where there is a second request with the same request address as the first request received in the miss status processing register, so that subsequent requests can smoothly enter the cache, which is beneficial to improving the overall performance of the cache. In addition, in the embodiment of the present application, the MSHR generates a processing task based on the merged third request and sends the processing task to the mainstream pipeline, that is, it only needs to enter the mainstream pipeline once, and the mainstream pipeline only needs to access the cache once based on the processing task, thereby reducing the pressure on the mainstream pipeline and the number of cache reads and writes, improving the first node's memory access efficiency to the cache, and helping to improve the overall performance of the computer system.

[0063] In an optional embodiment of the present application, if the request buffer in step 102 has a second request with the same request address as the first request address in the miss status processing register, the first request is transferred to the miss status processing register, and the first request and the second request are merged to obtain a third request, including:

[0064] Step S11: The request buffer queries, based on the first request address, whether a target register entry exists in the miss status processing register, where the target register entry is used to record second request information of a second request, and the second request address of the second request is the same as the first request address;

[0065] Step S12: If the target register entry exists in the miss status processing register, the request buffer adds the first request information of the first request to the target register entry to generate a third request.

[0066] In the MSHR, each request corresponds to an MSHR item, and the MSHR item records the request information of the request, wherein the request information may include the request address, request type, etc. of the request.

[0067] After receiving the first request, the request buffer can query the MSHR based on the first request address of the first request to determine whether there is a target register item in the MSHR. The MSHR item records the second request information of the second request. If the second request address of the second request is the same as the first request address of the first request, the MSHR item can be considered to be the target register item in this application. If the request buffer finds the target register item, it can be considered that there is a second case in the MSHR where the request address is the same as the first request address of the first request, that is, the first request meets the merge condition. Therefore, the first request and the second request recorded in the MSHR are merged: the first request information of the first request is added to the target register item where the second request is located. At this time, the target register item records the first request information of the first request and the second request information of the second request. For ease of distinction, the request recorded in the merged target register item is recorded as the third request. In this embodiment of the present application, the MSHR item of the second request, i.e., the target register item, can be directly reused, without the need to allocate a new MSHR item for the third request.

[0068] The following takes the first request as an acquisition request and the second request as a pre-fetch request as an example to illustrate the caching method provided in the embodiment of the present application.

[0069] In an optional embodiment of the present application, the first request includes a fetch request, and the second request includes a prefetch request; and in step 103, the miss status processing register generates a processing task according to the third request and sends the processing task to the main pipeline, including:

[0070] Step S21: the miss status processing register obtains a target data block from a downstream node according to the second request information in the third request;

[0071] Step S22: The miss status processing register generates a processing task according to the third request, and sends the processing task to the main pipeline when receiving the target data block returned by the downstream node.

[0072] Prefetching refers to loading data from main memory into the cache in advance to reduce the time the processor waits for data and improve cache hit rates. When the processor can predict data that may be needed in the future, it can start loading the data before it is actually needed, thereby improving overall performance.

[0073] In an embodiment of the present application, if the second request is a prefetch request, the MSHR needs to prefetch the target data block corresponding to the second request information into the cache from a downstream node, such as a main memory or a lower-level cache.

[0074] Taking the L2 cache as an example, a prefetch bit is used in the L2 directory to record whether a cache block is a prefetched block. When the MSHR receives an Acquire request, if the request address does not match the L2 or matches a prefetched block, the MSHR initiates a "trigger prefetch" request to the prefetcher. The prefetcher adds the optimal offset trained using the best-offset algorithm to the request address to generate a prefetch address. The prefetcher then sends the prefetch request to the cache bank where the prefetch address is located, along with the request type bit (Intent). The MSHR assigns an MSHR entry to the prefetch request. If the prefetched block is not in the L2, the MSHR is responsible for fetching the prefetched block from the L3 to the L2. When the MSHR completes a prefetch, it sends a response to the prefetcher, which uses this response to train the best-offset algorithm.

[0075] Next, the MSHR generates a processing task based on the third request and, upon receiving the target data block returned by the downstream node, sends the processing task to the main pipeline. After receiving the processing task, the main pipeline accesses the cache line and performs read and write operations. For example, it writes the target data block to the cache line and reads the written target data block to generate a processing result, which may include the target data block. The response buffer generates a first response to the acquisition request and returns the acquired data to the first node through the first response.

[0076] Optionally, in step 104, the main pipeline accesses the cache line of the cache according to the processing task and generates a processing result, including:

[0077] Step S31: The main pipeline determines a first cache line from the cache, and writes the target data block into the first cache line;

[0078] Step S32: The main pipeline updates the state of the target data block to a first state, where the first state is used to indicate that the acquisition request has been completed.

[0079] The main pipeline is responsible for reading and writing cache lines. Specifically, in an embodiment of the present application, when the first request is an acquisition request and the second request is a prefetch request, the main pipeline can determine a first cache line in the cache after receiving the processing task sent by MSHR, and write the prefetched target data block to the first cache line. It can be understood that when the cache is not full, the first cache line can be any unoccupied cache line. When the cache is full, the first cache line can be a cache line to be replaced determined based on a cache replacement algorithm. For example, the first cache line is the cache line that has been least recently accessed.

[0080] It should be noted that since the acquisition request and the prefetch request in the embodiment of the present application have the same request address, it can be considered that the data to be acquired by the acquisition request is the target data block to be prefetched by the prefetch request. Therefore, after the mainstream pipeline writes the target data block into the first cache line, the status of the target data block can be directly updated to the first status to indicate that the acquisition request has been completed.

[0081] In an embodiment of the present application, the acquisition request does not need to wait for the completion of the prefetch request before entering the mainstream pipeline for processing. Instead, it is merged with the prefetch request in the MSHR and enters the mainstream pipeline together. After the mainstream pipeline completes the data backfill, it can obtain the data required for the acquisition request and then respond to the acquisition request and the prefetch request at the same time, reducing the waiting time from the issuance of the acquisition request to the completion of the response, speeding up the response speed of the acquisition request, and improving the processor's memory access efficiency to the cache.

[0082] Optionally, in step S31, the main pipeline determines a first cache line from the cache and writes the target data block into the first cache line, including:

[0083] Sub-step S311: the main pipeline sends a read request to the directory of the cache to read the status information of each cache line in the cache, and determines a first cache line from each cache line of the cache using a cache replacement policy;

[0084] Sub-step S312: The main pipeline writes the target data block into the first cache line.

[0085] When the cache is full, the main pipeline needs to first select a first cache line to be replaced based on the cache replacement policy, and then backfill the target data block into the first cache line.

[0086] Specifically, the main pipeline can send a read request to the directory of the cache layer to read the status information of each cache line in the cache. The directory is a structure in the cache that stores the status of each cache line (for example, whether it is valid, whether it has been modified, replacement status, etc.) and the address tag. Each request to access the cache must read the directory to check whether the corresponding data block is in the cache. If not, the cache line to be replaced is selected based on the replacement status, and the data block is obtained from the downstream node and filled into the cache line to be replaced.

[0087] Based on the read status information, the mainstream pipeline adopts a cache replacement strategy, such as the PLRU (Pseudo Least Recently Used) replacement algorithm, to select the cache line that has been least recently accessed as the first cache line, and then write the target data block into the first cache line.

[0088] Among them, the PLRU algorithm uses a bit vector, each bit of which corresponds to a cache line in the cache set. Whenever a cache access occurs, the PLRU algorithm will update the corresponding bit vector based on the access situation to determine which cache line is the least recently used. The implementation logic usually includes: 1) Initialization: For each cache set, initialize a bit vector, and each bit in it is initialized to 0. 2) Access update: Whenever a cache line is accessed, update the corresponding bit based on the specific access situation. If the cache line is accessed, set the corresponding bit to 1, indicating that it has been used recently. If the cache line is replaced, set the corresponding bit to 0, indicating that it has not been used recently. 3) Replacement decision: When a cache line needs to be replaced, select the cache line corresponding to the bit with a value of 0 in the bit vector for replacement. In this way, the cache line that has been used least recently is selected.

[0089] Of course, in an embodiment of the present application, other cache replacement strategies may also be used to determine the first cache line, for example, an LRU (Least Recently Used) algorithm, an MRU (Most Recently Used) algorithm, an LFU (Least-Frequently Used) algorithm, and the like.

[0090] Optionally, the method further includes:

[0091] The main pipeline writes the original data in the first cache line back to the downstream node.

[0092] Before the main pipeline writes the target data block into the first cache line, the original data in the first cache line may be written back to the downstream node to prevent data loss.

[0093] The following uses the L2 cache as an example to illustrate the caching method provided in the embodiment of the present application.

[0094] In the L2 cache, the Acquire request of the upstream node will enter the Cache from the SinkA channel, and then determine in the request buffer (RequestBuffer) whether the request needs to be blocked. In an embodiment of the present application, the RequestBuffer can obtain the information of each MSHR item in the MSHR, and if blocking is required, determine the merge condition of the Acquire request. If the merge condition is met, the corresponding Acquire request will be passed to the MSHR item with the same address, and the item will be marked mergeA, and a new series of request status information will be added to include the contents of two requests, the Acquire request and the Prefetch request recorded in the MSHR item. When the downstream node returns the pre-fetched target data block, it wakes up the MSHR item, and the MSHR generates a processing task and enters the main pipeline (MainPipe) for processing. At this time, the main pipeline will select the first cache line and write the new data to the corresponding position of the first cache line, and the status of the target data block will be updated to the state it should be after the Acquire request is processed. The main pipeline then enters the response buffer (GrantBuffer) based on the processing results of the processing task, and the GrantBuffer is responsible for processing the responses to the two requests. For Prefetch requests, L2 needs to return a prefetch response to the upstream node that issued the prefetch, and can return it directly; for Acquire requests, L2 needs to return data and a response to the upstream node that issued the Acquire, and can respond sequentially through the response queue (grantQueue).

[0095] In summary, the embodiment of the present application provides a cache method. When there is a second request in the miss status processing register whose request address is the same as the first request address of the received first request, by merging the first request and the second request recorded in the MSHR into a third request in the MSHR, the situation where the request is blocked at the cache entry is reduced, so that other subsequent requests can enter the cache smoothly, which is beneficial to improving the overall performance of the cache. In addition, in the embodiment of the present application, the MSHR generates a processing task based on the merged third request and sends the processing task to the mainstream pipeline, that is, it only needs to enter the mainstream pipeline once, and the mainstream pipeline only needs to access the cache once based on the processing task, thereby reducing the pressure on the mainstream pipeline and the number of cache reads and writes, improving the first node's memory access efficiency to the cache, and helping to improve the overall performance of the computer system.

[0096] It should be noted that for the method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present application are not limited by the order of the actions described, because according to the embodiments of the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present application.

[0097] Device embodiment

[0098] 2 , there is shown a block diagram of a cache structure of the present application, wherein the cache includes a request buffer, a miss status processing register, a main pipeline, and a response buffer;

[0099] The request buffer is configured to receive a first request sent by a first node, the first request carrying a first request address; if a second request having the same request address as the first request address exists in the miss status processing register, transfer the first request to the miss status processing register, and merge the first request and the second request to obtain a third request; the third request carrying first request information of the first request and second request information of the second request;

[0100] The miss status processing register is used to generate a processing task according to the third request and send the processing task to the main pipeline;

[0101] The main pipeline is used to access the cache line of the cache according to the processing task and generate a processing result;

[0102] The response buffer is used to generate a first response to the first request and a second response to the second request according to the processing result, and send the first response and the second response to the first node in parallel.

[0103] Optionally, the request buffer is specifically used to:

[0104] querying, according to the first request address, whether a target register entry exists in a miss status processing register, the target register entry being used to record second request information of a second request, the second request address of the second request being the same as the first request address;

[0105] In a case where the target register entry exists in the miss status processing register, the first request information of the first request is added to the target register entry to generate a third request.

[0106] Optionally, the first request includes an acquisition request, and the second request includes a prefetch request; the miss status processing register is specifically used to:

[0107] Obtaining a target data block from a downstream node according to the second request information in the third request;

[0108] A processing task is generated according to the third request, and upon receiving the target data block returned by the downstream node, the processing task is sent to the main pipeline.

[0109] Optionally, the main pipeline is specifically used for:

[0110] Determining a first cache line from the cache, and writing the target data block into the first cache line;

[0111] The state of the target data block is updated to a first state, where the first state is used to indicate that the acquisition request has been completed.

[0112] Optionally, the main pipeline is specifically used for:

[0113] Sending a read request to a directory of the cache to read status information of each cache line in the cache, and determining a first cache line from each cache line of the cache using a cache replacement strategy;

[0114] The target data block is written into the first cache line.

[0115] Optionally, the main pipeline is further used to:

[0116] The original data in the first cache line is written back to the downstream node.

[0117] Optionally, the request buffer is specifically used to:

[0118] The first request information and the second request information are written into a new miss status processing register to obtain a third request, and the miss status processing register originally occupied by the second request is released.

[0119] Optionally, the first response is an acquisition response carrying acquired data, and the first response is used to indicate that data acquisition is complete;

[0120] The second response is a prefetch response, and the second response is used to indicate that the prefetch is completed.

[0121] Optionally, the main pipeline is specifically configured to: select, based on a pseudo least recently used replacement algorithm, a cache line that has been least recently accessed from each cache line of the cache as the first cache line.

[0122] In summary, the embodiment of the present application provides a high-speed cache. When there is a second request in the miss status processing register whose request address is the same as the first request address of the received first request, by merging the first request and the second request recorded in the MSHR into a third request in the MSHR, the situation where the request is blocked at the cache entry is reduced, so that other subsequent requests can enter the cache smoothly, which is beneficial to improving the overall performance of the cache. In addition, in the embodiment of the present application, the MSHR generates a processing task based on the merged third request and sends the processing task to the mainstream pipeline, that is, it only needs to enter the mainstream pipeline once, and the mainstream pipeline only needs to access the cache once based on the processing task, thereby reducing the pressure on the mainstream pipeline and the number of cache reads and writes, improving the first node's memory access efficiency to the cache, and helping to improve the overall performance of the computer system.

[0123] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0124] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0125] Regarding the processor in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method and will not be elaborated here.

[0126] Referring to Figure 3, which is a block diagram of an electronic device provided in an embodiment of the present application, the electronic device includes: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface communicate with each other via the communication bus. The memory is used to store executable instructions that enable the processor to execute the caching method of the aforementioned embodiment.

[0127] The processor may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable devices, transistor logic devices, hardware components, or any combination thereof. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0128] The communication bus may include a path for transmitting information between the memory and the communication interface. The communication bus may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, for example. The communication bus may be categorized as an address bus, a data bus, a control bus, etc. For ease of illustration, FIG3 shows only one line, but this does not imply that there is only one bus or only one type of bus.

[0129] The memory may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only), a CD-ROM (Compact Disa Read Only), a magnetic tape, a floppy disk, an optical data storage device, etc.

[0130] An embodiment of the present application also provides a non-temporary computer-readable storage medium, which, when the instructions in the storage medium are executed by a processor of an electronic device (server or terminal), enables the processor to execute the caching method shown in Figure 1.

[0131] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0132] The various component embodiments of the present application can be implemented in hardware, or in a software module running on one or more processors, or in a combination thereof. It will be appreciated by those skilled in the art that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the computing processing equipment according to the embodiment of the present application. The application can also be implemented as a device or apparatus program (for example, a computer program and a computer program product) for performing a part or all of the methods described herein. Such a program implementing the present application can be stored on a computer-readable medium, or can have the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0133] For example, FIG4 illustrates a computing device that can implement the methods according to the present application. The computing device typically includes a processor 1010 and a computer program product or computer-readable medium in the form of a memory 1020. Memory 1020 can be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, a hard disk, or ROM. Memory 1020 has storage space 1030 for program code 1031 for executing any of the method steps described above. For example, storage space 1030 for program code can include individual program codes 1031 for implementing various steps in the method described above. These program codes can be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, compact disks (CDs), memory cards, or floppy disks. Such computer program products are typically portable or fixed storage units, as described with reference to FIG5 . This storage unit can have storage segments, storage space, and the like arranged similarly to memory 1020 in the computing device of FIG4 . The program code can, for example, be compressed in a suitable form. Typically, the storage unit includes computer-readable codes 1031 ′, ie, codes that can be read by a processor such as 1010 , which, when executed by a computing device, cause the computing device to perform the steps of the method described above.

[0134] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0135] References herein to "one embodiment," "an embodiment," or "one or more embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present application. Furthermore, please note that instances of the phrase "in one embodiment" do not necessarily all refer to the same embodiment.

[0136] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and techniques are not shown in detail so as not to obscure the understanding of this description.

[0137] In the claims, any reference signs placed between brackets shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application may be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by one and the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names.

[0138] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

[0139] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, devices, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0140] The present application embodiment is described with reference to the flow chart and / or block diagram of the method, terminal device (system), and computer program product according to the embodiment of the present application. It should be understood that each process and / or box in the flow chart and / or block diagram and the combination of the process and / or box in the flow chart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device produce a device for realizing the function specified in one process or multiple processes and / or one box or multiple boxes of the flow chart.

[0141] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to operate in a predictable manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0142] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce computer-implemented processing, so that the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0143] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0144] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.

[0145] The above is a detailed introduction to a cache method, cache, electronic device and readable storage medium provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A cache method, applied to a high-speed cache, wherein the high-speed cache comprises a request buffer, a miss status processing register, a main pipeline and a response buffer; the method comprises: The request buffer receives a first request sent by a first node, where the first request carries a first request address; When the request buffer has a second request whose request address is the same as the first request address in the miss status processing register, the first request is transferred to the miss status processing register, and the first request and the second request are combined to obtain a third request; The third request carries first request information of the first request and second request information of the second request; The miss status processing register generates a processing task according to the third request, and sends the processing task to the main pipeline; The main pipeline accesses the cache line of the cache according to the processing task and generates a processing result; The response buffer generates a first response to the first request and a second response to the second request according to the processing result, and sends the first response and the second response to the first node in parallel.

2. The method according to claim 1, wherein: When there is a second request whose request address is the same as the first request address in the miss status processing register, the request buffer transfers the first request to the miss status processing register, and merges the first request and the second request to obtain a third request, including: The request buffer queries whether there is a target register item in the miss status processing register according to the first request address, the target register item is used to record second request information of the second request, and the second request address of the second request is the same as the first request address; When the target register entry exists in the miss status processing register, the request buffer adds the first request information of the first request to the target register entry to generate a third request.

3. The method according to claim 1, wherein: The first request comprises a get request, and the second request comprises a pre-fetch request; The miss status processing register generates a processing task according to the third request, and sends the processing task to the main pipeline, including: The miss status processing register acquires a target data block from a downstream node according to second request information in the third request; The miss status processing register generates a processing task according to the third request, and sends the processing task to the main pipeline when receiving the target data block returned by the downstream node.

4. The method according to claim 3, wherein: The main pipeline accesses the cache line of the cache according to the processing task and generates a processing result, including: The main pipeline determines a first cache line from the cache and writes the target data block into the first cache line; The main pipeline updates the state of the target data block to a first state, where the first state is used to indicate that the acquisition request has been completed.

5. The method according to claim 4, wherein: The main pipeline determines a first cache line from the cache and writes the target data block into the first cache line, including: The main pipeline sends a read request to a directory of the cache to read status information of each cache line in the cache, and determines a first cache line from each cache line of the cache by using a cache replacement strategy; The main pipeline writes the target data block into the first cache line.

6. The method according to claim 4, wherein: The method further comprises: The main pipeline writes the original data in the first cache line back to the downstream node.

7. The method according to claim 1, wherein: The combining the first request and the second request to obtain a third request includes: The first request information and the second request information are written into a new miss status processing register to obtain a third request, and the miss status processing register originally occupied by the second request is released.

8. The method according to claim 3, wherein: The first response is an acquisition response carrying acquired data, and the first response is used to indicate that data acquisition is completed; The second response is a pre-fetch response, and the second response is used to indicate that the pre-fetch is completed.

9. The method according to claim 5, wherein: The step of determining a first cache line from cache lines of the cache using a cache replacement strategy includes: Based on a pseudo least recently used replacement algorithm, a cache line that has been least recently accessed is selected from the cache lines of the cache as the first cache line.

10. A cache comprising a request buffer, a miss status processing register, a main pipeline and a response buffer; The request buffer is configured to receive a first request sent by a first node, wherein the first request carries a first request address; if a second request having a request address identical to the first request address exists in the miss status processing register, transfer the first request to the miss status processing register, and merge the first request and the second request to obtain a third request; The third request carries first request information of the first request and second request information of the second request; The miss status processing register is used to generate a processing task according to the third request and send the processing task to the main pipeline; The main pipeline is used to access the cache line of the cache according to the processing task and generate a processing result; The response buffer is used to generate a first response to the first request and a second response to the second request according to the processing result, and send the first response and the second response to the first node in parallel.

11. The cache of claim 10, wherein: The request buffer is specifically used for: querying whether there is a target register entry in the miss status processing register according to the first request address, the target register entry being used to record second request information of the second request, the second request address of the second request being the same as the first request address; In a case where the target register entry exists in the miss status processing register, the first request information of the first request is added to the target register entry to generate a third request.

12. The cache of claim 10, wherein: The first request includes an acquisition request, and the second request includes a pre-fetch request; the miss status processing register is specifically used for: Acquire a target data block from a downstream node according to the second request information in the third request; A processing task is generated according to the third request, and upon receiving the target data block returned by the downstream node, the processing task is sent to the main pipeline.

13. The cache of claim 12, wherein: The main pipeline is specifically used for: Determine a first cache line from the cache, and write the target data block into the first cache line; The state of the target data block is updated to a first state, where the first state is used to indicate that the acquisition request has been completed.

14. The cache of claim 13, wherein: The main pipeline is specifically used for: Sending a read request to a directory of the cache to read status information of each cache line in the cache, and determining a first cache line from each cache line of the cache by using a cache replacement strategy; The target data block is written into the first cache line.

15. The cache of claim 13, wherein: The main pipeline is also used for: The original data in the first cache line is written back to the downstream node.

16. The cache of claim 10, wherein: The request buffer is specifically used for: The first request information and the second request information are written into a new miss status processing register to obtain a third request, and the miss status processing register originally occupied by the second request is released.

17. The cache of claim 14, wherein: The main pipeline is specifically used to select, based on a pseudo least recently used replacement algorithm, a cache line that has been least recently accessed from each cache line of the cache as the first cache line.

18. An electronic device, comprising a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; the memory is used to store executable instructions, and the executable instructions enable the processor to execute the caching method as described in any one of claims 1 to 9.

19. A computer program, comprising computer readable codes, which, when executed on a computing and processing device, cause the computing and processing device to execute the caching method according to any one of claims 1 to 9.

20. A readable storage medium storing the computer program according to claim 19, wherein when instructions in the readable storage medium are executed by a processor of an electronic device, the processor is enabled to execute the cache method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Method and device for holding cache miss states of caches in processor of computer

    CN103399824A

  • Method and device for accessing fine-grained cache by GPU (Graphics Processing Unit) multi-granularity memory access request

    CN114661352A

  • Address translation method executed by processor and related product

    CN116860665A

  • Caching method, cache, electronic equipment and readable storage medium

    CN117609110A

  • Prefetcher with multi-cache level prefetches and feedback architecture

    WO2023287512A1

Cited By

  • Cache line data multiplexing method, electronic equipment and storage medium

    CN121051039A

  • Cache performance detection method and device and storage medium

    CN121434037A

  • Data processing method and device, electronic equipment, storage medium and program product

    CN121833554A

  • Cache operation method, cache, computing device and system

    CN121958146A

  • Cache device, operation method, electronic equipment and artificial intelligence processor

    CN121996574A